Bringing Existing Models into Sparx EA

Models in four places, none of them authoritative

This engagement took us to an organisation in the environmental services sector — waste collection, recycling and related public duties for a region — whose architecture knowledge had accumulated in four places over roughly a decade. The oldest and largest was a repository in a modelling tool whose vendor had stopped meaningful development years earlier; licences were still being paid, expertise was down to two people, and every year the export options looked a little more like a trap. Alongside it lived several hundred Visio drawings on a network share, an application inventory in Excel maintained by the service management team, and a younger collection of ArchiMate models that a recently arrived architect had started in Sparx EA because she could not face the old tool.

The four sources disagreed with each other, and everyone knew it. The same application appeared under three names; integrations existed in one source and not another; and when a question mattered, people phoned the two veterans rather than opening any of the models. The organisation had already chosen Sparx EA as the target — the new models were there, and the platform decision had been taken with procurement the year before. What they asked us for was the move itself: plan a migration, prove it would preserve what mattered, and leave them with one repository they could finally call authoritative.

We have written before about how we approach model migrations in general. This case is worth telling separately because of where the effort actually went: not into moving data, which is the easy half, but into deciding what the data meant and proving the move had not quietly changed it.

Taking stock of what actually existed

The first weeks were an inventory, and the inventory was itself a small automation exercise. The legacy tool could export its repository to XML, and we wrote scripts to profile the export rather than reading it by eye: counts by object type, counts by relationship type, properties in use, diagrams and their element references, and — importantly — last-modified dates. The profile put real numbers under the anecdotes. The repository held around five thousand elements and nine thousand relationships across some three hundred diagrams, but nearly half the elements had not been touched in five years, and whole branches belonged to programmes that had finished or been cancelled.

The Visio share was harder to profile and easier to judge. Sampling showed that most drawings were either presentation copies of diagrams that existed in the legacy tool or one-off sketches for meetings long past. We agreed early that Visio content would not be migrated mechanically: anything still valuable would come across because a named person claimed it and a modeller rebuilt it. Of the several hundred files, fewer than thirty were claimed, and none of the rest has been asked for since.

The Excel inventory, by contrast, turned out to be the most trustworthy source in the building — it was the only artefact with an owner who updated it on a schedule, because service management depended on it. That settled an argument about precedence: where the legacy repository and the inventory disagreed about an application's existence or name, the inventory won, and the repository entry was flagged for the mapping stage rather than imported as-is.

The inventory phase ended with a scoping decision the client made, on our recommendation, with the numbers in front of them: migrate the living half of the legacy repository, archive the rest as a read-only export kept for reference — with the split proposed by the profile and confirmed by people: every branch proposed for archiving was reviewed by the veterans or its former programme's owner, and anything still referenced by a living element came across regardless of its age — and treat the Excel inventory as the seed of the application catalogue in the target. Migrating everything would have been simpler to announce and worse to live with — dead elements do not become alive by changing tools, they just make the new repository as hard to trust as the old one.

Mapping source concepts to a target metamodel

The legacy tool had its own metamodel: around forty object types and a generous set of relationship types, used with the inconsistency that ten years and many hands produce. The target was Sparx EA's ArchiMate implementation, with a deliberately small set of extensions. Between the two sat the piece of work that determined everything downstream: the mapping table.

We built the mapping in working sessions with the two veterans and the new architect — the people who knew what the source types had actually been used for, as opposed to what their names suggested. Each source type got one of four rulings. Some mapped cleanly: the source's application concept to ArchiMate application component, its interface concept to application interface, and so on. Some mapped with a narrowing, where one source type had been used for several things and needed splitting by property or by package origin. Some had no ArchiMate equivalent worth pretending about — the source's matrix-style organisational links were the clearest case — and were carried across as stereotyped elements in a small extension profile, honestly labelled rather than forced into a near-miss ArchiMate concept. And some were ruled out of scope entirely, recorded with a reason.

Figure 1: The four source collections on the left, the mapping rules in the middle, and the target Sparx EA repository structure on the right, with provenance recorded on every migrated element
Figure 1: The four source collections on the left, the mapping rules in the middle, and the target Sparx EA repository structure on the right, with provenance recorded on every migrated element

Two conventions from this stage paid for themselves many times over. First, provenance: every migrated element carries tagged values recording its source identifier, source type and migration run, which made every later question of the form "where did this come from?" answerable in seconds, and made re-runs idempotent — an element is matched by source identifier and updated, never duplicated, while relationships and the generated aggregations carried the same run provenance and were replaced as a set, and content removed by a mapping change surfaced in the drop log rather than lingering silently. Second, the mapping table itself was data, not documentation: a spreadsheet the transformation scripts read directly, so the ruling the workshop made was, by construction, the ruling the migration applied. When a mapping changed, the change was one row, visible in version control, and the next run applied it everywhere. The same machinery handled consolidation across sources: where one application existed in the legacy repository, the inventory and the new architect's models, the match agreed in the naming workshops was recorded as additional source identifiers on a single canonical element, so an import from any source updated that element rather than spawning a twin, and the transformation redirected incoming relationships to it.

A pilot migration before any promises

Before committing to dates, we migrated one domain end to end: the collection-logistics domain, chosen because it was mid-sized, actively maintained, and owned by people willing to look hard at the result. The pilot existed to answer three questions. Could the pipeline run clean from export to import? Did the mapped result read correctly to the people who knew the content? And what would reviewing and correcting a domain actually cost, in hours, so the full plan rested on a measured number rather than an estimate defended in a meeting?

The pilot's answers reshaped the plan. The pipeline ran, but the review found a systematic problem no profile had caught: the source's habit of using diagram-level containment to imply ownership relationships that were never created as data. On the legacy diagrams, an application drawn inside a domain box belonged to that domain — visually, and only visually. Migrating the data faithfully would have migrated the loss of that information. We extended the transformation to read diagram containment from the export and generate explicit aggregation relationships, flagged with their own provenance value so nobody would later mistake them for deliberate modelling. The pilot review also priced the human side: about two hours of owner review per hundred elements, a number that held well enough across the full migration to plan by.

The migration pipeline and its checks

The full migration ran as a pipeline with five stages, each writing its output to disk so any stage could be re-run without repeating the ones before it: export from the legacy tool, profile, transform against the mapping table, import into a staging area in Sparx EA, and verify. The import used the automation interface rather than XMI — we wanted element-by-element control, provenance written as we went, and stable handling of re-runs, and the automation API gives exactly that at the cost of speed. A full run took a little under two hours, which for a one-time migration with re-runs was a price worth paying for the control.

Figure 2: The migration pipeline from legacy export to verified import, with the two gates at which a run stops: reconciliation of counts after transformation, and owner sign-off after staging review
Figure 2: The migration pipeline from legacy export to verified import, with the two gates at which a run stops: reconciliation of counts after transformation, and owner sign-off after staging review

Staging mattered more than it sounds. Migrated content landed in a staging package structure, separate from the live models the new architect had been building, and only moved to its final home after review. This kept the growing repository honest — nothing unreviewed sat next to trusted content — and it gave the review a clear unit: a domain moved when its owner signed off, not before. The verification stage produced a report per run: elements in by type, relationships in by type, elements updated versus created, generated relationships, and every source object that produced no target object, each with the mapping rule that said so.

Accounting for every relationship

The rule we held ourselves to, and the one we would recommend to anyone doing this, is that every source relationship must end the migration in exactly one of three states: migrated, transformed under a named rule, or dropped under a named rule. The verification report reconciled the arithmetic every run — source relationships in, target relationships out, transformations and drops itemised — and a run whose numbers did not reconcile was a failed run, full stop, regardless of how plausible the repository looked when you opened it.

This sounds bureaucratic and was, in practice, the cheapest insurance in the project. Twice during the domain migrations the reconciliation caught real defects — once a transformation rule that silently swallowed relationships whose endpoints had both been narrowed, once an export quirk that truncated a relationship table mid-file — caught not by the transform arithmetic, which can only balance what the export contains, but by the profile stage's independent cross-check of export counts against the legacy tool's own database. Both would have produced a repository that looked complete and was not, which is the failure mode people discover a year later, one missing dependency at a time, at which point trust in the new repository dies for reasons nobody can name. A migration is not just moving content; it is manufacturing the right to say "this is all of it", and the reconciliation is where that right comes from.

Counts are necessary but not sufficient. The reconciliation proves the pipeline lost nothing between export and import — the completeness of the export itself was the profile's cross-check against the source database; only the owner reviews prove the content still means what it meant. We never let a green report substitute for a person who knows the domain saying "yes, that is our landscape".

The toolkit behind the pipeline

Nothing in the toolkit was exotic, and that was a design goal: the client's own team had to be able to read, re-run and eventually retire every piece of it. The export profiling and transformation stages were Python scripts working over the legacy XML, with the mapping workbook read directly as their rule set. The import stage drove Sparx EA through its COM automation interface, writing elements, relationships, tagged values and staged diagrams, with interactive modelling clients closed during write runs — the automation session runs its own EA instance, and the discipline is about keeping human sessions out of the repository while scripts write; worth stating in the runbook rather than learning the hard way. The target repository sat on PostgreSQL from day one, sized for the organisation's future rather than the migration; our reasoning on that choice matches what we have written about database choices for Sparx EA repositories.

Every script, the mapping workbook and the run reports lived in the client's version control, and each verification report was archived with the run that produced it. That archive quietly became the migration's audit trail: when a question came months later about why a particular relationship type had been dropped, the answer was a mapping row with a date, an author and a workshop minute behind it. The stages, their outputs and their gates fit on one slide, which is roughly the test a migration plan should pass:

StageOutput keptGate
ExportSource XML, dated—
ProfileCounts and property usage report—
TransformImport-ready dataset plus drop logReconciliation must balance
Import to stagingStaged packages with provenance—
Verify and reviewRun report; owner sign-off recordOwner signs before promotion

Diagram fidelity, the honest version

Diagrams are where migration promises go to die, so we made a narrow one: data completely, diagrams selectively, layout approximately. The legacy export carried diagram composition — which elements appear on which diagram — and, for most diagram types, usable coordinates. It did not carry the visual grammar: colours with meaning, line routing, the careful geometry a modeller spends hours on. Regenerating every diagram would have produced three hundred technically accurate, uniformly ugly pictures, and we have seen what that does to adoption — people judge a repository by its diagrams long before they query its data.

Instead, the owners nominated the diagrams that mattered: about sixty of the three hundred, mostly domain landscapes and the integration views used in operational discussions. For those, the pipeline generated first-draft diagrams in Sparx EA from the exported composition and coordinates, and a modeller — budgeted at roughly an hour per diagram, funded by the project — brought each to the organisation's new diagram conventions. Each finished diagram went through a side-by-side review against a PDF snapshot of its legacy original, with the owner initialling the comparison. The remaining diagrams were archived as PDFs alongside the read-only export, findable but frozen. Nobody has asked for one to be resurrected, which we take as confirmation that the sixty were the right sixty — and that the coordinates-and-composition draft was the right level of automation to aim for, doing the tedious part and leaving the judgement to a person.

Cutover, freeze and the first weeks

The migration ran domain by domain over about four months, but there was still a cutover: the day the legacy repository stopped accepting changes. We negotiated a two-week change freeze on the source for the final reconciliation runs — long enough to re-export, re-run and re-verify every domain against a stable source, short enough that the veterans could grant it without a governance fight. The final runs followed the same staging path as every other run: each produced a per-domain delta report, domains with empty deltas promoted automatically, and the few with late changes went back to their owners for a short second sign-off rather than being written over promoted content. At the end of the freeze the legacy tool went read-only for everyone, with its licences kept for one further year purely as an archive viewer, and the Sparx EA repository became, formally and in a message from the CIO, the authoritative source.

The first weeks after cutover were deliberately over-staffed with support. A migration's reputation is set early: the first person who cannot find something they swear existed will either get an answer in an hour — with the provenance tagged values, usually a query away — or start telling colleagues the migration lost things. Most of the early questions traced to the scoping decision, not the pipeline: content people remembered was in the archived half, findable in the read-only export. Having a crisp answer for that case, including where the archive lived and how to request a resurrection, mattered as much as the pipeline's correctness. Three resurrections were requested in the first quarter; each took under a day, because the pipeline could migrate a single package on demand.

Communication ran on a simple rhythm throughout: a one-page status after every domain promotion, sent to the owners and the CIO, listing what had moved, what had been dropped under which rules, and what was next. Nothing in it was news to anyone involved in that domain — which was the point. Migrations lose their mandate through surprise more often than through error, and the cheap discipline of never letting a stakeholder learn something about their own domain second-hand kept the mandate intact for the four months the move needed.

What the organisation gained

A year on, the organisation runs a single repository with a little over three thousand living elements, package structure by domain, and conventions the new architect now owns. The Excel inventory has been retired — its owner signs off changes in the repository instead, which was the real test of authority — and service management, who never opened the legacy tool, read the repository through generated documents and views. The two veterans, whose knowledge the organisation had been one retirement away from losing, spent the migration turning what they knew into mapping rules and review comments; a good part of what lived in their heads now lives in data with provenance. And the architecture team makes its case for investment decisions from a repository it can defend, line by line, back to source — something the four disagreeing sources never allowed. The handover was gradual rather than ceremonial: the new architect ran the last two domain migrations herself with us reviewing, the veterans wrote the archive access guide, and the final deliverable was less a report than a working team that had already operated everything without us for a month. If you are weighing a similar consolidation, our Sparx EA consulting practice does this work often enough to have opinions on where it goes wrong.

What we would plan differently

Two things, honestly held. First, we scheduled the Visio triage as a side activity and it dragged for months, because "claim it or lose it" only works with a deadline someone enforces; we would now put a named end date in the plan and enforce it. Second, we under-planned the naming reconciliation between the legacy repository and the Excel inventory. Matching the repository's application population against a thousand inventory rows by name variants consumed more workshop hours than the metamodel mapping itself, and we would now start that matching in week one with fuzzy-match tooling and a standing weekly session, rather than treating it as a detail the mapping table would absorb.

One limitation deserves naming, because it is structural rather than a planning miss: a migration like this preserves information, but it cannot create authority. The repository became authoritative because the CIO said so, the inventory owner moved her process into it, and the archive stayed genuinely available — decisions taken around the tooling, not in it. The pipeline's contribution was narrower and still essential: it made those decisions safe to take, by ensuring the thing being declared authoritative deserved it.

If your organisation is looking at a repository it no longer trusts, in a tool it no longer wants, and wondering what a defensible move to Sparx EA looks like — we are happy to talk through what your version of this would involve, via our contact page.

This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.