From Scattered Diagrams to a Shared Architecture Repository

The situation we walked into

The organisation was a public administration with a few thousand staff, a portfolio of citizen-facing services, and an IT department that had grown by accretion over two decades. Architecture existed — that was never the problem. It existed in Visio files on a departmental share, in PowerPoint decks attached to steering committee minutes, in a handful of personal Sparx EA project files that individual architects kept on their own machines, and in the heads of three people who had been there longest.

The trigger was a transformation programme. The administration was consolidating several legacy service portals into a single digital counter for citizens, and the programme board asked a reasonable question: which applications does each service depend on today? Nobody could answer it from documentation. The answer that came back was assembled by hand in two weeks, was contested in the meeting where it was presented, and turned out later to have missed two systems. An internal audit the same year made a similar observation in more formal language: architectural documentation was fragmented, unversioned and person-dependent.

By the time we were engaged, the organisation had already decided that Sparx EA would be the repository — a decision partly made for them, since four of their six architects were already using it individually. What they asked us for was the part between the licence purchase and the working practice: one repository, a structure everyone could navigate, conventions that would survive contact with real projects, and a migration of the content worth keeping.

We have seen this starting point often enough at Sparx EA consulting engagements to know that the tooling is the easy half. The difficult half is deciding what the repository is for, because that decision determines what goes in, what stays out, and what "finished" looks like. We put that question on the table in the first week rather than the last.

What the inventory told us

We started with an inventory of every architectural artefact we could find, and we deliberately cast a wide net: the departmental shares, the document management system, the project archive, and the personal EA files. It took about three weeks, working alongside one of their architects who knew where the bodies were buried. The result was a spreadsheet of around four hundred artefacts, each with an owner if we could establish one, a date, a subject, and our judgement on whether it was current.

The judgement column was the sobering one. Perhaps a quarter of the artefacts were current enough to trust. Half were plausibly out of date — diagrams of systems that had since been replaced, process maps from a reorganisation two structures ago. The rest were duplicates, drafts, or slide-deck copies of other diagrams that had drifted from their source. The three personal EA project files overlapped heavily: the same applications appeared in all of them, under slightly different names, at different levels of detail, with different ideas about what an "application" was.

That last finding shaped the whole engagement. The organisation did not primarily have a storage problem, which a shared repository would have solved on its own. It had an agreement problem: no shared list of applications, no shared naming, no shared understanding of granularity. Consolidating the files without settling those questions would have produced one repository containing three architectures.

We presented the inventory to the architecture group with a simple proposal: treat the migration as an editorial exercise, not a lift-and-shift. Everything worth keeping would be re-homed deliberately, against an agreed application list and an agreed structure. Everything else would be archived where it was, findable but clearly retired. Nobody mourned the four hundred artefacts once they saw the duplication laid out in one table.

Shaping the target repository

The technical shape of the target was settled early. A file-based EA project on a share would have reproduced the old problems with a new file extension, so we set up a central repository on PostgreSQL, fronted by Pro Cloud Server, with the desktop clients connecting over HTTPS rather than a direct database connection. That choice also opened the door to WebEA later, which mattered because most of the people who needed to read the architecture were never going to install a modelling tool. We have written elsewhere about choosing a database for Sparx EA repositories; here PostgreSQL was the pragmatic answer, since the administration already ran it and could back it up with their standard tooling.

Figure 1: The consolidated target: one PostgreSQL-backed repository served through Pro Cloud Server to modelling clients and WebEA readers
Figure 1: The consolidated target: one PostgreSQL-backed repository served through Pro Cloud Server to modelling clients and WebEA readers

The package structure took more discussion than the infrastructure, as it usually does. We landed on six top-level packages: a strategy and capability package, a business architecture package, an application package holding the agreed application list, a technology package, a projects package where in-flight work could be modelled without polluting the baseline, and a reference package for imported standards and the modelling conventions themselves. The important decision was the separation between the baseline packages — describing the administration as it is — and the projects package, where anything could be sketched. Content only moved from projects into the baseline through a review, which gave the baseline a meaning it had never had on the file shares: if it is here, someone has checked it.

We resisted two structures that were proposed along the way. One was organising the repository by department, which would have encoded the current org chart into the model and broken at the next reorganisation. The other was organising it by project, which is how the file shares had rotted in the first place. Structure by architectural concern, with projects quarantined, has survived every reorganisation since.

Conventions before content

Before migrating anything, we wrote the modelling conventions with the architecture group — deliberately with them, not for them, because conventions imposed from outside get politely ignored. Four half-day workshops produced a document short enough to be read: which ArchiMate concepts were in scope for the baseline, what each one meant in this organisation's terms, how elements were named, and which relationships were expected between which layers.

The granularity question got its own session, because it had caused the divergence between the personal models. What one architect had modelled as a single case-management application, another had modelled as five components. We settled it with a working definition — an application is something with its own lifecycle, budget line or support arrangement — and walked the entire draft application list against that definition. The list came out at around two hundred and thirty applications, which was more than anyone had guessed and fewer than the pessimists feared.

We kept the mechanical side of the conventions in the tool rather than the document wherever possible. A small set of tagged values carried the attributes the administration actually needed to query — lifecycle status, owning unit, hosting model, criticality — and we configured these once so architects picked them from dropdowns instead of inventing values. Naming rules and required attributes were checked by a validation script run against the repository on a schedule, an approach we described in our piece on model validation in Sparx EA. The script produced a short exception report per package owner rather than a repository-wide wall of shame, which kept the tone corrective rather than punitive.

A convention that lives only in a document is a hope. A convention that a script checks every week is a practice. The document explains why; the script remembers.

Making the conventions the easy path

A convention document competes with muscle memory, and muscle memory wins unless the tool is arranged so that the convenient action and the correct action are the same action. We spent a focused week on exactly that arrangement, and it repaid itself more visibly than almost anything else in the engagement.

The concepts in scope were packaged as a small MDG Technology: a trimmed diagram toolbox showing only the agreed element types, with the agreed tagged values attached to each stereotype so a newly created element arrives carrying empty fields that expect to be filled, rather than requiring anyone to remember that fields exist. Architects opening a new diagram saw a toolbox of perhaps fifteen entries instead of ArchiMate's full palette, which quietly ended the era of exotic element types chosen because they looked right. The approach follows what we describe in building custom diagram toolboxes with MDG, scaled to a deliberately small profile.

Alongside the toolbox went a starter kit inside the repository itself: template diagrams for the three view types the conventions blessed, a worked example package modelling one real service correctly, and a set of saved model searches — applications without owners, elements missing lifecycle status, relationships crossing layers illegally — pinned where everyone could run them. The worked example did more teaching than the conventions document; people copy what they can see. None of this is glamorous work, and all of it is the difference between conventions that are followed and conventions that are cited.

Migrating in waves

With the target structure and conventions agreed, we migrated in four waves over roughly five months, each wave small enough to review properly.

Figure 2: The four consolidation phases, from inventory and triage through structured migration waves to governed operation
Figure 2: The four consolidation phases, from inventory and triage through structured migration waves to governed operation

The first wave was the application list itself, built fresh in the repository rather than imported, because it was the reference everything else would hang from. Each application got its agreed name, its tagged values, and an owner. The second wave brought in the three personal EA models. For these we used a combination of XMI export and scripted rework over the automation interface: elements were matched against the application list, survivors were re-parented into the new structure, duplicates were merged with their relationships re-pointed, and everything that matched nothing was parked in a quarantine package for its author to defend or delete. Roughly a third of the imported elements did not survive quarantine, and nobody has asked after them since.

The third wave was the selective redrawing of the Visio and PowerPoint material. We made a deliberate choice here not to chase imports. Visio import routes exist, but the diagrams worth keeping were worth redrawing against the consolidated element set, and the diagrams not worth redrawing were not worth keeping. About forty diagrams made the cut, redrawn by the architects themselves in facilitated working sessions — which doubled as training and surfaced a steady stream of "actually, that connection no longer exists" corrections that no automated import would have caught.

The fourth wave was the long tail: interface documentation, a security zoning model, and the standards library, brought into the reference package with their sources recorded. At the end of each wave we ran the validation script, walked the exceptions with the owners, and only then declared the wave closed. The discipline of closing waves mattered more than the wave boundaries themselves; it created a rhythm of finished things in an activity that otherwise never finishes.

Keeping it shared: ownership and review

A shared repository fails in one of two ways: it becomes read-only in practice because nobody dares touch it, or it becomes a sandbox because everybody does. The governance had to thread between the two, and we kept it as light as we could get away with.

Every top-level package got a named curator — not an approval bottleneck, but the person who walks the validation exceptions and decides what enters the baseline from the projects package. Project architects model freely in their own project packages, using elements from the baseline rather than copying them, which EA makes natural once people are shown the difference between dragging an existing element onto a diagram and creating a new one. The weekly exception report goes to curators; a monthly half-hour architecture board slot handles anything contested, and most months it handles nothing, which is the sign of a governance process sized correctly.

We also switched on EA's model security for the repository, with the require-lock-to-edit policy, and put group locks on the baseline packages so that changing them is a deliberate act that goes through a curator, while project teams lock and release their own packages freely. Locking gets grumbled about for the first month, then quietly prevents overwrites for every month after. Baselines are taken on the application and technology packages before each architecture board meeting, which gives the administration something it never had on the file shares: the ability to say what the approved application and technology architecture was on a given date, and to diff it against today. Our thinking on this kind of setup owes a lot to engagements described in architecture governance work in larger programmes, scaled down to fit an architecture group of six.

Opening the repository to readers

About four months in, once the baseline was trustworthy, we switched on WebEA for readers. This changed the repository's audience from six architects to a few hundred potential readers — programme managers, service owners, security officers — and it changed behaviour in a way we did not fully predict. Once service owners could look up their own service and see which applications it depended on, they started reporting errors. Some arrived indignantly, which we counted as success: indignation means the content is being read and is expected to be right.

We tuned what readers saw rather than exposing the whole repository. Worth knowing before you plan the same: on PostgreSQL, EA's visibility levels — the database feature that hides packages from whole classes of users — are not available, since they rely on Oracle and SQL Server schema features. So the projects package was kept out of the published view the plain way: a scheduled script performs a project transfer into a separate reader repository that simply does not contain it, and WebEA points at that copy. Half-finished project sketches read as facts to anyone outside the architecture group, so the exclusion was worth the extra moving part. Diagram titles and element notes were rewritten in places for a reading audience, which was less work than it sounds because the conventions had insisted on meaningful notes from the start. The lesson we keep relearning is that publication is an editorial act: the repository serves the modellers, the published view serves everyone else, and conflating the two audiences serves neither.

What changed for the organisation

The question that triggered the engagement — which applications does each service depend on — is now answered from the repository in minutes, with a model search rather than a two-week manual exercise. The transformation programme used the dependency views in its planning rounds, and the two systems the original hand-built answer missed were both in the model, found by the redrawing sessions.

The quieter changes matter as much. There is one application list, and when a new system is procured it enters that list on day one, because the procurement checklist now says so. The personal EA files are gone; their owners were their most enthusiastic retirers once the shared repository could do what the private files could not, which is show their work to other people with the context intact. The audit finding about fragmented documentation was closed the following year, with the auditors pointed at the repository, its baselines and its validation reports rather than at a binder of exported documents.

None of this made the architecture correct by magic. It made the architecture disputable, which is the actual prerequisite for correctness: a wrong entry in a shared, read, versioned repository gets found and fixed; a wrong entry in a personal file on a laptop is wrong forever.

What it took

Organisations considering this kind of consolidation reasonably ask what it costs, so here is the honest shape of it. The engagement ran about eight months end to end, at an intensity that varied a great deal: the inventory took three weeks of one consultant working with one of their architects; the conventions phase was four workshops spread over six weeks, with writing in between; the migration waves consumed roughly two days a week of our time and a comparable commitment from their architecture group, whose members did the redrawing themselves by design. The infrastructure — PostgreSQL, Pro Cloud Server, backups, security — was stood up by their own IT operations in a few days with a checklist from us, and has needed little attention since.

The distribution of effort surprised the client, and tends to surprise everyone. The tooling and infrastructure together were perhaps a tenth of the work. The editorial work — agreeing what an application is, naming things, deciding what deserved to survive — was well over half, and it is the half that cannot be delegated to consultants, because the agreements have to be theirs. What we brought to that half was structure, pace and an outsider's licence to ask why two teams maintained two lists of the same systems. Budgets for repository consolidations usually over-provision for tooling and under-provision for editorial stamina; if this case study transmits one planning correction, let it be that one.

A rhythm observation, too: the engagement worked in waves with closed ends rather than as a continuous background activity, and this was load-bearing. Each wave finished, was reviewed, and was declared done in front of the architecture board. Consolidations run as open-ended background work tend to be ninety percent finished forever.

What we would do differently

Two things, honestly. First, we would start the WebEA publication earlier, even with an imperfect baseline. We held it back out of caution about publishing unreviewed content, and the caution was reasonable, but the error reports from readers turned out to be the fastest quality mechanism we had. A visibly imperfect repository that improves weekly builds more trust than a curtain with promises behind it.

Second, we underestimated the pull of the old habits around slide decks. For a good six months after go-live, steering papers were still being decorated with hand-drawn PowerPoint architecture, redrawn from the repository rather than exported from it, because that is what people had always done. The fix was partly technical — decent document templates and diagram exports from EA so the repository version was the path of least resistance — and partly social, in that the architecture group simply started declining to review diagrams that did not come from the model. We now raise the document-generation question in the first month of similar engagements rather than the sixth.

The named limitation of the approach: a consolidated repository concentrates risk as well as knowledge. The administration now depends on one PostgreSQL database, its backups and its access control in a way it did not when the architecture was scattered and irrecoverable. That trade is worth making — scattered and irrecoverable is not a resilience strategy — but it is a trade, and it belongs in the operations conversation, not the small print.

If your architecture lives in twelve places and none of them is trusted, this is a well-trodden path out — and the first step, the honest inventory, costs three weeks and no software. If you would like help walking it, from repository design through migration to working governance, you can reach us through our contact page.

This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.