When Archi Models Outgrow Files

Nothing is wrong with Archi

It is worth saying this first, because the rest of the article reads like criticism and is not. Archi is a genuinely good ArchiMate tool: fast, free, standards-faithful, and pleasant to model in. Architects who use it tend to like it, which is not something you can say about every tool in this category.

The problem is not the modelling. It is that a desktop application reading and writing files is not a repository, and the moment more than one person is involved you need a repository whether you planned for one or not.

Figure 1: The progression, and roughly where each stage stops working
Figure 1: The progression, and roughly where each stage stops working

Stage one: one architect

Everything works. There is one file, it is on one laptop, and the person editing it is the person who knows what is in it. If they want history they copy the file with a date in the name, and that is genuinely sufficient.

Many organisations never leave this stage, and for them none of this matters.

Stage two: two architects

The first friction is social rather than technical. "Can you send me the latest?" Someone emails a file. The other person edits it. Now there are two latests.

Teams solve this with a convention — the file lives on the share, you tell the channel before you open it — and the convention works, right up until someone forgets, or is on leave, or opens it to look at something and forgets they have it open. The failure is silent: two people save, the second save wins, and the first person's afternoon is gone with no error message.

Stage two and a half: the sync drive

Between the email era and Git, nearly every team takes the same detour: put the model on OneDrive or SharePoint, because it is there, it is backed up, and it feels like a shared repository. For architecture models it is quietly the worst option on the whole progression, and the reasons are mechanical rather than a matter of taste.

Sync clients are built for documents that one person edits at a time and that tolerate an occasional conflicted copy. An .archimate file is a single large XML blob rewritten wholesale on every save, so two architects saving within the same sync window do not get a merge or even an error — they get model-Copy(1).archimate appearing silently next to the original, each file containing one person's afternoon, with no indication of which is which. The channel convention that more or less worked on a plain network share fails harder here, because sync lag means "I closed it, you can open it now" is not actually true for another thirty seconds — long enough, reliably, to produce the conflicted copy on the busiest day of the month.

The deeper trouble is that the drive's version history looks like a safety net and is not one. It versions the file on the sync service's own rhythm, not at meaningful modelling moments; restoring "yesterday 16:42" resurrects whatever half-state the client last pushed, and the copy-file conflicts live outside the history entirely. Teams discover all this forensically, an hour into reconstructing whose work survived. If the models are on a sync drive today, the practical advice is short: keep it for backup, take editing coordination out of its hands immediately — an explicit convention at minimum — and treat the first -Copy(1) file as the stage-three bell ringing early.

Stage three: a team

At four or five architects, the conventions stop holding. The symptom is the filename: model_v3_FINAL_jan.archimate, model_v3_FINAL_reviewed.archimate, model_current_DO_NOT_EDIT.archimate.

This is usually where Git enters, and Git does help — it gives real history and a real merge attempt. It also introduces a new failure: merge conflicts inside a large generated XML file, which nobody can resolve by reading. The honest resolution is usually "take mine", and somebody's work is lost anyway, just with more ceremony.

Stage four: a practice

The final stage is when architecture becomes something the organisation depends on rather than something the architecture team does. Delivery teams reference it. Risk cites it. Someone asks what was approved in March.

Figure 2: The four questions a file cannot answer
Figure 2: The four questions a file cannot answer

At this point the missing capabilities are not conveniences. They are the difference between architecture as a governed asset and architecture as a folder someone maintains.

The specific things you lose

CapabilityWith filesWhy it matters
Concurrency controlNone; last save winsSilent data loss, discovered later
AttributionFile-level at best"Who added this and why" is unanswerable
History of the modelCopies, if someone made themCannot reconstruct a past state reliably
Access controlFolder permissionsAll-or-nothing; no project scoping
SearchOpen each fileCross-model discovery is impractical
AuditNoneRegulated environments need this

What actually forces the change

In practice it is rarely a considered decision. Three events trigger it:

  1. Someone loses a day's work to a concurrent save, and it is visible enough to matter.
  2. An audit or regulatory question arrives that the current setup cannot answer, and the answer has a deadline.
  3. A new architect joins and cannot be given access to one project without being given access to everything.

If you recognise your organisation in stage three, the useful question is not whether to move but what you need on the other side. Concurrency control and history are usually urgent. Search and audit are usually what makes the case fundable.

Making the case in someone else's language

The capabilities table above convinces architects, who were already convinced. The people who fund the change respond to a different vocabulary, and the translation is worth doing carefully because every row translates well.

Concurrency control translates as an incident that already happened: the lost afternoon has a date, a name and a cost, and "this will recur, more often as the team grows" is a forecast nobody can argue with. Attribution and history translate as audit exposure — not hypothetically, but against whichever framework the organisation already answers to, where "we cannot show who changed the architecture or when" is a finding waiting for someone to write it down. Access control translates as the supplier, the joint venture or the regulated perimeter that currently cannot be given anything because the only grant available is everything. And search translates as the architects' own time: hours per week spent opening files to find things, multiplied by salaries the funder already knows.

Notice what is missing from that list: any claim about modelling quality or architectural excellence. Those arguments are true and they do not fund infrastructure, because the audience cannot verify them and suspects them of being enthusiasm. The fundable case is made of incidents, findings and hours — and it is strongest at a specific moment, just after trigger one or two from the previous section has fired, while the incident is still a fresh memory rather than an anecdote. Practices that wait for the perfect strategic moment usually end up making the case in the worst possible week instead: the one after the question arrived with a deadline attached.

What not to do

Two responses are common and both make things worse.

Splitting the model into many small files to reduce conflicts trades one problem for a harder one: cross-model relationships become unmanageable and nobody can see the whole picture.

Appointing a single person to make all the changes works and does not scale. It also creates exactly the bottleneck that architecture practices are usually trying to remove, and it fails completely the week that person is away.

What the move actually looks like

The word "migration" makes teams brace for a project; the reality, done in the right order, is closer to a good fortnight. The models themselves import cleanly — they are standard ArchiMate, which is the quiet dividend of having modelled in a standards-faithful tool all along. The work is in everything around them, and it sorts into a sequence.

First, inventory and triage: which files are the estate, which are copies of the estate, and which of the seventeen candidates for "latest" wins. This is archaeology, it takes longer than expected, and it is the last time anyone will ever have to do it — a sentence worth saying aloud when morale dips. Second, structure: the flat file collection becomes deliberately partitioned models, which is the moment to fix the boundaries the file era forced. Third, identity: accounts from the directory rather than invented ones, so access control starts correct instead of being retrofitted. And only fourth, the ceremony: a cutover date, the old share made read-only, and the first week run with deliberately generous support, because the first week decides what people say about the whole exercise.

What stays the same matters as much as what changes: Archi. The architects keep their tool, their notation and their diagrams — the repository replaces the files, not the modelling. This is the difference between this migration and a tool migration, and it is why the resistance, when it comes, is smaller than feared: nobody is being asked to abandon skills, only to stop emailing XML at each other. The day-two experience — open Archi, see the shared estate, edit under a lease, publish with a comment — is recognisably the old workflow with the failure modes removed, which is exactly the pitch it should have been sold as.

How big is too big?

A question that comes up constantly and has no single answer, because the limit is rarely file size. Archi handles large models better than people expect. What degrades first is human, not technical.

The practical thresholds, in the order teams meet them:

AroundWhat starts hurting
2 architectsCoordination — "is anyone in the model?"
A few hundred elementsFinding things without knowing where they are
4–5 architectsConventions stop holding without enforcement
Several thousand elementsLoad and save times become noticeable
More than one teamAccess control becomes a real requirement

Notice that the technical threshold appears fourth. The first three are organisational, which is why splitting the file — the instinctive technical fix — does not help with any of them.

The interim measures, and how long they last

Before committing to a repository, most teams try one or more of these. They are not wrong; they are time-limited.

  • A booking convention. Announce in the channel before opening. Works with three disciplined people, fails on the first holiday.
  • A single integrator. One person merges everyone's changes. Reliable and creates a bottleneck plus a bus factor of one.
  • Splitting by domain. Genuinely helps with contention and hurts anything crossing a boundary — which, in architecture, is most of the interesting content.
  • Git. The most durable of the four, and the subject of its own article.

Each buys six to eighteen months. That is worth something — it is often long enough to build the case for the real fix — provided the team is clear that it is an interim measure rather than the answer.

The questions teams ask at this point

"Can we not just be more disciplined?" — For a while, yes, and the interim measures above are exactly that discipline, formalised. The honest observation from watching many teams try is that discipline degrades with headcount and holidays, and the capabilities in the table — attribution, history, scoped access — are not discipline problems at all. No convention, however well kept, makes a file remember who changed it. The question conflates two gaps: the coordination gap, which behaviour can bridge temporarily, and the capability gap, which it cannot bridge ever.

"Is this not overkill for six architects?" — Size the answer by the readership, not the authorship. Six architects whose models are read only by each other can live on interim measures for years. Six architects whose models are cited by delivery teams, sampled by auditors and referenced in a regulator's filing are running a governed asset with hobbyist infrastructure, and the mismatch is invisible right up until the March question arrives. The table's rows are requirements someone else imposes; count the imposers, not the modellers.

"We tried a repository tool years ago and it died — why would this time differ?" — Usually because the last attempt replaced the modelling tool rather than the storage, and the architects rejected the downgrade. The shape described here keeps Archi and replaces the files, which is a different proposition entirely: nobody loses their editor, their diagrams or their muscle memory. Failed adoptions are almost always tool rejections; this is not a tool change. That distinction, made early and repeated often, is the difference between this migration and the memory everyone is bracing against.

"What happens to the old files?" — They become the archive, and they should be treated with the respect archives get: made read-only the day the repository goes live, kept intact with their timestamps, and recorded in the migration notes as the provenance of each imported model. What they must not remain is editable, because the one guarantee a migration needs is that there is exactly one place where the architecture can change — and a writable folder of familiar files is precisely where a stressed architect under deadline will revert to, six months in, with entirely good intentions and expensive consequences.

What good looks like a year later

The test of the whole exercise is not the cutover but the twelfth month, and the estates that made the move well share a recognisable profile. Nobody has emailed a model file in months, and the phrase "the latest version" has quietly left the vocabulary because there is nothing else to be. The Monday publication lands without anyone touching it, and delivery teams cite portal links in their design documents — links that will still resolve, and still say the same thing, when the design is audited. A new joiner got access to exactly their project on their first morning, through a group membership, and the supplier from last spring lost theirs automatically when the grant expired.

Most tellingly, the practice has stopped talking about tooling. The energy that went into coordinating files goes into conventions, quality and content — the work the team was hired for. The repository has become what infrastructure is supposed to be: invisible when healthy, loud when broken, and boring in between. Files were never the problem because they were files; they were the problem because they made every one of these ordinary capabilities extraordinary. A year in, the capabilities are ordinary again, which is the whole outcome worth buying.