Immutable Revisions: History You Can Rely On

Most architecture history is not history

Ask an architecture team what the model looked like six months ago and you usually get one of three answers: a folder of dated copies someone kept, a commit log that may have been rewritten, or a shrug.

None of these is history in the sense that matters — a record that can be consulted to establish what was true at a point in time, and that nobody can quietly alter afterwards.

Figure 1: A revision chain, a named baseline, and what restore does
Figure 1: A revision chain, a named baseline, and what restore does

What a revision has to carry

Five things, and each one is doing a job:

FieldWhy it is there
ContentThe model as it stood. Not a diff — a complete state.
Content hashDetects tampering and identifies identical states cheaply.
AuthorAttribution. "Who changed this" is the second question always asked.
TimestampThe first question always asked.
Parent revisionMakes the chain navigable and proves ordering.

Storing complete states rather than diffs is worth defending, because it looks wasteful. It means any revision can be reconstructed without replaying a chain, a corrupted revision does not invalidate everything after it, and the storage cost is genuinely small — architecture models are measured in megabytes, and a repository with thousands of revisions is still a rounding error against any other enterprise dataset.

Immutable means immutable

The property that makes history trustworthy is that a published revision is never modified. Not corrected, not amended, not tidied.

This has consequences people push back on. A typo published into a revision stays in that revision forever. An architect who publishes something they should not have cannot unpublish it. Both objections are real, and both have the same answer: publish a new revision. The mistake remains in the history, which is exactly what makes the history worth consulting.

A history that can be edited to remove embarrassing states is not evidence of anything. The moment it is editable, every question it answers has the caveat "as far as we know, unless someone changed it".

Restore creates, never replaces

The operation that most often breaks immutability in practice is restore. The naive implementation makes the current state equal to an old revision by rolling back — discarding everything since.

The correct implementation creates a new revision whose content is the old one's. Revision 44 contains what revision 41 contained. Revisions 42 and 43 still exist, still say what they said, and the fact that someone chose to go back is itself part of the record.

This costs nothing to implement and it is the difference between a version history and a version history you can testify about.

Retention

Immutable revisions accumulate. The question of what to keep arrives eventually, and the answer is usually "more than you think".

A model of a few megabytes, published a few times a week, produces perhaps a gigabyte a year. Against the cost of not being able to answer a regulatory question, that is nothing. If you must prune, prune the unremarkable middle — keep every revision for the last year, keep every baseline forever, and thin the rest.

What you must not do is let retention be implicit. A repository that silently drops revisions after ninety days will do so on the day before someone needs one.

What history unlocks

Beyond audit, three things become possible that architects tend not to expect:

  • Comparing two points in time. "What changed between the approved baseline and now" is a query rather than an exercise.
  • Attributing a decision. An element that nobody remembers adding has an author and a date, and usually a revision comment explaining why.
  • Recovering from a bad change confidently. Not "restore the backup and lose a week", but "publish the content of revision 41 as revision 44".

Revision comments are worth requiring

A small discipline with a large payoff: require a comment on publish.

The objection is that people will type "update" and the field will be worthless. Some will. Enough will not that the history becomes navigable, and the ones who write "added the three services agreed at the design authority on Tuesday" make the history genuinely valuable eighteen months later.

The alternative — an unannotated chain of revisions with authors and timestamps — answers when and who but never why, and why is the question people actually have.

Comparing revisions

Once history exists, the first thing people want is to compare two points in it, and a text diff of the serialised model is not it.

A useful comparison speaks in model terms: elements added, removed and renamed; relationships added and removed; diagrams whose contents changed. That is computable from two record-based revisions with a few set operations, and it is the feature that turns a revision list from an archive into a review tool.

If you build one thing on top of revisions, build this. "What changed between the approved baseline and now" is asked before every governance meeting, and answering it by hand is why those meetings have papers nobody trusts.

Content hashing, and what it is actually for

Hashing a revision's content sounds like a security measure and is mostly an operational one. It answers a question that comes up far more often than tampering: are these two revisions the same?

That question appears everywhere once you look. Did this publish change anything, or did someone open the model and save it unchanged? Is the model in the test environment the same as production? Did the migration produce identical content to the source? All of those are string comparisons if a hash exists and expensive diffs if it does not.

The detail that decides whether it works is what goes into the hash. Include the content and exclude anything that varies without meaning — timestamps, the modelling tool's version, the order of elements in a serialisation that has no defined order. A hash that changes when nothing changed is worse than no hash, because it produces confident wrong answers.

Who is allowed to write history

Immutability is a property of the data model, not a permission, and it is worth being precise about what it does and does not prevent.

It prevents a revision from being altered through the application. It does not prevent someone with database access from issuing an UPDATE, and any design that claims otherwise is overstating. What immutability plus hashing gives you is not prevention but detectability: an altered revision no longer matches its hash, and a hash chain linking each revision to its parent means altering one invalidates everything after it.

That distinction matters when the claim is made to an auditor. "Tamper-proof" invites a question about the database administrator that has no good answer. "Tamper-evident, with a hash chain and an independent integrity check" is both true and sufficient for almost every real requirement.

History that nobody can read

A repository with four thousand revisions and no way to navigate them has history in the same sense that an unindexed archive has records. The data is there and nobody will find anything in it.

What makes a long history usable is a small number of filters, and they are the same ones every time: by author, by date range, by package touched, and by whether the revision was baselined. The fourth is the one usually missing and the one people reach for most, because the question is rarely "what happened in March" and often "what were the approved states".

Revision comments matter here more than they seem to. A history where every entry says "update" is a history that has to be read by opening things. Requiring a comment is a small friction that pays for itself the first time somebody has to explain a change from eighteen months ago.

Branching, and why this design does not have it

Anyone arriving from source control asks for branches within the first week, and the answer — that there are none — sounds like a limitation rather than a decision. It is worth explaining as a decision, because the reasoning also explains why merges do not exist.

A branch is only useful if it can be merged back, and merging architecture models is the problem the whole design was built to avoid. A branch that cannot be merged is a copy, and copies of architecture models are how estates end up with four versions of the truth.

What people actually want when they ask for a branch is usually one of two things, and both have better answers. Speculative work — trying an option — is served by a separate model with a clear name and a lifespan, which is a copy but a deliberate and disposable one. Work that must not be visible yet is served by scoped permissions on a project rather than by hiding it in a branch.

Storage growth, in numbers

The objection to storing every revision in full is storage, and it is worth answering with arithmetic rather than reassurance because the arithmetic is decisive.

A substantial enterprise model — several thousand elements, a few hundred views — serialises to somewhere around five megabytes. An active model receives perhaps two hundred publishes a year. That is a gigabyte per model per year, uncompressed, and architecture serialisations compress extremely well because they are repetitive structured text.

Across twenty active models, that is a few tens of gigabytes a year before compression and a fraction of that after. This is a rounding error against any enterprise storage budget, and it buys the ability to read any historical state directly rather than reconstructing it. Diff-based storage saves an amount of money nobody will notice and costs an operation you will perform weekly.

When history becomes evidence

Most of the time revision history is an operational convenience. Occasionally it becomes the thing an external party relies on, and the requirements change sharply at that point.

Three properties that are optional for convenience and mandatory for evidence: the timestamp has to come from a trusted clock rather than the client, the author has to be an authenticated identity rather than a name typed into a field, and the integrity check has to be runnable by someone who does not trust the application.

That last one is the one to design for early. An integrity verification that can only be performed by the product that produced the data is a weaker claim than one that can be reproduced from an export with a published algorithm. Making the hash computation documented and reproducible costs nothing at design time and cannot be retrofitted convincingly.

Pruning without losing the thread

Even accepting that storage is cheap, there are reasons to remove revisions eventually: a deletion obligation, a migration, or simply a history so long that the useful entries are lost in it.

The rule that keeps a pruned history coherent is that the chain must not break. Removing revision 400 from a chain where 401 points at it leaves 401 referencing something that no longer exists, and every integrity check from that point onwards fails.

Pruning therefore has to work in whole prefixes — remove everything before a point and make that point a new root — or it has to leave tombstones that preserve the hash and the linkage while discarding the content. The second is more work and it is what preserves the ability to say that nothing between two dates was altered, which is usually the reason the history existed.

Baselined revisions should be exempt from any pruning rule, unconditionally. They are the ones something external depends on, and they are the ones a rule written in terms of age will delete first.

Explaining immutability to the people who live with it

The design's last requirement is not technical: the people using the repository have to understand what immutability means for them, because their first contact with it is usually a small shock. Someone publishes with a typo in the comment, asks how to edit it, and is told the comment now says that forever. The design is working exactly as intended, and it has just made its first enemy — unless the explanation lands well.

The explanation that lands: history is not a diary you keep, it is a witness statement others rely on. The typo stays because the ability to fix typos is the ability to fix anything, and a history that can be tidied is a history an auditor discounts entirely. The repository's promise to every future reader — including your future self, defending a decision in eighteen months — is that what it shows is what happened, unimproved. Most people accept this in one telling once it is framed as protection rather than punishment: immutability is what makes your March defensible, when someone else is asking the questions.

Two habits complete the onboarding. Teach the correction idiom — the follow-up revision whose comment says "corrects r412: owner was recorded wrongly" — so people know the system has a way to be wrong gracefully; an append-only history with a visible correction culture is more trustworthy than a clean one, not less. And teach restraint at the keyboard: since publish comments are permanent, the practice writes them as if a regulator will read them, because one day one will. Estates that internalise these two habits stop experiencing immutability as a constraint within a quarter; it becomes what it always was structurally — the reason anyone can rely on anything the repository says about the past.

History you can rely on is ultimately a gift the practice gives its future self: every defended decision, every settled dispute, every fast audit is drawn against the deposits made now, one honest revision at a time. The design's job is to make those deposits automatic and irreversible; the practice's job is to let it.

It is worth noticing, finally, how rare this property is in the tools around the practice: documents are edited in place, wikis rewrite silently, slides have no history worth the name. An architecture repository with immutable revisions is often the only system in the building whose account of the past is structurally trustworthy — which is why, once it exists, everything from audits to arguments quietly migrates toward it. Systems that cannot lie about the past end up being asked about it constantly. That is not a burden; it is the reputation working.