A file is one value
From the outside, an architecture model in a file is a single opaque value. You can copy it, version it, lock it and back it up โ all operations on the whole thing. What you cannot do is ask it a question without parsing all of it first.
That single property is the origin of almost every limitation teams hit. No partial locking, because there are no parts. No search, because there is nothing to index without a full parse. No scoped permissions, because permissions apply to a file.
What the relational core looks like
Storing the model as records means the obvious tables, and the obvious tables turn out to be enough:
Repository
โโโ Project
โโโ Model
โโโ Folder (organisation)
โโโ Element (id, type, name, documentation)
โโโ Relationship (id, type, source, target)
โโโ Diagram (id, name, viewpoint)
โโโ DiagramObject (diagram, element, x, y, w, h, z)
โโโ DiagramConnection (diagram, relationship, source, target)
โโโ Property (owner, key, value)
Alongside these sit the things a repository needs that a file cannot carry: revisions, baselines, locks, users, groups, scoped permissions, and audit events.
Normalise the core, JSONB the edges
A tempting shortcut is to store the whole model as a JSON document in a single column. Modern PostgreSQL makes this comfortable, and it is the wrong choice for the core.
The reason is that every capability you wanted from a database depends on the model being queryable. Finding all relationships whose target is a deleted element is trivial against a table and awkward against a JSON blob. Enforcing referential integrity is free in one and manual in the other.
Keep the core relational โ elements, relationships, diagram objects. Use JSONB for the genuinely open-ended parts: tool-specific metadata, extension properties, anything whose shape you do not yet know. That boundary tends to hold up well.
What becomes possible
Locking a model rather than a file
With records, a lock is a row referencing a model, held by a user, with an expiry. It can be inspected, listed, force-released by an administrator, and it expires on its own โ none of which is available when the lock is a file handle on a share.
Search without a full parse
An index over element names, documentation and properties is a query away. More importantly it can be filtered by what the requesting user is allowed to see, which is impossible when the unit of access is a file.
Meaningful diffs
Comparing two revisions of a file gives you an XML diff. Comparing two revisions of a set of records gives you "four elements added, one relationship removed, this diagram changed" โ which is what a reviewer actually wants.
The cost, honestly
This is not free. Two costs are real and worth naming.
Round-tripping fidelity. The model has to be reconstructed exactly on the way back out to the modelling tool. Anything the tool stores that your schema does not model is lost, and "lost" here means an architect's diagram layout quietly changing. This is where most implementations of this idea fail, and it is why an extension mechanism for tool-specific metadata is not optional.
Operational surface. You now run a database. It needs backups, monitoring, upgrades and someone who owns it. For an organisation that already runs PostgreSQL this is marginal. For one that does not, it is a genuine new commitment.
Round-tripping is where this succeeds or fails
Everything good about record storage depends on one unglamorous property: a model that goes into the database must come back out identical.
Not "semantically equivalent". Identical, including the things your schema does not care about โ diagram layout details, the modelling tool's own extension properties, ordering that affects nothing but which the tool will notice.
Architects detect violations immediately and lose trust permanently. A diagram whose shapes have shifted by four pixels after a publish is a small bug that reads as "this system alters my work", and there is no recovering from that impression cheaply.
Test round-tripping the boring way: import a real model, export it, and compare. Do it on every schema change, with the largest and ugliest model you can find rather than a clean example.
Where the extension mechanism earns its place
No schema anticipates everything a modelling tool stores. Vendors add properties, plugins add their own, and a repository that drops what it does not recognise is a repository that quietly damages models.
A JSONB column carrying unrecognised properties, preserved verbatim and returned on export, solves this completely and costs almost nothing. It is the difference between a repository that supports the tool and one that supports the subset of the tool its authors knew about.
The queries that justify the whole design
Arguing for a record store in the abstract convinces nobody. Arguing for it with four questions that a file cannot answer convinces everybody, because every architecture team has been asked all four and has answered them by hand.
What breaks if this system goes away? A transitive closure over the dependency relationships. In a record store it is a recursive query returning a list. In a file it is opening every diagram and reading arrows.
Which applications handle personal data and are hosted outside the EU? An intersection of two properties across the estate. Trivial as a query, impossible to answer confidently from diagrams, and it is a question that arrives from legal with a deadline attached.
What changed in this domain since the last board? A set difference between two revisions, scoped to a package. The file version of this is a visual diff nobody can read.
Which elements have no owner? A null check. The reason this one matters is that it is the question that makes every other governance activity possible, and it is unanswerable in a file estate without opening everything.
Indexing an architecture graph
A relational store holding elements and relationships behaves well up to a surprising size, and then behaves badly in a predictable place: recursive traversal. The dependency question above is a self-join repeated to an unknown depth, and the naive schema makes it slow enough to notice at a few tens of thousands of relationships.
Three measures cover almost every estate:
- Index the relationship table on both source and target, separately. Traversal goes in both directions and a single composite index serves only one of them.
- Index the type column too, because almost every real traversal is filtered by relationship type โ you want dependency, not association, and without the filter the closure explodes.
- Bound the recursion depth explicitly. Architecture graphs contain cycles, and a recursive query with no depth limit against a cyclic graph is a way to find out how your database handles running out of memory.
Beyond that, resist the urge to reach for a graph database on principle. Architecture estates are small by database standards โ a large enterprise landscape is tens of thousands of elements, not millions โ and the operational cost of a second database technology is real, while the query advantage at this size is not.
Migrations, and the schema you will regret
A record store has a schema, and a schema has to change. This is the cost people underestimate when moving off files, because a file format absorbs change silently and a table does not.
The regret is almost always the same: modelling a domain-specific concept as a first-class column because it seemed fundamental at the time. Two years later it is not fundamental, three-quarters of rows have it null, and removing it requires a migration on a table nobody wants to lock.
The discipline that avoids it is to keep the relational core to the things that are true of every element in every estate โ identity, type, name, container, timestamps โ and put everything organisation-specific in the extension mechanism. A property that turns out to be universal can be promoted later; a column that turns out to be specific is stuck.
What this does not solve
A record store fixes storage, concurrency, history and query. It does not fix any of the reasons architecture repositories usually fail, and it is worth saying so plainly because the technology arrives wrapped in claims it cannot support.
It does not make the model accurate. A well-indexed record of something that stopped matching reality eighteen months ago is a faster route to a wrong answer. It does not create the conventions that make an estate navigable; a database will happily store four names for the same system. It does not produce readers โ a repository with no publication path is still a thing only architects open.
What it does is remove the excuses. Once the store can answer the questions, the remaining problems are all about content and practice, and those are the ones worth having.
What a revision is, exactly
"Immutable revisions" is easy to say and the detail decides whether the history is usable. A revision has to be a complete, addressable state of a defined scope โ not a diff, and not a state of the whole database.
Scoping it to the model is what makes it tractable. A revision of one model contains every element, relationship, view and property in that model at that moment, identified by a content hash so that two identical states produce the same identifier and an altered one cannot masquerade as unaltered. It carries the author, the timestamp, and a pointer to the revision it was based on.
Storing that as a full copy per revision sounds wasteful and is not, at architecture-estate scale. A large model serialises to a few megabytes; a thousand revisions is a few gigabytes, which is nothing. Storing diffs instead saves space and costs the ability to read any revision without replaying every predecessor, which is the operation you will perform most often.
The pointer to the parent revision is the field that makes optimistic concurrency work at all. Without it, a publish cannot assert what it was based on, and the server cannot tell a legitimate update from one that would silently overwrite someone else's work.
Deletion in a store that never forgets
Immutability and deletion are in genuine tension, and the tension is not academic โ it arrives the first time someone models something they should not have, or the first personal data question.
The resolution used by most designs is that revisions are immutable and current state is not: deleting an element removes it from the working model going forward, while every revision that contained it still contains it. That is correct for architecture and it means the store retains data that a deletion request was supposed to remove.
For an architecture estate this is usually fine, because the personal data is a handful of owner names and the lawful basis for keeping them in an audit history is defensible. It stops being fine if someone models a customer journey with real customer data in it, which happens more often than it should.
The mechanism worth building before you need it is a hard-delete path that operates across revisions, is restricted to two or three people, and writes its own audit entry recording that a deletion occurred without recording what was deleted. It should be used approximately never. A store with no such path at all is a store that will one day have to be restored from backup with a gap in it.
Getting the model back out
The objection that carries most weight against a proprietary record store is lock-in, and it is a fair objection. A file estate can be read by anything; a database estate can be read by whatever the vendor provides.
What defuses it is a demonstrable export to a standard exchange format โ for ArchiMate estates, the Open Exchange File format โ that produces a file the original modelling tool can open. Not a documented capability. A file, produced from the real estate, opened in the tool, in front of whoever is asking.
The detail that decides whether it works is diagram layout. Element and relationship export is straightforward and every product does it. Preserving view geometry โ where the boxes sit, how the connectors route, which nesting was deliberate โ is where exports lose fidelity, and a round trip that returns the content with all the layouts destroyed has not really returned anything.
Test it on the way in, not on the way out. A migration is the natural moment to confirm that the estate can leave again, and an export that is verified once at the start is worth more than a guarantee in a contract.
Where to start, concretely
For the implementer convinced by all this and staring at an empty schema file, the right first milestone is deliberately small: one model, round-tripped. Tables for element, relationship and view with their identity columns; the JSONB property bag; an importer that reads a real .archimate file into them; an exporter that writes it back out; and a comparison script that proves the round trip lost nothing. That last piece is the milestone's actual deliverable โ not the schema, which any afternoon produces, but the demonstrated property that models survive the trip, because every future feature stands on that guarantee and no future week will be as good a time to establish it as the first one.
The instructive part is what the comparison script finds on its first honest run: element order differences that matter to nothing but the diff, whitespace the XML serialiser normalised, a default attribute the exporter omitted because the importer never stored "this was explicitly set to the default". Each is a small lesson in the gap between "same model" and "same file", and deciding โ consciously, in writing โ which differences count is the real schema design work, done against evidence instead of speculation.
From there the build order follows the value: revisions second, because append-only history changes how everything else is written and retrofitting it is misery; the read queries third, because they prove the storage is earning its keep; and only then the concurrency, permissions and audit layers that the sibling articles in this series describe. Records over documents is not a big-bang conviction โ it is one model, round-tripped, verified, and then never again wondering whether the store can be trusted with the estate.