Optimistic Concurrency for Architecture Models

The gap locking leaves

A model lock answers "is anyone else editing this right now". It does not answer "is what I am about to publish based on what is currently there", and those are different questions.

Figure 1: The sequence where a lock is respected and work is lost anyway
Figure 1: The sequence where a lock is respected and work is lost anyway

The sequence is unremarkable and happens regularly in practice. An architect opens a model, takes a working copy, and goes to a workshop. Their lease expires. A colleague takes a lease, makes a change, publishes. The first architect returns and publishes their version.

No lock was violated at any point. The colleague's change is gone.

Carry the base revision

The fix is standard and comes from HTTP: send the revision your change was based on, and let the server refuse if it has moved on. An If-Match header carrying an ETag, or a revision identifier in the request body — the mechanism matters less than the discipline.

PUT /api/v1/models/{id}
If-Match: "rev-41"
X-Lock-Token: 9f2c…
Authorization: Bearer …

→ 200 OK          published, now rev-42
→ 409 Conflict    server is at rev-42; your base was rev-41
→ 423 Locked      lock is held by someone else, or expired

Three distinct outcomes, three distinct status codes. That distinction matters because the client's response differs: a 409 means fetch and reconcile, a 423 means acquire a lease first.

Rejecting is the feature

It is tempting to treat a rejected publish as a problem to be engineered away — merge automatically, or take the newer timestamp. Both are worse than refusing.

An automatic merge of two architecture models has the failure mode described elsewhere: it produces something that parses and is semantically incoherent. Taking the newer timestamp is just last-write-wins with extra steps. A rejection, by contrast, is a correct answer delivered to a human who has the context to resolve it.

The measure of this design is not how few conflicts occur. It is that no publish is ever silently lost. A team that sees three rejections a month and loses nothing is in a far better position than one that sees none and cannot account for its history.

Making rejection survivable

A rejected publish is only acceptable if the architect does not lose their work. Three things make it so:

  1. The working copy persists. The client keeps the local model after a rejection. Losing an afternoon to a conflict message would be worse than the problem being solved.
  2. The message says what happened. "Someone published revision 42 while you were editing; you are based on 41" is actionable. "Conflict" is not.
  3. There is a route forward. At minimum, see what changed between the two revisions. Reconciling by hand is acceptable; reconciling blind is not.

Why not just lock harder

You could require the lease to be held continuously from open to publish, making a stale publish impossible. In practice this fails against how architects work: models are opened, examined, left open across meetings, and edited in bursts. Leases that never expire block the estate; leases that expire quickly interrupt real work.

Optimistic concurrency lets the lease be generous — long enough to be convenient, short enough to self-heal — because correctness no longer depends on it. The lock becomes a cooperation mechanism that prevents wasted effort; the revision check is the thing that guarantees nothing is lost.

A note on granularity

Revision checking at model level is coarse: any change to the model invalidates any other in-flight change, even if they touched unrelated parts. Finer-grained checking is possible in principle and rarely worth it, for the same reason element-level locking is not worth it — the semantics are global, so "unrelated" is harder to establish than it looks. Keep models bounded and the coarse check stops being a problem.

Presenting a conflict to a human

A rejection is only acceptable if the person receiving it can act on it. That is a user-experience problem more than a protocol one.

The minimum useful message names the other party, the time, and the difference: "Jan published revision 42 at 14:30, adding two elements to the Payments package. You are working from revision 41." That is enough for the architect to decide whether their change still applies.

The version without the detail — "conflict: please refresh" — produces the behaviour you would expect: people refresh, lose their work, and learn to publish more often in smaller increments not because it is good practice but out of fear.

Reducing conflicts without weakening the check

Three things reduce the conflict rate without compromising correctness:

  • Smaller models. The most effective by a distance. If two architects routinely conflict, the model is probably doing two jobs.
  • Publish early. A working session that publishes every hour conflicts less than one that publishes at 5pm, and loses less when it does.
  • Visible activity. If the client shows that someone else has the model open, most conflicts never start.

What does not help is lengthening leases, which merely moves the conflict from publish time to lease-acquisition time and makes it someone else's problem.

What the check protects that a lock does not

A lock stops two people editing at once. It does not stop the case this check exists for, which is one person editing a model they loaded before someone else's change landed.

The sequence is ordinary. An architect opens a model in the morning. A colleague, holding the lease, publishes a change at eleven. The first architect acquires the lease at two and publishes what they have — which is the morning's state, silently discarding the eleven o'clock change. Both people held the lease legitimately. Neither did anything wrong. The change is gone and nothing recorded that it happened.

Carrying the base revision closes exactly this gap. The publish asserts what it was built on, the server compares, and a publish based on a superseded state is rejected rather than accepted. The lease and the revision check solve different problems and an estate needs both.

Where the base revision has to come from

The check is only as good as the value the client sends, and there is a tempting shortcut that destroys it: asking the server for the current revision at publish time and sending that.

Doing so makes every publish succeed, because the assertion is trivially true. It looks like the check is working and it is checking nothing. This happens more often than it should, usually when someone is fixing an unexplained rejection under time pressure.

The base revision must be the one recorded when the model was loaded or last successfully published, held by the client across the editing session, and never refreshed except by an explicit user action that reloads the model. Writing that rule down near the code is worth doing, because the shortcut is not obviously wrong to someone encountering it later.

Rejection messages that do not waste an hour

A rejection is the system working, and the experience of receiving one is what determines whether people accept the design or campaign against it.

The unhelpful version says the publish failed. The architect does not know whether their work is safe, what happened, or what to do, and the first instinct — try again — either fails identically or, worse, succeeds after a reload that discarded their changes.

The useful version names who published the intervening change and when, states plainly that nothing was lost and nothing was written, and offers the two actions that exist: save this work locally, or reload and reapply. Three sentences. The difference between the two versions is an afternoon of work and it is the difference between a design people trust and one they route around.

Reapplying work after a rejection

The honest cost of this design is that a rejected publish means redoing work, and no amount of good messaging removes that. It is worth being direct about how much work, because the fear is usually larger than the reality.

With a short lease and frequent publishing, a rejection costs whatever was done since the last successful publish — typically under an hour. With long-held models and infrequent publishing it can cost a day, and that is an argument for publishing more often rather than for weakening the check.

What helps materially is being able to export the rejected state to a file before reloading. It converts "redo the work" into "compare and reapply", which is a different order of effort, and it costs one menu item.

Concurrency in the publication pipeline

The same problem appears in a place nobody expects it: two extractions running at once, or an extraction running while someone publishes.

An extraction that reads the model over several minutes while it is being changed produces a portal that is internally inconsistent — elements from before a change and relationships from after. Nothing fails, nothing is logged, and the output is subtly wrong in a way nobody will diagnose.

Pinning the extraction to a revision rather than to "current" fixes it completely, and it is the same idea as the publish check applied in the read direction. It also makes the publication reproducible, which is what the audit trail wanted anyway.

What happens under an unreliable network

The awkward case for any optimistic scheme is a publish whose response never arrives. The server accepted it; the client does not know. Retrying either duplicates the work or fails confusingly, depending on the design.

Making the publish idempotent removes the ambiguity. The client generates an identifier for the attempt, and a repeat of the same identifier returns the original result rather than performing the work again. The client can then retry freely, which is what an unreliable network requires.

Without it, the fallback is for the client to query the current revision after a timeout and compare it to what it expected to produce. Workable, more code, and it fails in the case where two publishes happened in the interval — which is exactly when getting it wrong matters most.

Testing that the check actually holds

This is a mechanism whose failure is silent: if the check stops working, everything succeeds and nobody notices until work goes missing months later.

It therefore needs a test that asserts a rejection, run on every release. Two clients, both loading revision N, both editing, both publishing — the first succeeds, the second must be rejected. If the second succeeds, the check is gone.

Add the pathological cases while you are there: a publish carrying no base revision at all, one carrying a revision that never existed, and one carrying a revision from a different model. All three should be refused, and at least one of them is usually accepted by an implementation that has only been tested on the happy path.

Explaining the design to people who wanted branches

Architects arriving from software development find this design restrictive, and the objection deserves a real answer rather than a statement of policy.

The answer is that merging works for source code because code has properties models do not: it is line-oriented, a human can read a diff, and the compiler and test suite tell you whether the merge was correct. An architecture model has none of these. Its serialisation is not meaningfully line-oriented, its diffs are unreadable, and nothing verifies that a merged model is coherent.

So the choice is not between merging and locking. It is between rejecting a conflict and guessing at one, and a guessed merge in an architecture model produces a graph that is wrong in ways nobody will find until they depend on it. Framed that way the design is a trade-off with a defensible side, which is a better conversation than a rule.

The invariant, and the test that pins it

Everything in this design defends a single sentence, and the sentence is worth writing at the top of the server module that implements it: no write is accepted whose base revision is not the current revision. Every mechanism above — the carried base, the rejection, the reapplication workflow — is scaffolding around that one invariant, and the practical question for an implementer is how to keep it true through years of maintenance by people who never read this article.

The answer is the same as for any invariant: a test that fails loudly when it breaks, written as the concurrency race it protects against. Open two sessions against a fixture model at revision N. Publish from the first; assert success and revision N+1. Publish from the second, still carrying N; assert rejection — not an error, not a merge, a clean rejection naming both revisions. Then the variants that catch real regressions: the same race through the bulk-import path, through the administrative edit path, through the API a future integration will use — because the invariant is only as strong as its least famous entry point, and the classic failure is the new endpoint added in year two that writes without checking, introduced by someone solving an unrelated problem.

Run the suite in CI and once against every release candidate of the real server, and the one-line invariant acquires the property that matters for a control: it cannot be silently lost. Concurrency designs do not fail at design time — this article's argument is straightforward and nobody disputes it in review. They fail in maintenance, quietly, through the side door. A test with two sessions and a stopwatch is the entire difference between a concurrency model and a concurrency intention, and it costs an afternoon to write against a lifetime of afternoons it protects.

The design earns its keep on the day nobody notices it: the stale publish that was rejected instead of silently destroying an afternoon, the conflict surfaced while both authors still remembered their intent. Concurrency control succeeds precisely to the extent that it stays out of the way until the moment it must not — and the one-line invariant, pinned by its test, is how it keeps that appointment for years.

Locks make conflicts rare; the revision check makes them survivable. An estate needs both, and it needs to know which job each one is doing — that division of labour is the whole design, and it fits on an index card.