Threat Modelling an Architecture Repository

Name the asset properly

Threat modelling starts with what you are protecting, and architecture repositories are usually mis-classified at this first step. They contain no customer records, no payment data, no personal information beyond a few owner fields — so they get treated as low-sensitivity internal tooling.

What they actually contain is a structured, current, authoritative map of how the organisation works: every system, every integration point, every trust boundary, every security control, and — by omission — every place a control is missing.

Figure 1: The asset, and the four paths it most often leaves by
Figure 1: The asset, and the four paths it most often leaves by

For an attacker who has established a foothold, this is the single most valuable internal document. It removes the reconnaissance phase entirely.

The four realistic paths

Over-broad grants

The most likely and least dramatic. Everyone gets Architect because scoping was fiddly, and now four hundred people can read the whole estate. No attacker required — this is a leak the organisation performs on itself, slowly, through convenience.

Disclosure through search and listings

Covered elsewhere but it belongs on the threat list: search results, inventories, result counts and facet values that are not permission-filtered disclose the existence and naming of things a user cannot open.

Export

Every repository needs an export, and export is where controlled access becomes an uncontrolled file. The model leaves as a file, the file gets emailed, and nothing tracks it. This cannot be prevented while export exists, which means it should be logged prominently and scoped to fewer people than read access.

Credentials on the desktop

A modelling plugin holding long-lived credentials — worse, holding them inside model files that are themselves shared — turns every architect's laptop into a route to the whole repository.

What to build against them

ThreatControlDetectable?
Over-broad grantsScoped grants, effective-permission reviewYes — periodic access review
Search disclosureFilter in the query, not the resultsOnly by testing with two users
Uncontrolled exportSeparate permission, loud audit eventYes, if exports are logged
Desktop credentialsShort-lived tokens in memory, never in filesPartly — token lifetime limits damage
Insider copyingAudit of reads on sensitive projectsYes, retrospectively
Backup theftEncrypted at rest, restricted downloadYes — downloads are auditable

The threat that is not an attacker

Worth including because it is more likely than any of the above: an architect deletes half a model and publishes.

This is not malicious and it does the damage an attacker would have to work for. The control is not security — it is immutable revisions, so that recovery is publishing the prior content as a new revision rather than a database restore that costs everyone else their week.

A threat model that only contains adversaries will produce a system that survives attack and not Tuesday.

Keep it short and revisit it

A threat model that runs to forty pages gets written once and never read. The useful artefact is two pages: the assets, the trust boundaries, the handful of realistic paths, and what each is mitigated by.

Revisit it when the architecture changes — a new integration, a new export route, a new class of user. Each of those adds a path, and the point of having written it down is that the addition is visible.

Supply chain

An architecture repository is a small application with a large dependency tree, and in a regulated environment the tree is now your problem.

Producing a software bill of materials in a standard format — CycloneDX or SPDX — is not paperwork. It is what makes the question "are you affected by this week's vulnerability?" answerable in minutes rather than by a developer reading manifests.

Two related expectations follow: a documented process for how quickly you respond to a critical finding, and reproducible builds, so that the artefact deployed corresponds to the source that was reviewed.

The plugin's threat surface

The desktop component deserves its own section because it is the part outside your infrastructure.

It runs on machines with varying patch levels, alongside other plugins of unknown provenance, in the user's security context. Its update mechanism is a distribution channel — one that an attacker who compromised it would find extremely useful.

Practical mitigations: sign the plugin, publish checksums, distribute through a channel the organisation controls rather than a public marketplace, and keep the plugin's privileges minimal — it should be able to do exactly what the authenticated user could do through the API and nothing more.

Who actually wants this data

Threat models drift into abstraction when the adversary is left as "an attacker". Three concrete ones want an architecture repository, and they want different parts of it.

An attacker with an existing foothold

The most valuable reader. They have a machine on the network and no map. Your repository is the map: it names the systems, the integrations between them, the trust boundaries, and the controls at each boundary. Reconnaissance that would have taken weeks and generated detectable noise becomes a search query.

A competitor, through a departing employee

Less dramatic and more common. An architect leaving for a competitor exports the estate on their last week. Nothing is breached; a legitimate user did a legitimate export. This is the threat that export controls and audit trails exist for, and it is the one most repositories cannot detect after the fact.

A supplier during a tender

A vendor given read access "so they can understand the landscape" now knows your integration weaknesses and your renewal dates. The access was granted deliberately and scoped carelessly.

Modelling security controls without documenting the gaps

There is a genuine tension here that security teams raise and architects tend to wave away. A model that records where authentication, encryption and monitoring sit is, by omission, a record of where they do not. The better the security modelling, the sharper the map of unprotected paths.

The wrong response is to stop modelling controls, which trades a confidentiality risk for a much larger governance one. The workable response is to treat control coverage as its own scope:

  • Model controls as properties on the relationships they protect, not as free-standing elements, so a control gap is a missing property rather than a labelled hole.
  • Keep the coverage view — the one that renders gaps in red — out of the published portal and inside the repository, where scoped permissions apply.
  • Give the security team their own perspective with the full view, and give everyone else the landscape without the control layer.

That is not security through obscurity; the controls still exist and still work. It is not publishing your own gap analysis to four hundred people.

What to log, and what logging costs you

Every repository proposal includes "full audit logging" and few specify what that means. The events worth keeping are narrower than the phrase suggests, and the ones people forget are the reads.

EventKeep it because
Permission granted or revokedThe only record of who could have seen what, at a given date. Reconstructing this from current state is impossible.
Export or bulk readThe departing-employee path. A single export of the whole estate looks nothing like normal use and is trivial to alert on — if it is recorded at all.
Failed authorizationOne is noise. Forty in an hour from one account is someone mapping what they can reach.
Publish and baselineGovernance rather than security, but it is the same log and it answers the audit questions.

The cost is that an audit log containing who read which parts of the architecture is itself sensitive, and in some jurisdictions it is personal data about employees with its own retention limits. Decide the retention period deliberately and write it down, because the default of "forever" is a decision too, just an unexamined one.

Reviewing it on a cadence that survives

A threat model produced once during procurement and never opened again is a document, not a control. The realistic cadence is annually, plus whenever one of four things changes: the repository gains an integration, the permission model is restructured, the estate takes on a new class of content such as personal data flows, or the deployment moves — most often from on-premises to a hosted environment.

Each of those changes the paths in and out. Nothing else usually does, which is why an annual review that finds no changes should take twenty minutes and not be treated as a failure of diligence.

Classifying the repository, and living with the answer

Most information classification schemes have three or four levels and a rule for assigning them based on the data inside. Architecture repositories break that rule, because their sensitivity comes from aggregation rather than from any individual record.

No single element is confidential. The name of a payment gateway is not a secret. The complete graph of which systems reach the payment gateway, through which boundaries, protected by which controls, is a different object entirely, and it is the object you are storing.

Two consequences follow, and both are unwelcome:

  • The classification has to be set at the level of the estate, not the element. Anything else lets an aggregate leak one non-sensitive record at a time.
  • Whatever level you choose applies to backups, exports and the published portal too. Teams routinely classify the repository as restricted and then publish an unrestricted portal derived from it, which makes the classification decorative.

Where the published portal changes the model

Publishing is the point at which a controlled repository becomes an uncontrolled artefact. The threat model has to cover the output, not just the store, and the output has properties the store does not.

PropertyConsequence
It is a directory of filesAnyone who can read the folder can copy the whole thing. There is no per-element authorization at the filesystem level.
It is trivially indexableIf it lands on any server reachable by a crawler, it will be crawled. Static portals have appeared in public search results this way more than once.
It has no expiryA portal copied to a laptop in March is still readable in December, showing an architecture that has since changed.
Its search index contains everythingIncluding elements the perspective was supposed to exclude, if the index was built before the exclusion was applied.

The last one is the subtle failure. Scoping that removes an element from the rendered pages but leaves it in the browser-side search index has not scoped anything — it has published the element with a slightly worse user interface.

Practical mitigations, ranked by what they cost

A threat model that ends in a list of controls nobody implements has failed. These are ordered by effort, and the first three are worth doing before the model is even finished.

  1. Scope the default role down. Most estates grant Architect to everyone because scoping was tedious once. An afternoon spent on project-scoped grants removes the most likely disclosure path entirely.
  2. Alert on bulk export. One rule, one log stream. It will fire twice a year and one of those times will matter.
  3. Publish to a path that requires authentication. Even basic network-level restriction removes the crawler path, which is the one that produces a genuinely bad afternoon.
  4. Build the search index after scoping, not before. A pipeline ordering question, and free if you catch it at design time.
  5. Separate the control-coverage perspective. Real modelling work, but it is what lets the security team model thoroughly without publishing the gap analysis.

A tabletop hour to make it real

Threat models on paper stay paper until the team walks one scenario together, so close the exercise with a sixty-minute tabletop, run with the people who would actually be in the incident: the platform administrator, the practice lead, someone from security operations.

Pick the scenario the model rates most likely — say, a phished architect account with edit rights on two repositories — and narrate it forward in fifteen-minute acts, asking at each step what the team would actually see, decide and do. Would the anomalous access pattern surface anywhere, or is the audit trail write-only in practice? Who can revoke the session, how fast, and out of hours? When the tampering is suspected, who decides whether Monday's publication runs — and does anyone in the room know how to hold it? When the account's edits are identified, does the revision history actually let someone enumerate them in minutes, and who re-reviews the models it touched? The gaps announce themselves as silences in the room, and each silence is a finding more actionable than anything the STRIDE table produced.

Write down the three worst silences, fix them, and rerun the hour next year with a different scenario — the leaked portal copy, the malicious insider, the compromised runner. The tabletop is where the threat model stops being the security section's document and becomes the team's shared reflex, and reflexes are what incidents are actually handled with. An hour a year is the entire cost; the alternative pricing is set during a real incident, at rates nobody enjoys.

A repository's threat model ends where it began: with the recognition that the asset is the map. Guard the map with the same seriousness the estate guards what the map describes — scoped access, honest logging, rehearsed responses — and the repository can be what it should be: widely read, confidently shared, and never the chapter of the incident report where the attacker learned the layout.

Keep the whole exercise in proportion, finally, because proportion is what makes it repeatable. The repository's threat model should fit on three pages, take an afternoon to revise, and be readable by the practice lead without a security dictionary — a document sized to be maintained rather than admired. The estates that get this wrong in the cautious direction produce forty-page models that are obsolete before their first review and defended by nobody; the estates that get it right treat the threat model like any other living artefact in the repository: owned, dated, versioned and small. Security documentation follows the same law as architecture documentation — the value is in the currency, not the thickness — and a three-page model revisited every spring protects more than a treatise revisited never.