Publishing Without Leaking

What a repository actually contains

An enterprise architecture repository is a description of how your organisation works, which systems hold which data, where the trust boundaries are, which controls exist and — by omission — which do not. Handed to an attacker, it is a reconnaissance document of unusual quality.

This is not an argument against publishing. It is an argument for deciding, deliberately and once, what leaves the repository, rather than discovering the answer after it has.

Figure 1: The scoping decision, made per content type rather than per element
Figure 1: The scoping decision, made per content type rather than per element

Four categories to think about

Work in progress

The most common accidental disclosure and the least dangerous. A half-designed target architecture published to a wide audience gets quoted as a decision. Exclude packages that are explicitly a working area, and make that exclusion structural — a package convention — rather than a judgement someone makes each time.

Security detail

Control implementation notes, network zone specifics, key management arrangements, named gaps. A modelled gap is exactly the sentence an attacker wants. Security content usually belongs in a separate publication with a narrower audience, not in the general one.

Third-party and commercial detail

Supplier names, contract references, licence positions. Often fine internally, rarely fine if the portal is reachable by suppliers, and sometimes contractually restricted in ways the architecture team is unaware of.

Personal data

Named individuals in ownership fields. "Owner: Jan Peeters" is personal data; "Owner: Payments Engineering" answers the same question and is not. This one is worth fixing in the model rather than filtering at publication, because it improves the model.

Exclude by structure, not by judgement

The mechanism matters more than the list. An exclusion that depends on someone remembering to tick a box each week will eventually not happen.

Make it a property of the model. A package naming convention, a tagged value, a perspective definition — anything the publication reads and acts on automatically. Then the question at review time is "is this element in the right package", which is a modelling question with a clear answer, rather than "should this be published", which is a judgement call under time pressure.

The publication report is the check. If it lists what was excluded and why, someone can review the exclusions periodically without reviewing every element. If it does not, nobody will ever verify that the scoping is still correct.

A static portal on an internal share inherits the share's access control. A static portal on a web server frequently inherits nothing — the URL is the access control, and URLs get forwarded.

If the portal is on a web server, it needs the same authentication as any other internal application. "It is only on the intranet" is not an access control, and an unauthenticated portal reachable from a VPN is one credential-stuffed account away from being public.

Two publications beat one compromise

The common failure is trying to make one publication safe for the widest possible audience, which strips it of enough content to be useless for the audience that actually needed it.

Generating twice is cheap. A full internal publication for people who need the detail, and a narrower one for suppliers or a wider internal audience, costs one more scheduled job and no additional modelling. It is almost always the right answer when the audiences differ in what they may see.

Before the first publication

Sit down once with whoever owns information security and walk the scope together. Half an hour, before anything is published, with a list of what will be included. This conversation is straightforward at that point and unpleasant afterwards.

Reviewing the scope periodically

Scoping decided once will be wrong within a year, because the repository grows and new packages are created by people who were not part of the original conversation.

The mechanism that keeps it honest is the publication report's exclusion list. Reviewed quarterly, it answers two questions cheaply: is everything that should be excluded still excluded, and has anything new appeared that nobody classified?

The second is the one that catches real problems. A new package created last month, populated with security control detail, included by default because the rule was an explicit exclusion list rather than an explicit inclusion list — that is the shape of most accidental disclosures.

Prefer explicit inclusion over explicit exclusion where the content is sensitive. Excluding by list means anything new is published by default; including by list means anything new is withheld by default. The second failure is recoverable, the first is not.

Suppliers and the third publication

Once you have two publications, a third for suppliers is a common request, and it needs a different kind of scoping decision.

Internal scoping is mostly about sensitivity. Supplier scoping is about contract: what have you committed to share, what would disclose another supplier's position, and what would tell a competitor how you operate. That is a conversation with procurement and legal rather than with security, and it is worth having before the request arrives rather than under time pressure when it does.

The exclusion that has to happen first

Scoping a publication is a pipeline ordering problem before it is a policy problem, and getting the order wrong produces a portal that looks correctly scoped and is not.

The sequence that works: extract, apply exclusions, then build everything downstream from the reduced set. Rendering, catalogues, matrices and the search index all consume the filtered model and never see the full one.

The sequence that leaks: extract, build the index, then exclude elements from the rendered pages. The pages are correct, and the index file — which is downloaded by every visitor — contains the names, types and often the descriptions of everything that was supposed to be removed. This is the most common real leak in generated portals and it is invisible from the user interface.

Checking the output rather than trusting the configuration

Exclusion rules are configuration, and configuration is believed rather than verified. The check that matters operates on the generated files.

It is not sophisticated: pick a handful of terms that must not appear — the name of a confidential programme, a system that was supposed to be out of scope, a supplier name — and search the entire output directory for them before deployment. Including the index, including the diagram files, including anything the pipeline emitted.

Running this automatically as the last pipeline step, with a failure blocking deployment, converts a policy into a control. It takes seconds, it has no false negatives for the terms it knows about, and it is the only test that would catch a rule that silently stopped matching after someone reorganised a package.

Diagrams leak differently

Element-level exclusion removes an element from the model. It does not remove it from a diagram that was rendered as an image, and this is the second common leak.

A view showing twelve applications, three of which are excluded, renders as an image containing twelve boxes unless the renderer knows about the exclusion. If diagrams are exported from the modelling tool rather than rendered by the pipeline, the renderer cannot know, and the exclusion applies to everything except the pictures.

Two workable positions. Render diagrams in the pipeline from the filtered model, which is the correct answer and requires the pipeline to control rendering. Or exclude whole views rather than elements — any view containing an excluded element is dropped entirely. The second is cruder and it is safe, which is the property that matters here.

Removing content is only half the problem. The other half is that a scoped portal can disclose the existence of what it excluded, without showing any of it.

Result counts do this. So does a facet listing a domain with zero results, an autocomplete built from the unfiltered corpus, or a navigation tree that shows a package containing nothing. Each tells a reader that something exists which they were not meant to know about, and together they can outline a programme fairly precisely.

The rule is that an excluded thing should be indistinguishable from a thing that never existed. That means empty containers are removed rather than shown empty, counts are computed after filtering, and suggestions come from the filtered set. None of this is difficult; all of it is skipped by default.

Who decides what is excluded

Exclusion decisions are made by whoever builds the pipeline, which is the wrong person, because they are commercial and legal judgements rather than technical ones.

A pending acquisition, an unannounced restructure, a supplier relationship under renegotiation — the architect building the portal may not know any of these are sensitive, and will publish them reasonably. The people who do know are not usually asked, because nobody thinks of a publication pipeline as something to consult about.

The fix is a named owner for the scope, sitting outside the architecture team, who reviews the exclusion list rather than the portal. Reviewing a list of rules is a ten-minute job. Reviewing a portal is not, which is why the review that is proposed is usually the one that never happens.

Reviewing the scope when the estate changes

An exclusion rule written against a package path is correct until somebody moves the package. Then it silently matches nothing, the pipeline reports no error, and the next publication includes everything it was supposed to omit.

Two defences. Make rules fail loudly when they match nothing — a rule excluding a package that no longer exists should stop the build, not be ignored, because the second behaviour is indistinguishable from success. And prefer rules written against a property rather than a path, since a confidentiality marker on the element travels with it when someone reorganises.

The property-based approach has its own weakness: it depends on the marker being applied to new elements, which is a modelling discipline rather than a technical control. In practice both are used, with path rules as the coarse boundary and properties for anything that genuinely must not escape.

The conversation to have before the first publication

The most useful hour in the whole exercise is spent before anything is published, with three or four people who are not architects, asking one question: what in here would be a problem if it were seen by everyone in the company?

The answers are consistently things the architecture team would not have flagged. Project code names that reveal intent. A system named after the vendor being replaced. A capability model that shows which functions are being consolidated. Ownership fields that name individuals in a reorganisation.

Doing this once, before the first publication, is worth more than any amount of subsequent review, because it establishes what categories exist. After that the ongoing question is much narrower: has anything new appeared in one of those categories?

Recovering from a leak

It is worth deciding in advance what happens if something is published that should not have been, because the instinct — quietly republish without it — is inadequate and slow.

A static portal is copyable. Anyone who opened it may have a local copy, and a browser cache certainly does. Republishing corrects the source and does not retrieve anything. The response therefore has two parts: remove it from the source and the deployed output including any retained archive, and treat the disclosure as having happened rather than as averted.

The second part is the one people skip. Whether it needs escalating depends on what leaked, and that judgement belongs to the same person who owns the scope. Having named them in advance means the question has an obvious recipient on the afternoon it matters, which is worth more than any amount of process written afterwards.

The red-team read

Configuration review checks that the rules are right; one exercise checks the thing that matters, which is what the published output actually discloses to a motivated reader. Once a year, or after any scope change, hand the portal to someone technical who was not involved in publishing it — a security colleague is ideal — with a hostile brief: you are reconnoitring this organisation; spend two hours; report what you learned.

The results are reliably humbling in specific ways the configuration never shows. The scope excluded the security package, but a diagram in the integration view still carries the firewall vendor's product name in an element label. The infrastructure models were withheld, and the tagged values on an application quietly include the hostname convention and the data-centre site code. Search-index behaviour reveals that excluded projects exist, through the counts, exactly as the earlier section warned. The retired-systems catalogue — published for lifecycle transparency — reads, to hostile eyes, as a list of the software least likely to be patched. None of these came through the front door of the scope configuration; all of them are one architect's helpful habit meeting one attacker's reading list.

Route the findings the same way as any other publication defect: some become scope changes, some become gate rules — "no hostnames in element names" is mechanically checkable — and some become modelling-convention updates, because the leak was in what got written, not in what got published. Then repeat next year, with a different reader, because the estate changed and so did what an attacker would want from it. Two hours of adversarial reading is the cheapest penetration test the organisation will ever commission, and it is the only one that tests the publication rather than the infrastructure around it.

Publishing widely and publishing safely are not in tension; they are the same discipline applied at two boundaries. The structural exclusions decide what leaves the repository, the red-team read verifies what actually left, and between them the practice earns the right to be generous with everything else — which is, after all, the point. The safest portal is not the smallest one; it is the one whose scope somebody deliberately chose and adversarially checked.

Scope decisions age like everything else here: the perimeter drawn for last year's estate is one acquisition, one divestment or one new regulation away from wrong. Put the scope review on the same calendar as the red-team read, and the publication stays as deliberate as the day it launched.