Search That Respects Authorization

Why search is the payoff

Ask an architecture team what they gained from centralising models and the answer is rarely locking or audit, though those are why the project was funded. It is that they can finally find things.

"Does anything already do address validation?" is unanswerable across forty files on a share and trivial against an indexed repository. That single capability changes how architecture gets used, because it turns the repository from an archive into a reference.

The leak nobody plans for

Search is also where scoped permissions quietly fail.

Figure 1: Filtering results is only half the job
Figure 1: Filtering results is only half the job

The obvious requirement is that a user cannot open a model they have no grant for. The subtler one is that they must not learn it exists. A search that returns "Project Falcon — Payment Reconciliation" to someone with no access to Project Falcon has disclosed the project's existence, its name, and the fact that it concerns payments.

In most organisations that is a minor embarrassment. In one where unannounced project names correspond to acquisitions, restructures or regulatory remediation, it is a genuine incident — and it happens without anyone doing anything wrong, because the search worked as built.

Filter at query time, not display time

The implementation detail that determines whether this is safe: the permission filter has to be part of the query, not applied to the results afterwards.

Post-filtering leaks through the edges. Result counts include what was removed. Pagination behaves oddly — page two has three items because five were filtered. Aggregations and facet counts reflect data the user cannot see. Each individually looks like a bug rather than a disclosure, which is why they survive review.

The test: for two users with different grants, does the same query return different counts and different facet values, not just different rows? If not, you are post-filtering.

What to index

For architecture models, the useful fields are narrower than for general text search:

FieldWeightNote
Element nameHighestMost queries are a half-remembered name
Element typeFacetCheap and heavily used for narrowing
DocumentationMediumTruncate; the first paragraph carries most signal
PropertiesSelected keysOwner and lifecycle, not every key
Model and projectFacetAnd the permission boundary
Relationship namesLow or noneLarge, rarely searched

Keeping the index current

An index that lags is worse than no index, because users trust it. If someone publishes a model and cannot find their own element five minutes later, they stop believing search results, and that belief does not come back easily.

Update on publish, synchronously if the model is small enough that it does not slow the publish noticeably. If it is not, update asynchronously but surface the lag — a "last indexed" timestamp is a cheap and honest way to keep trust.

Search is also an audit surface

Worth noting because it is rarely considered: search queries are themselves interesting. Repeated searches across projects a user has no grant for, by someone about to leave, is a pattern worth being able to reconstruct.

Logging queries has obvious privacy implications and should be a deliberate decision rather than an accident of implementation. But if you are building a repository for a regulated organisation, expect to be asked whether you can answer "what did this person look for" — and decide in advance what the answer is.

Ranking for a small, structured corpus

Search over an architecture estate is not web search, and importing web search intuitions produces poor results.

The corpus is small — tens of thousands of elements at most — highly structured, and heavily duplicated by nature: many elements legitimately share similar names across projects. Term-frequency ranking, tuned for large heterogeneous corpora, performs badly here because the signal is not in word frequency.

What works better is simple and explainable: exact name match first, then prefix, then substring, then documentation. Within each tier, prefer elements in projects the user works in. Readers can predict this behaviour, which matters more than marginal relevance gains.

Making "not found" useful

The most common search outcome in a young repository is no results, and what happens next determines whether people keep using search.

An empty result page is a dead end. A result page that says "nothing matched in the two projects you can see; three matches exist in projects you cannot access — request access" is genuinely useful, and can be shown without disclosing names.

That disclosure boundary needs deciding explicitly. In many organisations a count is fine and a name is not; in some, even the count reveals too much. It should be a configuration decision rather than an accident of implementation.

The three architectures, and what each costs

Authorization-aware search has three implementations in wide use. They differ in where the filtering happens, and that choice determines both the leak surface and the query latency.

ApproachHow it worksWhere it hurts
Post-filterSearch everything, then drop results the user cannot see.Result counts and pagination leak. Page two of a ten-result search returns three items and the arithmetic tells the user what they were not shown.
Pre-filter by permission setAttach the user's permitted scopes to the query and let the index restrict.Correct, and the permission set has to be resolved on every query. Cheap only if scopes are coarse.
Per-user indexBuild a separate index per user or per role.Fastest queries, worst freshness, and the index count grows with the org chart. Viable at a handful of roles, not at hundreds of users.

For an architecture estate the second is almost always right, because permissions are scoped to repositories and projects rather than to individual elements, so the filter is a short list of identifiers rather than a per-element check.

Aggregates leak even when documents do not

Filtering the result list correctly and leaving the aggregates unfiltered is the subtlest version of this failure, and it survives review because the visible results are right.

Facet counts are the usual culprit. A sidebar reading Domain: Payments (47) tells a user who cannot open a single one of those 47 elements that the Payments domain exists and roughly how large it is. The same applies to type breakdowns, date histograms, and the total-results figure.

Autocomplete is the other one. Suggestions built from the full corpus will happily complete the name of a project the user has no access to, which is a cleaner disclosure than any search result — it happens before they even press enter.

The fix is not clever: every aggregate has to be computed over the same filtered set as the results. It is worth stating as an explicit test, because the natural implementation computes aggregates from the index for speed and that is exactly the version that leaks.

What a static portal can and cannot do here

A generated static portal has no query-time authorization, because it has no query time — the search index is a file the browser downloads. Anyone who can fetch the portal can fetch the index, and the index contains whatever was put in it.

This is not a flaw to be worked around; it is a property to design with. It means scoping must happen at generation time, and it means one portal cannot serve two audiences with different entitlements.

  • Generate one portal per audience, each with its own index built from its own scope. Duplication of output, not of source.
  • Build the index after exclusions are applied, never before. An index built from the full model and then paired with filtered pages has published everything.
  • Check the generated index as an artefact, not as an implementation detail. Grep it for a term that should not be there before every release. It takes a second and it is the only test that would have caught the ordering mistake above.

Latency, and why it decides adoption

Search that takes two seconds is search people stop using, and they do not report it — they go back to asking an architect, which looks like a modelling problem rather than a performance one.

For a structured corpus of a few thousand elements the achievable target is well under 100ms, and the things that break it are predictable: resolving the permission set on every keystroke rather than once per session, computing facets over the whole corpus before filtering, and shipping a browser-side index large enough that the first search waits on a download.

The last one has a simple ceiling worth knowing. An index much past a couple of megabytes turns the first search into a network event, and on a corporate VPN that is not a small number. Trimming what gets indexed — names, types, owners and descriptions, not full documentation bodies — usually solves it without any cleverness.

Resolving a user's scope without a query per element

Pre-filtering only works if the filter is cheap to compute, and the naive implementation is not: for each result, look up whether this user may see it. That is a permission check per candidate, on every keystroke, and it collapses at a few thousand elements.

What makes it fast is that architecture permissions are not per-element in any sensible design. They are granted on repositories and projects, and elements inherit from their container. So a user's entire entitlement is a short list of container identifiers — usually under a hundred, often under ten — and the filter becomes a set membership test against a field already in the index.

That list should be resolved once when the session starts and cached for its duration, not recomputed per query. The cost of doing so is that a permission revoked mid-session keeps working until the session refreshes, which is a real gap and worth bounding deliberately: a cache lifetime of a few minutes makes the window small enough for most estates, and a revocation event can invalidate it explicitly where the estate is sensitive enough to warrant the plumbing.

The design to avoid is per-element access control lists. They make every one of these problems harder, they are almost never needed for architecture content, and estates that adopt them spend the following two years explaining why search is slow.

Search as the thing that reveals modelling quality

Nothing exposes a badly maintained model faster than a search box. Browsing a package tree hides inconsistency, because you see one branch at a time and each branch is internally coherent. Searching flattens the estate and puts the inconsistencies next to each other.

Four things surface within a week of turning search on, and all four are modelling findings rather than search findings:

  • The same system under three names, because three teams modelled it independently and nobody reconciled them.
  • Elements whose names are sentences, which rank badly and read worse. These are usually documentation that ended up in a name field.
  • Whole domains that return nothing, revealing scope everyone assumed was covered.
  • Duplicate relationships between the same pair of elements, which the tree view never showed because they render on top of each other.

This is uncomfortable and it is the most valuable thing search does in the first month. It is worth setting the expectation before launch, because the instinct when the list appears is to treat it as a defect in the search implementation.

Synonyms are the cheapest quality improvement available

The single highest-return piece of work in an architecture search is a synonym list, and it is usually the last thing anyone builds because it feels like a hack rather than architecture.

The mismatch it fixes is universal: the model holds the name IT uses and the reader searches for the name the business uses. Nobody searches for CRM-PROD-01. They search for the vendor name, or the internal nickname, or the process the system supports. A model that only answers to its own vocabulary is a model only its authors can search.

The list should come from evidence rather than from a workshop. Log the queries that return nothing for a month, sort by frequency, and the top twenty will be obvious. Mapping those twenty takes an afternoon and measurably changes whether people come back, which is more than most modelling work can claim.

Keep the list in the repository as element properties rather than in the search configuration. Aliases are facts about the elements, they are useful outside search, and configuration files are where knowledge goes to be forgotten.

The test suite that proves it

Authorization-aware search is a security control, and controls want tests rather than assurances. The suite that keeps this honest is small and adversarial, and it runs in CI against a fixture estate with a deliberately awkward permission matrix: a user with one project, a user with everything, a user with nothing, an expired grant, an overlapping pair of scopes.

Each test asks the question an attacker would. Search as the one-project user for an element name that exists only outside their scope: assert zero results — not an error, not a redacted stub, zero. Ask for the facet counts: assert the totals equal what their scope contains, because aggregates leak even when documents do not. Fetch the index artefact directly, bypassing the UI, with their session: assert it is their scope's index and nothing more. Then the regression that matters most: re-run the entire suite after every change to the permission model, the index builder and the search endpoint, because the leak this design fears is never introduced by the search team — it is introduced by an innocent refactor two layers away, in the week nobody was thinking about search at all.

Write the failure messages for the audit as much as for the developer. "Scoped user received 3 results from outside scope; expected 0" is simultaneously a broken build and a finding that never happened — and a folder of these tests, green, dated, is the cheapest answer the practice will ever give to the reviewer who asks how it knows the search does not leak.

Search that respects authorization is, in the end, a promise about silence: what the index does not say to the wrong reader matters as much as what it says to the right one. Design for that silence explicitly, test it adversarially, and the estate's most powerful discovery feature stops being its most dangerous disclosure channel.