The repository nobody trusted
A construction and engineering group โ active across several countries, with an architecture community of around twenty-five people modelling in a shared Sparx Enterprise Architect repository โ asked us a question we hear more often than any other: how do we get our repository back? Seven years of growth had left them with somewhere around forty thousand elements, and the accumulated sediment of seven years of habits. A search for their ERP system returned three elements with slightly different names, each with its own set of relationships, none marked as the real one. Diagrams referenced elements whose owners had left. Half the application elements had no lifecycle status; a third had no owner at all.
The consequences were the usual ones, and they compounded. Architects stopped trusting search results, so they created new elements rather than risk reusing the wrong existing one โ which made the duplication worse. New joiners copied whatever patterns they found first, so inconsistency reproduced itself. Impact analysis, the whole point of keeping relationships in a repository, quietly stopped being offered to projects because nobody would stand behind the answers. The repository had crossed the line from asset to liability without anyone deciding anything.
What made this engagement different from a one-off cleanup โ which we have also done, and which fails predictably โ is that the client asked for the right thing: not a clean repository, but a repository that stays clean. That is an automation problem and a governance problem in roughly equal measure, and the automation is the easier half.
Why manual review had already failed
They had tried the obvious remedy first, as most organisations do. A quarterly model review board had run for over a year: two senior architects sampling a quarter's changes in an afternoon, writing findings in a spreadsheet, mailing them out. It failed for reasons worth naming, because they are structural rather than personal.
The volume was wrong by two orders of magnitude. Twenty-five modellers touch thousands of elements a quarter; two reviewers sampling for an afternoon see a fraction of one per cent of the change, so findings read as arbitrary bad luck to whoever received them. The feedback loop was too slow โ a naming mistake flagged eleven weeks after it was made has already been copied into other people's work. And the reviewers were senior enough that their findings felt like verdicts, which made the whole exercise faintly punitive and much resented. By the time we arrived, the review board had quietly stopped meeting, and its spreadsheet had joined the repository on the list of things nobody trusted.
The lesson we drew for the design that followed: quality feedback has to be fast, complete rather than sampled, boringly consistent, and delivered by a machine โ because a machine applying an agreed rule is impersonal in a way a senior colleague never can be.
Agreeing what good means
Automated checks are only as legitimate as the rules behind them, so the engagement started with people, not scripts. Two workshops with the architecture community produced a rule catalogue โ about thirty rules in five families โ and, just as importantly, produced them in the community's own words, with the arguments had in the open.
The five families: naming rules, one pattern per element type and stereotype; completeness rules, a short list of mandatory tagged values โ owner, lifecycle status, description โ on the element types that carry portfolio weight; structural rules, covering orphaned elements, relationship types that violate the agreed metamodel, and elements that appear on no diagram and participate in no relationship; diagram hygiene, including a ceiling on elements per diagram and a ban on free-floating elements with no relationships; and duplicate candidates, detected by normalised-name comparison across packages.
| Rule family | Example rule | Disposition |
|---|---|---|
| Naming | Application elements follow the agreed pattern; no environment suffixes in names | Mostly auto-fix |
| Completeness | Owner, lifecycle status and description present on portfolio-bearing types | Flag |
| Structure | No orphaned elements; relationship types conform to the agreed metamodel | Flag |
| Diagram hygiene | Element ceiling per diagram; no free-floating elements | Flag |
| Duplicates | Normalised-name candidates across packages | Flag, quarantine list |
Every rule got three attributes in the catalogue: a severity, an owner who can change the rule, and a flag saying whether a script may fix violations automatically or only report them. The last of these caused the longest argument and became the backbone of the design, so it gets its own section below. We wrote the catalogue up as a page per rule, with examples of compliant and non-compliant elements pulled from their actual repository โ recognition beats abstraction when you want a community to accept rules as theirs. The approach leans on Sparx EA's own machinery where it can: several structural rules are enforceable through EA's model validation framework, and we used it where it fit, adding scripted checks where the rules outgrew it.
How the tooling works
The mechanics are deliberately unexciting. A scheduled job runs every night at two in the morning on a small utility server, works through the rule catalogue against the full repository โ SQL for detection, a headless EA instance driven over the automation interface for the correction pass โ and publishes its findings before anyone arrives at work. On their repository โ SQL Server behind Pro Cloud Server, forty thousand elements โ the full run takes around twenty minutes.
Detection and correction take different paths, for performance reasons that anyone who has scripted against Sparx EA will recognise. Detection runs as SQL against the repository database โ set-based queries over the element, connector and tagged-value tables find every violation of a naming pattern or every missing owner in seconds, where walking the object model over the automation API element by element took hours when we timed it at this scale. Corrections, the small subset of changes scripts are allowed to make, go through the automation API instead, inside a nightly maintenance window โ writing through the API keeps EA's own bookkeeping intact, which direct SQL writes do not, and we treat that as non-negotiable.
The output is two artefacts from the same run. An HTML quality report, published on the intranet, shows the trend per rule family and per package area โ the practice-level view. And an Excel extract per package owner lists that owner's open findings, each with the element GUID, the rule violated, and a one-line remedy. Everything is traceable to the run that found it, and every run is logged with its counts, which mattered later when the numbers started being quoted in management meetings and someone reasonably asked where they came from.
What a finding looks like
The unit of the whole system is the individual finding, and we spent more care on its shape than on any algorithm. A finding names the rule violated in plain language, the element by name and GUID, the package path, the date first detected, and a one-line remedy โ what compliant looks like, not just what wrong looks like. "Missing owner on 'Site Logistics Planner' (Applications/Operations); set the Owner tagged value to the accountable product manager" is a finding someone can act on between meetings. "Completeness violation, severity 2" is a finding someone learns to ignore.
The GUID matters more than it appears to. Element names change โ sometimes because a naming rule was just enforced โ and a finding keyed by name orphans itself the moment its subject is fixed. Keyed by GUID plus the rule that fired โ one element can carry several findings โ a finding tracks its element through renames and moves, which is what allows the system to tell the difference between "fixed" and "renamed to dodge the check", and to close findings automatically when the nightly run no longer reproduces them. Where the repository's WebEA publication allows it, the digest links each finding straight to the element, so the distance from reading a finding to standing in front of the offending element is one click. Friction, we have learned, is the real enemy of repository hygiene โ every extra step between the finding and the fix loses a percentage of the fixes.
Auto-fix versus flag
The workshop argument about automatic correction settled on a principle that has survived contact with two years of production use: a script may fix what is mechanical, and must only flag what is meaningful. The test is whether a fix could possibly be wrong in a way that matters. Trailing whitespace and doubled spaces in element names, stereotype capitalisation matched against the profile's published list, a tagged-value date written in the wrong format โ these have exactly one correct outcome, so the nightly job applies them silently and logs each change with its before and after values.
Everything else is a flag, however tempting the automation. Duplicate candidates are the sharpest example: two elements named almost identically in different packages are usually a duplicate, but occasionally they are a genuine distinction โ a test environment and its production twin, a planned replacement modelled alongside its predecessor. A script that merges them destroys information silently and permanently. Ours writes them to a quarantine review list instead, and a human disposes of each pair: merge, rename, or mark as intentionally distinct, the last recorded with a tagged value naming the accepted counterpart, which the detector respects for that pair on every subsequent run, so accepted distinctions do not reappear as findings forever.
Every automated change is logged with element GUID, field, old value, new value and run date โ and the log is append-only. The first time an architect asks "what touched my element?", the answer decides whether the community trusts the tooling or starts working around it. There is no second first impression.
The ratio surprised the client: of the thirty-odd rules, only six qualified for auto-fix. That is the honest shape of repository quality โ most of what is wrong with a model is a judgement call that was never made, not a keystroke that went astray, and tooling should make the calls visible rather than pretend to make them.
Keeping the scripts themselves honest
Tooling that polices quality acquires authority, and authority needs its own checks. The scripts live in version control next to the client's other utilities, and rule changes go through the same small review as any code change โ a habit that earned its keep the first time a tightened naming pattern was rolled back cleanly after it misfired on a legacy package.
More unusually, the rules have tests. We maintain a small fixture repository โ a deliberately broken model containing one example of every violation the catalogue describes, plus the near-misses that should not fire: the legitimate near-duplicate pair, the element whose odd name is grandfathered by an exception tag. Every change to a rule runs against the fixture before it runs against production, and the fixture grows a new specimen every time a real-world case fools a rule in either direction. It is a modest discipline, perhaps a day of work up front, and it converts rule tuning from superstition into engineering. The alternative โ editing a regular expression and waiting to see what Monday's digests do to the community's patience โ is how quality tooling loses its licence to operate.
The nightly run also watches itself in one simple way: each run compares its finding counts to the previous run, and a swing of more than a few hundred in either direction flags the run for a human look before the digests go out. A rule misfiring against forty thousand elements produces a very large number very quickly, and the one time it happened โ the legacy-package incident above โ the digest hold meant twenty-five architects never saw the noise, which is why they still read their digests.
Routing findings to the people who can fix them
A quality report addressed to everyone is addressed to no one, so ownership does the routing. Every top-level package carries an owner in a tagged value โ establishing this was itself a fortnight of archaeology at the start of the engagement, and four package areas genuinely had no owner to name, which was a finding in its own right and went to the practice lead to resolve.
Each Monday, every owner receives their digest: open findings in their packages, sorted by severity and age, with the week's delta on top โ new findings, fixed findings, and anything past the agreed age threshold. The digest is deliberately short and deliberately weekly; a daily nag would be muted within a month. A forty-five-minute monthly quality clinic, run by the practice lead, walks through the stubborn residue: findings nobody knows how to fix, rules producing noise, disputes about whether something is actually wrong. Rules get amended in the clinic โ several were, in the first months, which did more for the tooling's legitimacy than any accuracy could.
One thing we advised against, and still do: a leaderboard. Ranking owners by finding count converts a quality practice into a blame instrument overnight, and the response is not better models โ it is quarantine packages stuffed out of sight and elements deleted rather than fixed. The trend chart per area, shown without names attached in the clinic, applies all the social pressure that is useful and none that is destructive.
Rolling it out without a revolt
Switching thirty rules on against seven years of sediment would have produced a number so large it demoralised everyone and a flood of digests that taught people to delete them. The rollout ran in three phases over about five months, and the phasing was as much a part of the design as any script.
Phase one ran report-only for six weeks. The first full run found a little over four thousand findings, which we published with the explicit message that the number was a starting line, not a verdict โ nobody was asked to fix anything yet. The six weeks tuned the rules against reality: three rules generated most of the noise and were tightened or rewritten, and the duplicate detector's normalisation needed two rounds of adjustment before its candidate list was worth a human's time.
Phase two was the paydown. Each package area ran one or two fix sprints โ a focused week where the owner and the architects working in that area cleared their backlog, with the mechanical fixes already done for them by the newly enabled auto-fix rules. The baseline dropped from four thousand to under a thousand across three months, and the trend chart of that descent, updated nightly, turned out to be the single most motivating artefact of the engagement.
Phase three made quality structural: a gate. Baselines and the weekly HTML publication both run through the scheduled tooling rather than by hand, and the tooling refuses to produce either if any critical-severity finding is open in the packages being published โ publication being the moment model content reaches people who cannot see the caveats. Thresholds ratchet: they tighten after each clean month and are never loosened without a decision recorded in the clinic. The gate has been genuinely unpopular exactly twice, both times a day before a deadline, and both times the finding it blocked on was real.
What changed
Two years on, the steady state is a few hundred open findings โ the working residue of an active repository, not a backlog โ with criticals typically at zero and never older than a week. But the number was never really the point. Search became usable, and the duplicate rate for new elements fell to nearly nothing once architects could trust that reuse was safe, which broke the vicious circle that had created the mess. Impact analysis is offered to projects again. New joiners inherit the conventions from the rules rather than from whichever package they open first, and their first Monday digest teaches them more about the house style than the onboarding deck does.
The quality reports also acquired an audience nobody designed for: management. The per-area trend became part of the quarterly IT review, and the practice lead reports that being able to show the repository's health as a maintained curve โ rather than assert it as an opinion โ changed the budget conversation around the architecture practice noticeably. A repository with evidence of its own quality is a different negotiating position from a repository with a reputation.
The honest edges
Three limitations deserve naming, because the pattern is sometimes sold without them. First, semantic quality is out of reach: the scripts verify that an element has an owner, a lifecycle and a compliant name, but not that the owner is the right person or that the model tells the truth about production. Automated checks police form; only review polices meaning. The monthly clinic and ordinary peer review still carry that load, and always will.
Second, the duplicate detector finds candidates, not duplicates. Its precision improved with tuning, but roughly one candidate in five is a legitimate distinction, which is exactly why merging was never automated. Anyone who claims their tooling resolves duplicates without a human in the loop is describing data loss with confidence.
Third, the rules encode today's conventions, and conventions age. Two of the original naming patterns were revised within the first year as the metamodel evolved, and each revision briefly spiked the finding count against existing content until a targeted auto-fix migrated the old pattern forward. A rule catalogue without an amendment process would have calcified into exactly the kind of resented bureaucracy the review board had been; the clinic's authority to change rules is what keeps the machine legitimate.
Where they are now
The client runs the whole arrangement without us, which is the exit we design for. The scripts are theirs, documented and versioned alongside their other tooling; the rule catalogue has a standing owner; and the clinic has met monthly for two years, latterly with better attendance than the review board ever managed. Our remaining involvement is an annual afternoon reviewing the rule set against how the repository has evolved โ the sort of check-in described in our Sparx EA consulting practice, and deliberately small.
If your own repository has crossed the line from asset to liability โ if your architects have started routing around search, and impact analysis is a service your team no longer dares to offer โ you can reach us through our contact page.
This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.