A programme with three versions of the truth
The organisation was a transport operator in the middle of a multi-year modernisation programme: a new ticketing back office, new validators on vehicles, a rebuilt mobile app and a real-time passenger information service, each running as its own delivery stream with its own suppliers. Just under two thousand requirements had accumulated across the programme, and they lived in three places at once. The original procurement annexes held them in Word. Each stream kept an Excel tracker that had started as a copy of the annex and then evolved. The delivery teams worked from Jira, where requirements had been rewritten as epics and stories in language that no longer matched either of the other two sources.
The architecture team, meanwhile, kept a well-maintained Sparx EA repository describing the application landscape — current and target — that none of the requirement sources referred to. Architecture decisions were recorded in slide decks, one deck per steering meeting, filed by date. Everything existed; nothing was connected.
The moment that turned this from an irritation into a mandate came during a programme review, when an accessibility audit asked which solution components satisfied the accessibility requirements in the procurement annex. Producing an answer took the better part of three weeks, involved four people, and produced three lists that did not agree with each other. The programme director's question to us was simple: make it so that this class of question can be answered from one place, in minutes, with an answer we can defend.
What we found when we looked
Before proposing anything we spent time reading the actual artefacts, and the problems were more specific than "requirements are scattered". The Word annexes and the Excel trackers had drifted apart: requirements had been reworded, split or quietly dropped in the trackers, and because rows had been inserted and deleted, the row numbers people used as identifiers no longer meant anything stable. A requirement cited as "F-117" in a design document might be F-121 in the current tracker, because the numbers had been re-derived from row order more than once. There were no identifiers that survived editing, which meant every cross-reference in every document was built on sand.
The decision decks had a different problem. They recorded real decisions with real rationale, but they referred to systems by names that had since changed, and they never referred to requirements at all. Reconstructing why the validator platform had been chosen meant finding the right deck, then finding someone who remembered the meeting. Two of the people who remembered had already left the programme.
We also found something worth protecting. The analysts had a working routine in Jira that the delivery teams depended on, with its own review states and its own rhythm. Any approach that required analysts to abandon that routine and author requirements inside Sparx EA instead was going to fail on adoption, whatever its technical merits. That observation shaped the whole design of the solution.
Setting up the work
We agreed to prove the approach on one stream before touching the other three, and picked the validator replacement stream: around three hundred and forty requirements, active design work, and a supplier contract that made traceability contractually interesting. We asked for three things — write access to a dedicated area of the repository, two working sessions a week with the lead analyst and the stream architect, and a standing rule that any decision taken in those sessions went into a decision log in the model from day one. The last point was partly about eating our own cooking: if recording decisions in the repository was too clumsy for us, it would certainly be too clumsy for them.
Six weeks were planned for the pilot: two for the requirement backbone and the import, two for linking the existing design, and two for reviews, corrections and the first generated outputs. This kind of scoping matters in Sparx EA consulting work because traceability initiatives fail by boiling the ocean far more often than they fail on technique.
A backbone for requirements in the repository
The backbone is deliberately unexciting. Each requirement became a Requirement element in Sparx EA, in a package structure with one package per stream and sub-packages by topic — fare media, validation, device management, reporting. The requirement text went into the element notes. The element name carried a short readable label prefixed by a stable identifier, and that identifier is the important part: it lives in a tagged value we called SourceID, it is issued once, and it never changes, whatever happens to the wording, the ordering or the grouping of requirements. The EA GUID anchors the element internally; the SourceID anchors it in every conversation with people.
Requirement type — business, functional, non-functional, constraint — was applied as a stereotype, and status as a tagged value with an agreed set of values. We resisted the temptation to carry over the fifteen ad-hoc columns from the Excel trackers; most of them were duplicates of each other or notes in disguise. Four tagged values survived. Everything else stayed behind, on purpose.
The first load used EA's CSV import to get moving, but we replaced it within the week with a script over the automation API, because the load had to run unattended and match on SourceID — an identifier the trackers carry and EA's GUIDs do not — rather than be a one-off migration. The script matches on SourceID, creates what is missing, updates text and status on what exists, and writes a short log of everything it touched. The automation API is what makes this a living arrangement rather than a snapshot that starts rotting the day after the import.
The metamodel on a page
Everything the pilot used fits on one page, and keeping it that small was a fight we are glad we won. Early workshops produced proposals with a dozen requirement categories, five kinds of decision and a separate element type for assumptions. Each addition sounded reasonable in isolation; together they would have produced a repository that only its authors could use. The version that shipped has four requirement stereotypes, one decision stereotype, and three connector rules, and after six months nobody has asked for more.
| Element | Modelled as | Carries |
|---|---|---|
| Requirement | Requirement element, stereotyped business, functional, non-functional or constraint | SourceID, status, text in notes |
| Decision | Stereotyped element in the stream's decision package | Status, date, forum, rationale in notes |
| Satisfaction | Realization from component or service to requirement | Nothing — its existence is the information |
| Decomposition | Aggregation between requirements | Nothing |
The connector rules matter more than the element list. Realization always points from the solution to the requirement, never the other way, so the Traceability window reads the same way for every reviewer. Aggregation is only used between requirements, so decomposition trees stay inside the requirement world. And associations from decisions are deliberately loose — a decision may touch anything — because the value of a decision link is recall, not rigour. We wrote these rules into the repository's model guidance and, later, into a validation script that flags violations overnight, in the spirit of the checks described in our piece on model validation in Sparx EA.
Linking requirements to the design
With the backbone in place, linking is conceptually simple: a realization connector from an application component or application service to each requirement it satisfies. Decomposition between requirements uses aggregation, so a coarse procurement-level requirement can be broken into the finer statements the designers actually work against, without losing the thread back to the contract.
The discipline is in who creates the links and when. We put link creation into the designers' definition of done: a design package is not review-ready until every component in it either realises at least one requirement or carries a note saying why it exists anyway. Reviewers check the links in the Traceability window during the review itself, walking up from a component to the requirements above it, which turns an abstract obligation into a two-minute habit. The Relationship Matrix, scoped to one package of requirements against one package of components, became the standing review view: realisations appear as marks in the grid, and an empty row is visible to everyone in the room at the same moment.
One modelling choice deserves a mention because it saved arguments later. Requirements link to application services where behaviour is what matters, and to components only where the requirement genuinely constrains a specific building block. Linking everything to components feels concrete but ages badly — components get replaced, services survive replatforming, and the traceability should survive with them.
Non-functional requirements cut across everything
Functional requirements link naturally to the service or component that does the work. Non-functional requirements do not: a response-time obligation on validation touches the validator firmware, the on-vehicle gateway and the back office at once, and an availability target touches nearly everything. The lazy options are both bad — link the NFR to every component and the matrix becomes a wall of marks that nobody reads, or link it to nothing and the audit question comes back in a year.
The pattern we settled on links each non-functional requirement to the small set of application services whose behaviour it genuinely constrains, and lets the service-to-component realisations carry the implication downward. The three-hundred-millisecond validation target links to the fare validation service alone; anyone who needs the affected components walks one step down the Traceability window. For the handful of true cross-cutting constraints — data protection obligations, the accessibility set — we accepted links to a deliberately curated group of services, agreed in a workshop, recorded with a decision element explaining the scoping. It is not perfectly precise. It is defensible, readable and maintained, which for the audit conversations that prompted the whole engagement is worth more than precision.
The accessibility requirements that started everything got one more treatment: a dedicated diagram per user-facing stream showing the requirement set and the realising services on one page. When the audit returned mid-programme, that diagram and the scoped matrix export were the answer — prepared before the meeting in which the question was asked, rather than three weeks after it.
Recording decisions next to their consequences
Architecture decisions moved out of slide decks and into the model as stereotyped elements — one element per decision, carrying tagged values for status, date and the forum that took it, with the rationale in the notes. Each decision element is associated with the components it affects and the requirements that drove it. The decks did not disappear; they are still how decisions get discussed. But the deck now quotes the decision element's identifier, and the model is where the decision lives after the meeting ends.
The payoff is that a decision sits in the Traceability window two clicks from anything it touched. When the validator firmware supplier proposed a change to offline validation behaviour, the stream architect pulled up the affected component, walked up to the decision that had fixed the offline requirements, and had the original rationale — and the names of the people who agreed it — in front of the supplier the same afternoon. Under the old regime that conversation would have started with an archaeology project.
A model search keeps the decision log honest: any decision element still in proposed status after a month, and any decision with no associations at all, appears on a list that the weekly session reviews. A decision linked to nothing is either unfinished bookkeeping or a decision about nothing; both deserve attention.
Keeping authoring where the analysts work
The analysts never stopped working in Jira and Excel, and this was a deliberate design decision, not a compromise we settled for. The stream tracker is the source of record for requirement text — the Jira stories carry the SourceID and are reconciled against the tracker inside the analysts' own routine — and a nightly script reads the trackers, matches rows to elements on SourceID, and updates text and status in the repository. The flow is strictly one way: authorship stays outside, linking happens inside. To make the boundary impossible to cross by accident, the requirement packages are locked with package-level security under a group only the sync account belongs to, so requirement text cannot be edited interactively — designers draw their realization connectors from diagrams in their own design packages, and the script alone writes text and status.
It is worth being honest about the limitation underneath this arrangement. Sparx EA can author requirements — the Specification Manager gives a document-style editing surface built for exactly this — but these analysts lived in Jira, their review workflow lived there with them, and no editing surface was going to outweigh that. And with EA's auditing switched off, as it was in this repository, there is no per-requirement text history of the kind a dedicated requirements management tool keeps by default. The one-way sync accepts all of that: the tool each role already trusts remains its home, and the repository's job is to hold the connections no other tool was holding.
The rule that made the whole arrangement stable: requirement text is authored outside, links are authored inside, and no tool owns both. Most failed traceability attempts we have seen broke this rule in one direction or the other.
Coverage reviews that fit in an hour
The recurring outputs are unglamorous and heavily used. A model search lists requirements with no incoming realization — the orphans — counted after rolling decomposition up, so a parent whose children are all satisfied does not appear, and withdrawn requirements are excluded. Another lists components that realise nothing, which is where gold-plating goes to be found. A third lists stale decisions, as described above. The three lists together, plus the Relationship Matrix for whichever package is under review, fit comfortably into a one-hour session.
The first orphan run on the pilot stream listed just over eighty of the three hundred and forty requirements — a number that caused a brief silence in the room. Working through the list took two sprints. Most orphans were genuine gaps in linking rather than in design, a handful exposed design work that had quietly never happened, and about a dozen turned out to be obsolete requirements that nobody had had the authority to delete from a contractual annex. Those were marked as withdrawn in the tracker — the flow stays one-way — with a decision element in the model recording who agreed and why, which is exactly the kind of small governance the model makes cheap.
For documents, we configured a template so that each stream's solution design pack is generated from the repository with a traceability annex: every requirement in scope, its status, the services realising it, and the components behind those services, produced by walking the realisation chain one step down. Document generation from the maintained model means the annex is a by-product, not a document anyone maintains — regenerating it after a model change takes about a minute, and a regenerated annex cannot disagree with the model it came from.
What changed for the programme
The accessibility question that took three weeks now takes minutes: scope the Relationship Matrix to the accessibility requirements, export it, done. More telling is what happens with change. When the programme descoped one of the fare media options, the impact list — affected requirements, affected components, decisions that mentioned it — was generated the same afternoon the change was tabled, and the steering discussion worked from that list rather than from recollection.
Supplier conversations changed tone as well. The validator supplier receives the generated traceability annex with each design iteration and returns comments against requirement identifiers that both sides agree on. The argument about which version of which spreadsheet was current has simply ended, because nobody negotiates from a spreadsheet any more.
Perhaps the best signal: the second and third streams were brought onto the approach by the operator's own architecture team, using our runbook, with us in a review seat rather than doing the work. The fourth stream is scheduled but not started at the time of writing — we would rather say that plainly than round it up to a completed rollout.
Six months in
Some of what we watch to judge whether an arrangement like this has taken root. The nightly sync has run for about half a year with a handful of failures, each one caught by its own log and each one traceable to a renamed worksheet or a tracker saved in the wrong folder — human events, not technical ones, and each fixed the next morning. The orphan list, which began at just over eighty for the pilot stream, sits at a steady handful across three streams; it never reaches zero, and a standing agenda item keeps it from growing, which is the realistic definition of success.
Link volume tells its own story. The three active streams hold around nine hundred requirements and a few thousand realization connectors, created by a dozen designers as part of normal design work rather than by anyone whose job title contains the word traceability. That distinction is the one we care most about, because a traceability layer maintained by a dedicated person is a traceability layer that stops the day that person moves on. The weekly session that once needed us now runs without us; our involvement has settled into a monthly review of the searches and the occasional argument about whether something deserves to be a decision.
And the decision log has quietly become the most consulted part of the repository. New joiners are pointed at it in their first week. Suppliers quote decision identifiers back in their own documents. Nobody has opened the old decks in months.
What we would do differently
Three things. First, we would fight harder to start the decision log programme-wide in week one, not stream by stream. Decisions turned out to be the highest-value element type per hour invested, and every week of delay was a week of rationale evaporating into slide decks. Second, our pilot stream choice flattered the approach: validator requirements were relatively stable, so the nightly sync had an easy life. A stream with heavy requirement churn would have stress-tested the matching rules earlier and exposed edge cases — split requirements, merged requirements — that we instead met in month four. The matching logic handles them now, but we would rather have met them in the pilot.
We would also set expectations about the identifier debate. Agreeing the SourceID scheme consumed a surprising amount of week one, because identifiers are the one thing everyone in a programme has an opinion about. It is worth settling in a single session with the right people in the room, and it is genuinely important — but it is a one-hour decision, not a one-week one.
If your organisation is facing a similar situation — requirements in several tools, design in Sparx EA, and nobody able to say with confidence which parts of the solution answer which obligations — you can reach us through our contact page.
This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.