Creating a Shared Language for Business Data

Where every report needed a translator

The organisation is a pensions administrator in the Benelux region, responsible for several schemes and several hundred thousand members. The administration platform at the centre of the estate is more than twenty years old, and around it sit the systems that grew up over those years: a payments engine, an actuarial modelling environment, a document output platform, a member portal, and a data warehouse programme that was supposed to bring the reporting side of the house into one place.

The warehouse programme is what brought us in, although not directly. The programme was staffed, funded and technically competent, and it was losing weeks to a problem that had nothing to do with technology: nobody could agree what the most ordinary words in the business actually meant. A "member" might include deferred members and pensioners, or only active contributors, depending on which team you asked. A "scheme" in one system was a "section" in another and a "plan" in a third. "Contribution" sometimes included transfers-in and sometimes pointedly did not.

Every mapping workshop between the warehouse team and a source-system team began with an hour of terminology archaeology, and some of them never got past it. Worse, the ambiguity had started to reach the outside world: a regulatory return had been queried because two sections of the same submission counted members differently, and the explanation took a fortnight to assemble.

The ask, when it finally reached us, was refreshingly plain: help us agree what our words mean, and put the result somewhere more durable than a spreadsheet. The organisation already ran a Sparx EA repository for its application architecture, which settled the "where" before we started.

Definitions lived in spreadsheets and in heads

The discovery phase was short, because there was not much formal material to discover. There were three competing glossaries: a mapping spreadsheet maintained by the warehouse team, a data dictionary left over from a platform migration seven years earlier, and a wiki page the operations team kept for training new administrators. They overlapped on perhaps half their terms and contradicted each other on a good share of the overlap.

The administration platform's database was documented by its column names and by two long-serving analysts who could explain, from memory, why the member table had four different date fields that all looked like a start date. The actuarial team, sensibly, had built their own extract layer with their own names for everything, which insulated them from the platform and added one more dialect to the estate.

There had also been a previous attempt at exactly this exercise. It had produced a ninety-page business glossary document, reviewed once, approved, filed, and never opened again. That document taught us more than any interview: whatever we produced could not be a document. It had to live in a tool people already had open, it had to be small enough that maintaining it was nobody's second job, and it had to be connected to the systems people argued about — a definition on its own settles nothing if you cannot see which system holds the thing defined.

Workshops with the people who use the words

We ran the definition work as a series of half-day workshops — around a dozen of them spread over four months. The room mattered more than the method. Each session had scheme administrators who use the words with members on the phone, someone from actuarial, someone from finance, a modeller from the warehouse programme, and one enterprise architect. Deliberately absent: anyone whose only contribution would be how a particular system happens to store things today. Physical schemas were evidence, not authority.

We scoped tightly. The first sessions covered contributions and the money coming in; later ones covered benefits, entitlements and the money going out. Membership itself — the hardest and most political domain — we left until the group had built the habit of disagreeing productively. That ordering was one of the better decisions of the engagement.

Figure 1: The working cycle, from definition workshops through the Sparx EA model to published glossary views
Figure 1: The working cycle, from definition workshops through the Sparx EA model to published glossary views

Two rules kept the sessions honest. First, no system screenshots in the room: the moment a screen appears, the discussion becomes about what the system does rather than what the business means. Second, every term left the room in one of exactly three states — agreed, parked with a named owner and a date, or split into two terms because it turned out to be two concepts wearing one name. "Parked indefinitely" was not a state, and the parking list was read out at the start of the next session.

Definitions before diagrams

The workshops themselves ran on examples rather than abstractions. A candidate definition went up on the screen, and the room tried to break it. Is a transfer-in a contribution? Is a pensioner who returns to work an active member, a pensioner, or both? Does an entitlement exist before it is calculated, or only after? Ten minutes of counter-examples told us more about a definition than an hour of wordsmithing, and the counter-examples themselves were worth keeping — we recorded the sharpest ones alongside each definition as part of the term's record.

The discipline we enforced on the definitions was brevity: two sentences, written in business language, naming no system. Anything longer is a description hiding an unresolved disagreement. Where two teams genuinely meant different things by one word — as with "member", which operations used for people and finance used for benefit records — we refused to crown a winner. The concept was split, each half got a precise name, and the ambiguous bare word was recorded as a synonym pointing at both, with a note explaining the trap.

Each agreed term carried the same small set of facts: the name, the two-sentence definition, a named steward, its status, the examples and counter-examples, and the synonyms it absorbs. Sparx EA has a built-in Project Glossary, and we chose not to use it for this — its classic list-based entries cannot hold tagged values, appear on diagrams or be traced to systems, and although recent versions can back the glossary with model elements, we wanted the terms to live inside the architecture itself rather than in a glossary structure beside it. So every agreed concept became a real element in the repository, which is what made everything in the rest of this case study possible.

The conceptual model in Sparx EA

The agreed terms became ArchiMate business objects in a dedicated conceptual model area of the repository, one package per domain: membership, contributions, benefits, schemes and employers. Where a term named a thing the business tracks — a member, a contribution, an entitlement — it became a business object. Where it named a classification or a state, it became a note on the object it classifies rather than an object of its own; a conceptual model with four hundred boxes is a glossary that has learned to draw, and we were determined to keep the model small enough to argue with.

Figure 2: The core of the conceptual model — business objects for the pensions domain and the logical data objects that realise them
Figure 2: The core of the conceptual model — business objects for the pensions domain and the logical data objects that realise them

Relationships between business objects were associations with verb phrases — a member accrues an entitlement, an employer remits a contribution, a scheme defines a benefit structure. The verb phrases came out of the workshops, not out of a modelling convention, and reading a diagram aloud as sentences became our standard test of whether it was right. Specialisation we used sparingly, only where the business genuinely reasoned about the subtype differently: deferred members and pensioners earned their place; a taxonomy of contribution types did not, and is recorded on the contribution object itself instead.

Every business object carries the same tagged values: steward, status (draft, agreed or retired), the date it was agreed, and the source of the definition. The two-sentence definition itself lives in the element's notes field, so it is at hand wherever the element is used — in the docked notes window, in search results, and in every generated document whose template pulls the notes field. After four months the conceptual model held around sixty business objects, with the membership domain still ahead of us. That number was a design goal, not an accident — sixty concepts is a language; six hundred is a landfill.

One meaning, two languages

Operating in Belgium added a wrinkle that deserves its own section, because it changed the shape of the model. The administrators work in Dutch and French, member communications go out in both, and several core terms carry legal weight in each language — the Dutch and French names of a benefit category are not translations of convenience but the words that appear in scheme rules and on statements. A glossary that picked one working language would have been dead on arrival in half the offices.

The model therefore treats the concept as the anchor and the names as attributes of it. Each business object carries its Dutch and French names as tagged values alongside the repository name, and the two-sentence definition is maintained in both languages in the element's notes, agreed as a pair in the workshops — which occasionally exposed real differences hiding inside supposed translations. The sharpest example: the Dutch term in daily use for one benefit type covered two legally distinct arrangements that the French term kept separate. That was not a translation issue but an undiscovered homonym, and it was split like any other, with both language communities naming the halves.

The glossary documents come out per language from the same elements — a document template per language pulls the matching name tagged value and that language's definition from the notes — so there is one model and two publications rather than two vocabularies drifting apart. For a Benelux organisation this pattern — concept once, names per language, definitions maintained as a pair — has become our default, and we would now recommend it even where the second language feels optional.

From shared language to logical models

The conceptual model earns its keep when it meets the systems, and that is what the logical layer is for. For each platform in scope — the administration platform, the warehouse, and the actuarial extract layer — we modelled the significant data structures as data objects in per-platform packages, and connected each one to the business object it realises. Attributes appear only at this level: the conceptual model says what a contribution is; the logical model says which fields the warehouse holds for one.

Those realisation links turned out to be the most valuable relationships in the repository, because they answer the two questions every mapping meeting had been stumbling over. "Where does this concept live?" became a traceability query: follow the realisations from the business object and you get every store in the estate that claims to hold it. When the answer for the member concept came back as five data objects across three systems, two of them in the warehouse itself, the warehouse team's duplicate-record suspicion acquired a map: five claims to check against the actual systems, and three confirmed as genuine duplicates within the fortnight.

The reverse question — "what business concept does this table serve?" — was just as productive. We used Sparx EA's Relationship Matrix with business objects on one axis and data objects on the other to review coverage domain by domain. Empty rows exposed concepts with no modelled store — sometimes a modelling gap, sometimes a real absence, always a question for the owning team to answer; empty columns exposed storage no concept claims, which is how a forgotten interim table from the previous migration finally got an owner and, three months later, a decommissioning date. The matrix review became a standing agenda item, ten minutes at the end of each workshop, and it kept both layers honest.

Repository mechanics: who edits, what is baselined

The unglamorous decisions about repository hygiene did as much for the model's survival as anything in the workshops, so they belong in the record. The conceptual and logical layers live in their own package branch, separate from the application architecture, with package-level security enabled: stewards propose, but only two people — the client's architect and, during the engagement, ours — apply changes. That is not distrust of the stewards; it is what keeps sixty elements from becoming three hundred, because every addition passes one editorial gate where "is this genuinely a new concept?" gets asked out loud.

Before each monthly review session, the branch is baselined. The baselines turned out to matter sooner than expected: within the first year, a dispute about a warehouse mapping turned on what the agreed definition of a contribution had said six months earlier, before a revision. Sparx EA's baseline comparison answered it in minutes: the pre-revision baseline held the old wording, and the dated modification note on the element supplied when it changed and why. For a model whose entire purpose is settling arguments about meaning, being able to settle arguments about the model's own history is not a luxury.

Diagram discipline was the last piece. Each domain has exactly one core diagram, kept to what fits on a screen legibly, plus the realisation views per platform. We declined, twice, to produce the wall-sized everything-diagram that programme sponsors tend to request, and offered the Relationship Matrix and model searches for completeness questions instead. Diagrams are for reading; the repository is for querying; confusing the two is how models become murals.

Keeping the language alive

A glossary dies the day it stops being maintained, so the maintenance had to be cheap and slightly automatic. The change process is deliberately small: anyone can propose a change to a term through its steward; proposals are reviewed in a short monthly session — usually under an hour — and applied to the model in the session itself, with a dated note on the element. No change advisory board, no forms. The heavier ceremony of the earlier glossary attempt was, we were told firmly, one of the reasons it died.

The automatic part is a validation script that runs against the repository on a schedule through the automation API. It checks that every business object has a definition in its notes, a steward, and a status; that every agreed data object realises at least one business object; that naming conventions hold; and and that anything agreed whose modification date is newer than its last dated note is flagged for a look. Failures land in a short report to the architect, not in anyone's way. It is the same approach we describe in our article on model validation in Sparx EA, scaled down to a glossary's needs.

Publication mattered just as much, because administrators and analysts were never going to open a modelling tool to look up one word. The model is exported on a schedule to a browsable HTML view on the intranet, and Sparx EA's document generation produces a per-domain glossary annex directly from the repository — definitions, stewards, diagrams and all. The ninety-page document exists again, in a sense, but nobody maintains it; it is generated, and it is always as current as the model it came from.

The test we set at the start: a new warehouse analyst should find the agreed meaning of any core term, and every system that stores it, in under a minute, without asking anyone. That test now passes — and the "without asking anyone" clause is the part that changed daily life.

What changed for the warehouse programme

The warehouse programme felt the difference first. Mapping workshops stopped opening with terminology archaeology, because the terminology was on the screen with a steward's name attached. Disagreements did not vanish — but they changed shape, from circular arguments about what a word means into either a lookup or a change request against a specific term. One of the two is over in a minute and the other lands with the person accountable, and both are progress.

The regulatory return that started everything got its answer as well. When the next query came, the reporting team traced the disputed figure to its warehouse table, the table to its data object, and the data object to the agreed definition of the member concept it realises — with dates and stewardship attached. The response went out in two days rather than a fortnight, and the underlying inconsistency between the two sections of the return had by then already been found and fixed using the same trail, walked in the other direction.

The quieter change was in the source-system teams. The two long-serving analysts whose memories had been the real data dictionary spent a career answering the same questions; their knowledge is now largely in the model, examples and warnings included, and they were the most enthusiastic contributors we had — partly, one of them said, because retirement stops being a governance risk when what you know is written down somewhere that will outlive the spreadsheet.

What we would do differently

Two things, and we say both to every client who asks for this kind of engagement now. First, we would start even smaller. Our first domain, contributions, absorbed nearly half the workshops, because the group was learning how to disagree as much as what to agree. The second and third domains went three times as fast with the habits in place. Front-loading one narrow, unglamorous domain as a training ground is not lost time; it is what makes the rest of the calendar believable.

Second, we would bring the regulatory reporting team into the room from the first session rather than the fourth. They turned out to be the biggest consumers of precise definitions, and the sharpest counter-example generators in the building — a regulator's question is a counter-example with a deadline. Their late arrival cost us rework on terms that had been agreed without their edge cases.

One limitation belongs in the record. Sparx EA held the conceptual and logical layers well, but it is not a data catalogue: it does not profile data, sample values, or observe what actually flows at runtime, and we drew the modelling line at the logical level on purpose. When the organisation later evaluated catalogue tooling for column-level lineage and quality metrics, the EA model gave the evaluation its requirements and its business vocabulary — but it could not, and should not, pretend to be that tool. Knowing where the model stops is part of keeping it trusted.

Where they are now

Eighteen months on, the model is still maintained, which is the only measure of success that matters for a glossary. The monthly review continues under the client's own architect; the validation script still runs; the membership domain — saved for last — was agreed without us in the room, which is how it should be. The warehouse programme shipped its first regulatory data mart mapped end to end against the shared language, and the sixty business objects have grown to just over seventy, each one argued into existence the same way.

The work behind this case study sits close to our Sparx EA consulting practice: repository design, modelling conventions, and the scripting that keeps a model trustworthy after the workshops end. If your organisation is losing weeks to words that mean different things in different rooms, you can reach us through our contact page.

This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.