Publishing Consistent Architecture Documents from the Model

The document factory nobody wanted to run

This engagement took us to an engineering consultancy โ€” several hundred engineers designing infrastructure and industrial installations, with an internal IT architecture group of about eight people supporting a steady stream of delivery projects. Every project of any size owed the governance board a solution design pack: a Word document describing context, requirements, the logical design, deployment and interfaces, somewhere between thirty and eighty pages.

The architecture group modelled seriously in Sparx Enterprise Architect. Their repository held the application estate, the integration landscape and, for each project, a solution package with genuinely maintained diagrams. And yet the design packs were assembled by hand: diagrams exported as screenshots and pasted into Word, element descriptions copied out of the model, catalogue tables retyped. An architect told us in the first interview that a design pack cost him two days of writing, of which perhaps half a day was thinking and the rest was transcription.

Transcription is not just slow; it forks the truth. The moment a diagram is pasted into a document, the document starts ageing independently of the model. Reviewers commented on the Word file, architects fixed the Word file, and the repository โ€” the thing the next project would build on โ€” quietly fell behind the documents describing it. When we compared three recent packs against the repository, all three contained at least one diagram that no longer matched the model it was copied from.

The brief was to make Sparx EA generate the packs: same governance-approved structure, current model content, and the writing time spent on design rather than assembly.

What we found in the packs

Before automating a document, it pays to read a stack of them coldly. We collected a dozen recent design packs and catalogued their structure. Nominally they followed a company template; practically, each architect had evolved a dialect. Section names drifted, interface catalogues appeared as tables in some packs and prose in others, and two packs had dropped the deployment chapter entirely because the deadline arrived first. Reviewers on the governance board confirmed what the variation implied: they spent the first ten minutes of every review finding where things were, rather than reading.

The repository side was more encouraging. The solution packages were consistently structured โ€” the group already had a package convention per project โ€” and diagram hygiene was good. The gaps were in the prose fields: element notes were rich for applications the group cared about and empty for ones they considered obvious, and tagged values for things like criticality and data classification existed but were unevenly filled. This mattered, because a generated document is merciless about gaps. A hand-written pack papers over a missing description with a sentence the author improvises; a generated one prints a blank cell where the description should be, in front of the governance board.

So the engagement, honestly scoped, was two-thirds document engineering and one-third model hygiene โ€” and we said so at the start, because clients who expect a template to fix a content problem are disappointed by week three.

The shape of the work followed that scoping. We ran it over roughly three months at two to three days a week: a fortnight of structure workshops with the architects and reviewers, then template and fragment construction in short iterations โ€” each iteration generating a real pack from a real project and putting it in front of a reviewer, because template work judged against imaginary content proves nothing. A pilot project ran the full loop before anything was declared standard, and the last fortnight went on the runbook, the validation script and a handover session where the client's own template owner made a change while we watched rather than the other way around.

Agreeing one structure before touching a template

The first workshops had nothing to do with Sparx EA. We sat the architecture group and two governance board reviewers together and negotiated the structure of the pack itself: six chapters โ€” context, requirements summary, logical design, deployment, interfaces, decisions โ€” with agreed content obligations for each. The reviewers drove out sections that existed by tradition rather than use; the architects drove out obligations no model could honour. The result fitted on one page, and it is the single most load-bearing artefact of the engagement. A generator can only enforce a structure that people have actually agreed.

One decision from those sessions shaped everything downstream: which parts of the pack are model content and which are authored prose. Element catalogues, diagrams, interface tables, decision logs โ€” model content, generated. The executive summary and the design rationale narrative โ€” authored, by a human, every time. We resisted the temptation to generate imitation prose from notes fields; a design pack whose "rationale" is stitched-together element documentation reads exactly like what it is. The authored sections live in document artifact elements beneath the solution package root โ€” one artifact per authored section, since an element carries a single linked document โ€” so even the prose is stored in the repository and captured by the baselines taken at each gate rather than on someone's desktop.

The decisions chapter earned a special mention in those workshops. Decisions had previously lived wherever the deciding meeting's notes lived, which is to say nowhere findable; the packs restated whatever the author remembered. We moved them into the model as stereotyped decision elements โ€” status, rationale, the alternatives considered, and who decided โ€” created at the moment of deciding rather than at documentation time. The generated chapter then simply prints the record. Reviewers gained the ability to ask "what changed since the last gate" and get an answer filtered on the decision dates and statuses recorded on the elements; the architects gained the stranger benefit of watching old decisions resurface intact when a design question reopened a year later, instead of being relitigated from scratch.

Figure 1: What each chapter of the design pack draws from the repository: packages, diagrams, element notes, tagged values and linked documents mapped to the six agreed sections
Figure 1: What each chapter of the design pack draws from the repository: packages, diagrams, element notes, tagged values and linked documents mapped to the six agreed sections

Mapping the model to the document

Figure 1 shows the mapping we ended up with. Each chapter of the pack declares its sources in the model. Context draws the business context diagram and the stakeholder catalogue. The requirements summary pulls requirement elements with their status and priority properties. Logical design renders the solution's application diagrams with element notes beneath each. Deployment walks the deployment packages. Interfaces is a generated table โ€” provider, consumer, protocol, data classification โ€” built from connector tagged values. Decisions prints the architecture decision elements with rationale and status.

The mapping is enforced by convention in the package structure: every project starts from a template package (created through the group's model wizard pattern) whose sub-packages correspond one-to-one with pack chapters. An architect who models in the agreed places gets a correct document without thinking about documents at all. An architect who invents their own structure discovers the generator is indifferent to improvisation โ€” which, we admit, is the point.

Cover-page metadata comes from tagged values on the solution package root: project code, version, status, author, review date. Nothing about the document is typed into the document.

The interfaces chapter deserves a note, because it forced a convention change that outlived the documents. The generated table needed protocol, direction and data classification per interface, and those facts lived โ€” when they lived anywhere โ€” in connector notes written as prose. We moved them to tagged values on the connectors, defined once with the group, and adjusted the diagramming habit so an interface is drawn once and reused rather than redrawn per diagram. That change was mildly unpopular for a fortnight and quietly transformative afterwards: the same tagged values that feed the design pack now feed an estate-wide interface catalogue the integration team had wanted for years and never had the data to build. Documents are demanding customers, and a model that can feed a document honestly is a better model everywhere else too.

Building the templates and fragments

The implementation uses Sparx EA's document generation machinery the way it works best: one master template that owns page layout, styles and chapter ordering, and a set of template fragments doing the specialised work inside chapters. The catalogue tables โ€” interfaces, requirements, applications โ€” are fragments backed by custom SQL, because the fragment queries can join through connectors and tagged values in ways the ordinary template selector cannot, and because SQL-backed fragments keep the column set stable no matter what a diagram happens to show. Where a chapter mixes diagrams with per-element notes, the fragment iterates the package's diagrams and renders each with its documentation block beneath.

A model document element at the root of each solution package binds the master template to that project's packages, so generation is one right-click for an architect โ€” or no clicks at all: a script drives the automation API's document generation interface nightly for projects in active review, writing dated output to the project's shared folder. The nightly copy means a reviewer is never more than a day behind the model, and the dated filenames quietly ended the era of SolutionDesign_final_v3_REALLY.docx.

We kept the Word styling deliberately restrained โ€” the company's fonts, colours and cover page, applied through the template's linked styles โ€” after burning more hours than we care to admit discovering which Word constructs survive EA's RTF-based generator and which do not. Multi-level numbered headings, portrait tables, image scaling: fine. Landscape sections mid-document and deeply nested tables: fragile in the combinations we tested, and we designed them out rather than fighting the generator. That constraint is real, and we name it in the limitations below.

The pilot pack, and what it taught us

The pilot was a mid-sized integration project โ€” a new laboratory data platform with a dozen interfaces โ€” chosen precisely because its solution package was in average condition rather than showroom condition. The first generated pack was sobering in the useful way: forty-one pages of perfectly formatted gaps. Eleven applications in scope, four with empty notes; an interface table with a blank protocol column three rows deep; a decisions chapter containing one decision, which everyone in the room knew was off by about six. Nothing in that list was a template defect. The template had simply printed the truth about the model, which no hand-written pack had ever done.

The pilot architect spent a day and a half filling the gaps โ€” in the model, not the document โ€” and regenerated. The second pack went to a governance reviewer cold, with no explanation that it was generated, and came back with ordinary content comments and one telling remark: it was the first pack that year in which the diagrams and the catalogue tables agreed with each other. That sentence did more for internal adoption than our closing presentation did; the other architects asked for the template rather than being handed it.

The pilot also set a rule we now consider standard: the generator is adopted one project at a time, alongside its quality gate, and never by decree across in-flight projects. A project mid-review has a document dialect already negotiated with its reviewers; forcing a regeneration into that conversation creates resistance that outlives the engagement. New projects started on the template; running ones finished as they were; within two quarters the question had answered itself.

The quality gate in front of the generator

A generated document is only as good as the model behind it, so we put a gate in front of the generator. Before any pack is produced for review, a validation script โ€” a model search plus a short automation script โ€” walks the solution package and reports: elements in scope with empty notes, interfaces missing protocol or classification tagged values, diagrams not linked into any chapter package, requirement elements without status. The script writes a one-page checklist, and the group's rule is simple: the pack that goes to governance is generated only when the checklist is clean or every remaining gap has a named owner and a date.

The quality gate changed behaviour more than the generator did. Empty notes fields used to be invisible; now they surface as named lines on a checklist two days before a review. Architects fill them in because the alternative is watching the gap print itself in front of the board.

This is the model hygiene third of the engagement paying for itself. Within a few months, the notes coverage in active solution packages went from patchy to near-complete โ€” not because anyone was lectured about documentation, but because the model became the visible source of a visible document. We have written elsewhere about model validation in Sparx EA; this engagement is the pattern applied with a deadline attached.

Mechanically, the gate is unglamorous: a model search per rule, collected by a script that runs from a menu entry any architect can reach, plus a scheduled weekly run across all active solution packages whose results go to the group lead. We wrote the rules to name their consumer โ€” "interface rows missing protocol (interfaces chapter, table column 3)" โ€” so that every finding explains which part of which document it would have disfigured. Rules without a consumer did not go in. The temptation to grow a quality gate into a general model-beauty contest is real, and each additional rule spends the credibility of the ones that matter; we held the line at fourteen checks, and the client has added only two since, both tied to their new data-protection chapter.

The review loop, and where comments go

Figure 2: The generation cycle: the quality gate checks the solution package, the generator assembles the pack from templates and fragments, reviewers mark up the Word output, and corrections flow back into the model or the linked documents before regeneration
Figure 2: The generation cycle: the quality gate checks the solution package, the generator assembles the pack, reviewers mark up the Word output, and corrections flow back into the model before regeneration

Figure 2 shows the loop as it now runs. Reviewers review Word documents. We accepted that as a fact of organisational life rather than a battle to win: the governance board marks up the generated pack with tracked changes and comments, the way they always have. What changed is what happens next. The architect triages every comment into one of two destinations: content comments are fixed in the model โ€” a note corrected, a missing interface added, a decision's rationale expanded โ€” and prose comments are fixed in the document artifacts. Then the pack is regenerated. Reviewer markup lives only in the review copy โ€” nobody corrects content by editing the generated file, because any edit there is erased by the next generation run.

The regenerated pack goes back with a short changes note, and because regeneration costs minutes rather than days, review cycles tightened from weeks to days. One reviewer's habit made us smile: he began checking the generation date on the cover against the model's last-modified date, having realised the two together gave him a fair sense of whether he was reading something current. Trust in documents, it turns out, is rebuilt by making their provenance checkable.

The dated generation outputs earned an unplanned second job as evidence. When a dispute arose months later about what a design had said at approval time, the answer was the pack the board had actually reviewed โ€” kept from the gate and referenced in the approval minutes, retrieved from the project folder in minutes โ€” no reconstruction, no argument about which "final" file was final. The group now keeps every pack that went to a governance gate, which costs nothing and has already paid for itself twice. Because the packs are kept alongside the minutes that accepted them, what was approved is a matter of record rather than recollection.

What changed for the client

The arithmetic first. A design pack that cost roughly two days of architect time now costs roughly half a day, nearly all of it in the authored sections โ€” the thinking, which was always the valuable part. Across the group's project throughput that is several weeks of architect capacity returned per year, and the group used it to clear a backlog of estate documentation that had waited years for a free week.

The governance board received identical structure in every pack for the first time. Reviews start at the content because nobody hunts for the interfaces chapter any more. And the repository is no longer the thing that lags the documents; it is the thing the documents are made of, which changed its status inside the group more than any mandate could have. When leadership later asked for a quarterly estate summary, the group answered with another template over the same repository rather than another writing exercise โ€” the capability compounds, which is the quiet win. A year on, the same repository content feeds three document families โ€” project design packs, the estate summary, and a lightweight interface catalogue for the integration team โ€” from three templates over one maintained model, which is the arithmetic that makes model maintenance easy to defend in budget discussions. When a department asks what it gets for the architects' modelling time, the honest answer used to be an appeal to good practice; now it is a list of documents the department already reads. Our overview of EA's document generation covers the machinery in more depth.

Where the approach has edges

Three limitations, stated plainly. First, the generator's Word fidelity has a ceiling: EA's RTF-based engine handles disciplined layouts well, but a communications-department document full of landscape spreads and intricate tables is not what it produces. We designed the pack to live inside what the generator does reliably, and that was the right trade โ€” but it was a design constraint, not a free choice.

Second, the review loop depends on human triage. Word comments do not flow back into the model by themselves, and an architect who skips the triage discipline reintroduces the forked truth we removed. The runbook makes triage an explicit step with a checkbox in the review workflow; discipline held for the eighteen months we can see, but the mechanism is procedural, not technical.

Third, generation quality is bounded by model quality, permanently. The quality gate catches gaps; it cannot supply insight. A solution modelled shallowly produces a shallow pack with perfect formatting โ€” faster than before. We told the client plainly that the generator amplifies whatever the repository contains, and the investment in notes and tagged values is not optional overhead but the actual content of the documents their governance runs on.

Where they are now

The template set is maintained by the client's own team โ€” we handed over the master template, nine fragments, the validation script and a runbook, and walked through a template change end to end so the first modification would not happen under deadline pressure. Since handover they have added a data-protection chapter (a new fragment over existing tagged values) without our involvement, which is what a successful handover looks like. The nightly generation still runs; the checklist still gates the board pack; and the two-day document is not coming back.

If your architects are spending their days transcribing a repository into Word โ€” the model is probably closer to generating those documents than you think, and the gaps that stop it are usually countable on one checklist. We can help you find out through our Sparx EA consulting practice, or via the contact page.

This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.