A pilot stuck between enthusiasm and refusal
The client ran customer support operations for several consumer brands — a few hundred agents answering email and chat across three sites. Management wanted to pilot an assistant that drafts replies to incoming customer emails for an agent to review, edit and send. The vendor was chosen, the business case was written, and the pilot had then spent almost half a year going nowhere, because the risk and data protection function refused to approve something nobody could precisely describe.
Their refusal was, in our view, entirely reasonable. When the risk team asked which data the drafting system could read, they received three different answers from three different people. When they asked whether a draft could reach a customer without human review, the answer was "no, obviously", supported by nothing in writing. The pilot team experienced this as obstruction; the risk team experienced the pilot as a black box being pushed at them with a deadline attached. Both were describing the same gap: there was no architecture anyone could point at.
We were brought in to close that gap — not to build the pilot and not to assess the vendor, but to produce a model of the pilot's data sources, components, human review steps and system interfaces that both sides would accept as the description of what was actually being proposed. The client already used Sparx EA in its IT function, which made the repository the natural home for the work.
What we found when we looked
The pilot's documentation consisted of the vendor's product sheet, a statement of work, and photographs of two whiteboards. The internal integration developer held most of the real knowledge in his head and was refreshingly candid about what was undecided. Piecing the picture together took about two weeks of short interviews with the pilot team, the vendor's solution engineer, the CRM administrator and the knowledge base owner.
The facts that emerged were mostly reassuring and occasionally not. The drafting service was hosted by the vendor; the client's data reached it at request time rather than being bulk-copied, which was better than the risk team had feared. But the set of data sent with each request had grown quietly during development. What began as "the customer's email and relevant knowledge base articles" had accumulated the full conversation history, CRM case fields, and — the finding that mattered most — free-text agent notes, a field where agents recorded anything and everything, including health details and payment disputes in the customers' own words. Nobody had decided to send sensitive data to the vendor. It had simply come along with a convenient API object.
This is the recurring lesson of integration archaeology: systems exchange what is easy to exchange, not what was agreed, and without a model the drift is invisible until someone goes looking.
Scoping the work to one honest use case
The temptation in an engagement like this is to model the whole support operation. We deliberately did not. The model covered one use case — email reply drafting — end to end, with everything else appearing only where the pilot touched it. A narrow model that is actually accurate beats a broad one that is approximately right, especially when its purpose is to be examined by sceptical readers looking for the gap between description and reality.
We agreed the modelling vocabulary with both audiences before drawing anything. ArchiMate application components for the systems involved, data objects for what they exchanged and stored, business processes and roles for the human steps, and a small set of tagged values that carried the compliance-relevant facts: whether a data object contained personal data, its source system, its retention rule, and whether it crossed the boundary to the vendor. The risk team helped define those tags, which mattered — the model was built to answer their questions, so their questions shaped its metadata. Designing the vocabulary around the audience is a part of architecture work that repays the half day it costs.
How the workshops ran
The modelling was built in eight working sessions over about five weeks, and the composition of the room was the most carefully designed part. Every session had the integration developer, the support operations lead and one of us; the vendor's solution engineer joined the sessions about the components and flows, under the existing confidentiality agreement; and the data protection officer joined the sessions about data sources and the review workflow. Having the DPO in the room while the model was drawn — rather than reviewing a finished artefact — shortened the whole engagement by an amount we would struggle to overstate. Questions that would have been a two-week correspondence cycle were answered on the diagram, in the room, while everyone watched.
We modelled live in Sparx EA on the shared screen, which we do on almost every engagement of this kind. The alternative — whiteboard now, transcribe later — introduces a quiet editorial step in which whoever transcribes resolves every ambiguity alone. Live modelling forces the ambiguity to be resolved in the room, where the people who know the answer are sitting. Between sessions we tidied layouts, completed tagged values and published a read-only HTML export, so the risk team could watch the description forming week by week instead of receiving it as a fait accompli. By the fifth session they had stopped being reviewers of the model and become contributors to it, which changed the tone of the eventual approval discussion entirely.
One rule kept the sessions honest: anything the room could not confirm was modelled with an explicit open-question marker rather than a plausible guess. The marker list ended every session on the screen, each item with a name and a date against it. Eleven questions went onto that list across the engagement; the last one — about how long the vendor retained request logs — took three weeks and a contract addendum to close, and would never have surfaced from a diagram drawn to look finished.
Modelling the data sources first
We modelled the data estate before the components, because in this pilot the data was the risk. Four source systems fed the drafting flow: the support platform holding the incoming email and conversation history, the CRM holding customer and case records, the knowledge base holding around two thousand articles, and a product database consulted for order details. Each contribution was modelled as a data object with access relationships showing which component read it, and the tagged values recorded what the risk team needed to know about each one.
Laying it out this way turned an abstract anxiety — "customer data goes to an AI vendor" — into a finite, inspectable list. Some entries were uncontroversial: knowledge base articles contained no personal data at all. Some needed care: conversation history was personal data but was the entire point of the use case. And one entry failed inspection outright. The free-text agent notes served no purpose the pilot team could defend; they had been included because the CRM's case object included them by default. Seeing the field on a diagram, labelled as personal data crossing the vendor boundary, settled its fate in a single meeting. It was removed from the integration before the pilot processed its first live email.
The request package, field by field
Because the request package was where the risk actually lived, we took it below the usual altitude of an architecture model and itemised it: every field that left the client's estate on every drafting request, as its own attribute on the outbound data object, each tagged for source system and personal-data content, each with a recorded outcome from the minimisation review. The exercise took one long session and settled more anxiety than any other single artefact.
| Field | Source | Personal data | Outcome |
|---|---|---|---|
| Incoming email body | Support platform | Yes | Kept — the use case |
| Conversation history | Support platform | Yes | Kept, truncated to the current thread |
| CRM case fields | CRM | Partly | Reduced to case type, status and product |
| Agent free-text notes | CRM | Yes, unbounded | Removed entirely |
| Knowledge base articles | Knowledge base | No | Kept |
| Order details | Product database | Partly | Reduced to order status and dates |
The pattern in the outcomes column is the pattern of every good minimisation review: the fields that survived intact are the ones the use case is made of, and everything else was either trimmed to the specific attributes the drafting actually used or shown the door. The CRM case object is the instructive entry. The integration had been passing the whole object because the API returned the whole object — some thirty fields, most of them irrelevant, several of them sensitive. Nobody had chosen that; the default had. Reducing it to three named fields cost the developer an afternoon and removed more personal data from the outbound flow than every other decision combined, with one exception.
That exception, the agent notes field, is worth dwelling on because of how it was defended before it was removed. The argument for keeping it was that the notes sometimes contained context that improved the drafts — which was true. The argument that won was made by the DPO from the diagram: the field was unbounded, unclassifiable in advance, and carried whatever a customer had ever said in their worst moment. A field whose contents cannot be characterised and whose purpose cannot be defended cannot be minimised, only excluded. The email body and conversation history were unbounded free text too, of course — the difference was necessity. The use case cannot exist without the message it is answering, so that exposure was accepted knowingly and recorded as such, where the notes field carried the same class of risk for no defensible return. Having the flow drawn, named and tagged made that argument concrete in a way no policy document had managed, and the decision took one meeting.
Modelling the AI components and the vendor boundary
The application layer of the model drew one line more carefully than any other: the boundary between what ran inside the client's estate and what ran at the vendor. Inside sat the support platform, an integration service built by the client's developer, and a retrieval component that selected relevant knowledge base articles for each request. At the vendor sat the drafting service and the language model behind it, which we modelled as an external application service — deliberately not decomposing the vendor's internals, because the client could neither see nor verify them, and a model should not imply knowledge nobody has.
Every flow relationship crossing that boundary carried a name for the information it moved: the request package of email, history, selected articles, the reduced case fields and order details going out; the draft reply and a confidence indicator coming back. The tagged values on those flows recorded transport encryption and the vendor's contractual retention commitment — the model held the claim and referenced the contract clause, so an auditor could follow the chain from diagram to obligation. What the model could not do, and did not pretend to do, was verify the vendor's behaviour; we return to that under the limitations below.
One modelling decision proved unexpectedly useful. We gave the confidence indicator its own data object rather than treating it as a technical detail, because the review workflow depended on it: low-confidence drafts were flagged for closer attention. Making it a first-class element meant the dependency was visible, and when the vendor later changed how the indicator was calculated, the client knew exactly which process step was affected — a small, concrete case for keeping model and change management connected.
The human review steps, modelled explicitly
The claim doing the most work in the pilot's business case was "a human reviews every reply before it is sent". A claim like that belongs in the architecture, not in a slide, so we modelled the review workflow as business processes with assignment relationships to the roles performing them — and modelled it as it would actually run, not as the idealised version.
The send action itself sat behind the agent's explicit approval in the support platform — no draft could reach a customer without an agent action on it. What the platform could not show was whether that action was considered or reflexive, which is the gap the rest of the workflow design addresses. The workflow had four paths. An agent could send a draft as-is, edit it and send, reject it and write from scratch, or escalate — the escalation path existing for drafts that were fluent, plausible and wrong, which the agents had already learned to watch for during the vendor's demonstration. Alongside the per-email flow sat a weekly sampling process: a quality lead reviewed a random selection of sent replies against the drafts they started from, looking for the failure mode nobody catches in the moment, which is the agent who stops reading before approving.
Modelling the workflow surfaced a gap that a slide would have hidden. The pilot design measured nothing about review behaviour: approval rates, edit distance, time spent per draft. The sampling process had no feed. The workshop that found the gap also closed it — the integration service would log draft-versus-sent differences and time-to-send per draft, and the tagged values on the sampling process recorded where those logs lived and how long they were kept. The risk team's later approval cited that specific mechanism, which we take as evidence that modelling the sociotechnical part of the system, not just the plumbing, is what made the approval possible.
Traceability from risk register to running system
The risk team maintained a register of eleven concerns — data minimisation, retention, the possibility of fabricated answers, agent deskilling, and so on. We imported each as a requirement element in Sparx EA and linked it with realisation relationships to the elements that addressed it: the concern about sensitive data flowing outward traced to the reduced request package and the removed notes field; the fabrication concern traced to the review workflow and the escalation path; the retention concern traced to the boundary flows and their contractual tags.
The Relationship Matrix then gave us the view the whole engagement had been aiming at: concerns down one axis, architectural responses across the other, with every cell either populated or conspicuously empty. Two concerns had no populated cells at first pass, which was the honest result — one concern about long-term vendor lock-in had no architectural answer and was accepted explicitly by the steering group as a commercial risk, and one about model behaviour drift became a new operational commitment rather than a modelling artefact. From the populated matrix we generated the approval pack with Sparx EA's document templates: one section per concern, the responding elements, the diagrams they appear in, and the owning role. The risk team received about thirty generated pages instead of a slide deck, and approved the pilot three weeks later.
The matrix outlived the approval, which was the design intention rather than a happy accident. The eleven concerns stayed in the model as living elements, and when the second use case came along, eight of them applied unchanged — the traceability work for that use case consisted largely of drawing new realisation links to existing concerns, plus two new concerns specific to chat. The review effort fell accordingly, and the risk team could see precisely which parts of the new proposal were genuinely new. A concern register that lives in the same repository as the architecture compounds in value with each use case; one that lives in a document has to be reassembled by hand for every new proposal.
What changed, and what got removed
The pilot went live roughly two months after the modelling work concluded, on one brand's email queue, with the reduced data package and the instrumented review workflow. The approval was conditional in exactly the way good approvals are: the conditions referenced elements and processes in the model, so everyone could see what had been promised and where it lived.
The most consequential outcome was subtractive. A field containing customers' unstructured personal detail stopped flowing to an external processor before the pilot started, which is the kind of change that generates no headline and prevents one. The most durable outcome was procedural: when the pilot team proposed adding chat transcripts — a new contribution from the support platform — some months later, the change arrived as a model update — new data object, new flows, tags filled in — and went through review in under a week, because the structure for describing it already existed. The second use case was designed in the model before a line of integration code was written, which is the working pattern the first use case never had.
The pilot was not approved because the architecture was elegant. It was approved because every claim in the business case — what data goes out, what comes back, who reviews it, how anyone would know if review stopped working — pointed at something inspectable in the model.
Where the model stops
A model like this describes design-time intent, and we were explicit with the client about the three places that stops being enough. The model asserts that the vendor deletes request data within the contractual window; the model records that obligation, but only the contract, its audit rights and the vendor's demonstrated compliance can make it credible — an obligation on paper is not yet a behaviour. It asserts that every reply passes human review; the platform's approval step enforces the action, and only the logging the workshops added can show how carefully it is exercised. And it captures the system as approved, not as it evolves — the vendor ships changes on its own schedule, and no repository notices a model behind an API getting quietly better or worse. Each assertion in the model that depended on runtime behaviour was therefore paired with an operational commitment, and the quality lead's sampling review is the mechanism that keeps the paper and the practice connected.
We regard naming those limits as part of the deliverable. An architecture model that claims to be the whole story is exactly the kind of artefact a competent risk function learns to distrust, and rightly so. What the model can honestly claim is narrower and more useful: at the moment of approval, everyone argued about the same picture, and the picture was accurate. In our experience that is rarer than it should be, and it is the precondition for every argument that follows being productive.
If your organisation is trying to get a pilot like this through review — or is the review function asking for a description that does not yet exist — this is work we do regularly. Our Sparx EA consulting and free Sparx EA assessment are good starting points, and you can reach us through our contact page.
This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.