Three proposals that all said yes
An organisation in the aviation sector — running scheduled passenger operations with a few thousand crew across several bases — was replacing the system at the centre of its crew operation: rostering, day-of-operations crew tracking, and compliance with flight-time limitation rules. The incumbent system was two decades old, written around rules that had since changed twice, and held together by the two people who still understood its batch jobs.
The procurement had reached a shortlist of three vendors when we were brought in, and it had stalled in a way the readers of vendor proposals will recognise. All three responses answered essentially every requirement with yes. All three included an integration section listing "standard adapters" for the surrounding systems. All three were, on paper, roughly the same product at three different prices — which everybody in the room knew was false, and nobody could demonstrate. The evaluation team was being asked to recommend a ten-year commitment on the basis of prose.
The question we were asked was whether architecture could give the evaluation something firmer to stand on. Our answer was a specific proposal: model the integration reality the winner would have to live in, model each vendor's proposal against it in Sparx Enterprise Architect, and force the differences into the open where they could be scored. Roughly six weeks of work, run alongside the commercial evaluation rather than instead of it.
Modelling the ground truth first
Vendor claims cannot be assessed against an environment nobody has written down, so the first two weeks went into the current landscape — not all of it, only the part the new system would touch. With the integration team and the operations department we built the interface inventory around the incumbent: what data crosses each interface, in which direction, how often, in what format, and which business process stops working when it fails.
The inventory settled at fifteen interfaces that mattered, connecting the crew system to flight scheduling, operations control, the HR master, payroll, the training and qualifications register, the crew mobile app, hotel and travel booking, and the data warehouse. We modelled these in ArchiMate — application components, the flows between them, and the business processes each flow ultimately serves — with each interface carrying tagged values for frequency, mechanism and criticality.
Two of the fifteen interfaces surfaced only in workshop conversation; they existed in no documentation at all. One was a nightly payroll extract that left the crew system through a shared network folder and had done so, invisibly and reliably, for eleven years. Finding undocumented interfaces is not a bonus outcome of this kind of modelling — in our experience it is close to the point of it, because the undocumented interface is precisely the one a migration breaks. Both went into the model with a tagged value marking their provenance, and both went, prominently, into the requirements.
Criticality was classified with the operations people rather than the technologists, on one question: how long can the airline fly normally with this interface down? Two interfaces earned the top class — crew tracking to operations control, and rostering to the crew app — where the honest answer was measured in minutes. Payroll, by contrast, tolerates a day of latency without a passenger noticing, whatever its political weight inside the building. Writing those classifications into tagged values, with the operations director's name on the session that agreed them, gave the later scoring its weighting scheme and removed an entire genre of argument from the evaluation: the loudest stakeholder no longer set the priority, the model did.
Requirements as model elements, not a spreadsheet
The procurement had inherited a requirements spreadsheet of the usual kind: several hundred rows, accreted from stakeholder wishlists, with duplicates, contradictions and a good deal of wishful thinking. Before it could anchor an assessment it needed to become a requirements model, and the difference is more than storage.
We brought the requirements into Sparx EA as Requirement elements, merged duplicates, and split the compounds — a row demanding "full integration with HR and payroll including historical data migration" is three requirements with three different costs and, as it turned out, three different answers per vendor. The consolidated set landed at around a hundred and twenty requirements, each carrying tagged values for priority on a strict must/should/could scale, source (who asked for it and in which session), and theme. Traceability is what the model adds over the spreadsheet: using the Relationship Matrix, every integration requirement was linked to the application components and interface elements it constrains — each interface in the inventory being an element in its own right, not merely a line on a diagram, so "payroll interface fidelity" pointed at the actual payroll flow in the landscape model, undocumented folder transfer and all.
The consolidation itself changed the procurement before any vendor saw it. Forcing must/should/could discipline onto a wishlist is a political act — around thirty rows claimed to be mandatory until the operations director asked, for each, whether the airline would genuinely walk away from a contract over it. Twenty-two survived as musts. That conversation happened because the model demanded a value in a field, which is a modest illustration of something we see repeatedly: the discipline a repository imposes is often worth more than the repository.
One model per vendor, same shape
The centre of the method is a structural trick, simple to describe and demanding to keep: every vendor's proposal is modelled as an overlay on the same baseline, in the same shape, at the same level of detail. We created a package per vendor, each containing a copy of the integration landscape, and worked each vendor's written proposal into its copy: which of our components their product replaces, which interfaces they serve natively, which they serve through configuration or custom adapters, and which they do not serve at all.
Every deviation from the baseline was recorded the same way — a tagged value on the affected interface or component recording the vendor's mechanism, the evidence for it (a proposal page reference, later a workshop statement), and our confidence in it. The discipline that matters is symmetry: the moment one vendor's model is prettier, more detailed or more charitably interpreted than another's, the comparison is dead, and a diligent vendor's honesty about a gap will read worse than a careless vendor's silence. We kept a short internal checklist for it — same diagrams, same viewpoints, same assumptions written down for all three — and reviewed the three packages side by side weekly for drift.
Sparx EA earns its keep here in an unglamorous way: three parallel variants of one landscape, with shared requirements traced into each, is exactly the bookkeeping a repository does well and slideware does catastrophically. The alternative we most often see — three PowerPoint decks in three vendors' house styles — is not a comparison at all; it is three advertisements filed together.
Inside the repository
The package structure carried more of the method than any diagram, so it is worth describing plainly. One package held the ground truth — the landscape and interface inventory, owned by the client's integration lead and writable by nobody else once agreed. One held the consolidated requirements. Three sibling packages held the vendor overlays, deliberately identical in internal structure: the same sub-packages, the same diagram names, the same viewpoints in the same order, so that reviewing vendor B after vendor A felt like turning the page rather than starting a new book.
Baselines gave the process its audit trail. We took a baseline of each vendor package at three moments: when the overlay was first built from the written proposal, when the vendor confirmed it at the end of their workshop, and when scoring closed. The differences between the first two baselines are, in effect, a record of what each written proposal had left unsaid — instructive reading for the procurement team, and occasionally uncomfortable. Sparx EA compares a baseline against the current model rather than two baselines against each other, so we ran the comparison at each capture point and kept the reports — which made those gaps reviewable line by line, and the evaluation report cites two of them by name — the workshop-confirmation baseline for what each vendor agreed, the scoring-close baseline for the state the scores were awarded against — so there was never a question about which state of the model a score referred to.
Package-level security enforced the etiquette. During the assessment window, each overlay was writable only by the two of us doing the modelling; scores went in through the trace relationships, not by editing vendor claims; and after confirmation, an overlay was sealed — later corrections happened as annotated additions, never silent edits. None of this is exotic machinery. It is ordinary repository discipline pointed at a procurement, and it is precisely what a folder of slide decks cannot provide.
Walking the model with each vendor
Each vendor then got a half-day workshop, and the agenda for all three was the same: walk the diagrams, interface by interface, and correct us. Not a demo, not a presentation — their solution architect, our model, and the standing question "this is what we believe your proposal means; where are we wrong?"
The format changes the conversation in a way that is difficult to overstate. A vendor can agree with a paragraph in the abstract; a vendor cannot leave a diagram wrong when it is projected in front of them with their name on the package, because the diagram will be scored. One vendor's "standard HR connector", walked through against our flow, turned out to be a nightly file import whose worst-case lag, for a change landing just after a run, was most of a day — perfectly workable for the HR master, disqualifying if it had been the crew-tracking flow, and invisible in the written proposal. Another vendor, to their considerable credit, looked at our undocumented payroll extract and said plainly that replicating it would be custom work on their side, priced accordingly; that honesty was captured in the model with the same neutrality as everyone else's claims, and the confidence tagged value did the rest.
Every correction was made in the model during the workshop, projected live, and each session ended with the vendor confirming their overlay as amended. That confirmation mattered later — twice during contract negotiation, a claim was settled by opening the model rather than by exchanging recollections of a meeting.
Scoring in the open
Scoring ran against the confirmed models, requirement by requirement, vendor by vendor. Each of the hundred and twenty requirements received a score per vendor on a deliberately blunt scale — met natively, met with configuration, met with custom work, not met — recorded as tagged values on the trace relationships, with a one-line justification pointing at the model or workshop evidence. Blunt scales are a choice: a ten-point scale invites false precision and negotiating over decimals, while a four-state scale forces the honest question, which is what kind of yes this is.
| Score | Meaning | What it costs later |
|---|---|---|
| Met natively | Works as shipped, evidenced in the walk-through | Licence, mapping and testing |
| Met with configuration | Within the product's own configuration surface | Setup effort, usually survives upgrades |
| Met with custom work | Vendor builds or adapts; priced as such | Build, test dependency, ongoing maintenance |
| Not met | Outside the product, no committed path | A workaround the client owns forever |
Integration effort was estimated separately, per interface, from each vendor's confirmed mechanism — an interface served natively still costs mapping and testing, but little more; one served by a custom adapter carries a build, a test dependency on a third system, and a maintenance line for as long as it lives. The comparison tables and the per-vendor annexes for the evaluation report were generated from the repository with Sparx EA's document templates, on the pattern we describe in our piece on document generation from the model — which meant the evaluation committee's pack and the model could not quietly diverge, and a late correction in the model flowed into the tables with a regeneration rather than a Friday evening of manual editing.
The method does not make the decision objective, and it should not be sold that way. Judgement decides the weights, judges the evidence and owns the recommendation. What the model changes is that every judgement is written down, attached to the thing it judges, and arguable — which is what "structured" should mean in a structured assessment.
What the comparison actually showed
The comparison produced the shape we see more often than any other in vendor assessments: the best functional answer and the best integration answer were different vendors. The functional front-runner — strongest rostering optimisation, best day-of-operations tooling — needed custom work on five of the fifteen interfaces, including one of the two criticals. The runner-up covered rostering slightly less impressively and touched only two interfaces with custom work, both peripheral.
The third vendor's proposal, which had read as confidently as the others, thinned under the walk-through: crew tracking came from a partner product with its own contract, its own roadmap and its own support organisation, a structure their written response had not volunteered. They scored honestly and finished a clear third, and the model shows precisely why — which is also what makes a structured assessment defensible to a losing bidder, a property procurement teams learn to value the first time a challenge letter arrives.
The recommendation went to the steering committee as a one-page argument standing on the generated annexes: the runner-up, because the functional gap was closable by configuration over two seasons, while the front-runner's integration exposure sat on the flows where the operation bleeds money by the minute when they fail. The committee accepted it in a single meeting. Members told the evaluation lead afterwards that what convinced them was not the recommendation but the fact that, for the first time in the procurement, the three proposals looked different from each other.
What the models did after the decision
Most evaluation artefacts die on contract signature. These did not, and the afterlife was designed in. The winning vendor's confirmed overlay — their mechanisms, their workshop statements, their custom-work commitments per interface — was attached to the contract as a technical annex, giving the delivery phase an agreed description of what "integrated" means, interface by interface, written before anyone had an incentive to reinterpret it.
The same package then became the implementation baseline in the repository. As delivery proceeded, the project architects recorded actuals against it, and the deviations tell their own story: two interfaces turned out easier than assessed, one materially harder, and the custom payroll adapter came in as scoped — a small vindication for the vendor who had priced their honesty. The landscape model, meanwhile, stayed behind as the permanent record of the crew domain's interfaces, undocumented folder transfer now documented, and has since been reused by two unrelated projects that needed the same ground truth. This is the quiet compounding value of doing procurement work in a repository rather than in a deck: the by-products persist, and connect to the broader architecture practice instead of evaporating with the transaction.
When this approach is worth its cost
Honesty about the price: this was roughly six weeks of focused work — two of ground truth, one of requirements consolidation, one of vendor overlays, one of workshops, one of scoring and reporting — plus real hours from the client's integration and operations people, and half a day of goodwill from each vendor. That is proportionate for a ten-year operational commitment with fifteen live interfaces. It would be absurd for a departmental tool with two.
The scaling rule we give clients: the method pays when the integration surface, not the feature list, is where the risk lives. If the honest worry is whether the product works, run a proof of concept. If the honest worry is whether the product survives contact with your landscape — and for core operational systems it usually is — then the landscape is the test rig, and it has to be built before it can be used. A second limitation is equally worth naming: the models assess what vendors claim and confirm, not what their software does. The confidence tagged values and the workshop confirmations narrow the gap between claim and reality; only delivery closes it. We position the method as making a paper-based decision honest, not as replacing the due diligence a reference visit or a proof of concept provides.
What we would do differently
Two adjustments, both about sequence. We would consolidate the requirements before the shortlist rather than after it. The must/should/could reckoning that thinned the mandatory list happened with three vendors already selected against the unconsolidated version — no harm resulted here, but the shortlist could in principle have looked different against the honest list, and that is a risk worth removing by doing the cheap work first.
And we would schedule the vendor workshops with more air between them. Running all three inside eight days was efficient for diaries and hard on symmetry: by the third workshop we had learned better questions, and keeping the earlier vendors level meant sending each a short follow-up round with the questions their sessions had not been asked. The fairer design is a deliberate two-pass structure for everyone — walk-through, then a written confirmation round a week later — rather than pretending a single afternoon per vendor lands identically three times running.
One habit we would keep exactly as it was: writing our confidence in every vendor claim into the model alongside the claim itself. It felt bureaucratic in week three. It was the difference between an assessment and an advertisement by week six, and it is the single element of the method we now refuse to run without.
Where they are now
The new crew system went live base by base over a season and a half, against the integration annex, with no interface surprises of the kind that had been the organisation's quiet dread — the undocumented extract having been, in the project manager's phrase, the cheapest finding of the decade. The evaluation method itself has been reused twice since, by the client's own architects with the templates and package structure we left behind, most recently for a maintenance-system procurement we heard about only afterwards, which is the kind of afterwards we prefer.
If your organisation is facing a selection where every proposal says yes and a decade of integration reality will be lived with the answer, you can reach us through our contact page.
This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.