Mapping the Road to Cloud Migration

A hosting contract with an end date

The organisation was a media company โ€” broadcast and digital publishing โ€” running its estate across two hosted data centres under a contract with a firm end date a little under two years away. The board had decided not to renew: the destination was public cloud, the business case was signed, and a migration partner was being procured. What the organisation could not produce, when the programme asked for it, was a reliable answer to the first question any migration asks: what do we actually run, what does it talk to, and what breaks if we move it?

There was a CMDB, of a sort. There were architecture diagrams, of various ages, in various drawers. There was a great deal of knowledge in the heads of a small number of long-serving engineers, several of whom were within sight of retirement. What there was not was a single place where an application, its dependencies, its infrastructure and its migration fate were connected โ€” and without that, wave planning was going to be done by hunch, which is how migrations end up moving a scheduling system in March and discovering in April that it fed the playout chain nobody moved with it.

Our brief was the model, not the migration: build the dependency and deployment picture in Sparx EA, make the migration choices expressible and visible in it, and hand the programme an instrument for planning waves and tracking transition states. The migration partner would execute; the model would keep everyone honest about sequence and impact.

Reconciling the inventory

We started from the CMDB export, knowing it would be wrong in both directions, because CMDBs usually are. The export listed rather more than two hundred applications; workshops with domain leads shortened and lengthened the list at the same time. A few dozen entries were duplicates, decommissioned systems or things that were really modules of something else. A comparable number of real applications were missing entirely โ€” mostly departmental tools and, tellingly, several of the systems closest to the broadcast chain, which had grown up outside IT's line of sight.

Each surviving application became an application component in the repository, and the reconciliation decisions were recorded as we went: an element either carries a tagged value linking it to its CMDB identifier, or carries a note explaining why it exists in the model without one. Alongside the identifier we kept a deliberately short set of tagged values โ€” business owner, technical owner, criticality, data sensitivity, current hosting location โ€” populated in the workshops rather than by survey, because surveys about ownership return fiction and workshops return arguments, and the arguments are the point. Getting two names against every application took about six weeks of sessions threaded between people's day jobs, and was some of the most valuable time in the engagement.

Mapping what talks to what

Dependencies came next, domain by domain: newsroom, playout, media asset management, advertising, corporate. In each session we put the domain's applications on a working diagram and asked the people who run them one question repeatedly โ€” when this system does its job, what does it need, and what needs it? Interfaces went into the model as information flows between application components, each with a short name, a direction, a tagged value for mechanism โ€” API call, database link, message queue, file transfer โ€” and a tagged value naming the dependent end, because the direction data moves is not always the direction dependency runs. The file transfers took the longest to surface and mattered the most; scheduled file drops are the connective tissue of broadcast estates, and they are precisely the dependencies nobody documents because they have worked unattended for a decade.

Figure 1: The baseline landscape โ€” applications in two hosted data centres with the information flows between them, including the cross-site dependencies that constrain migration order
Figure 1: The baseline landscape โ€” applications in two hosted data centres with the information flows between them, including the cross-site dependencies that constrain migration order

By the end of this phase the repository held several hundred flows, and the first honest picture of the estate existed: not a diagram someone drew once, but a model that could answer questions. The cross-site flows โ€” application in one data centre depending on something in the other โ€” got particular attention, because each one is a latency and sequencing constraint the moment one end of it moves to cloud and the other does not.

Keeping several hundred flows healthy needed housekeeping from the start. A nightly script checked for the standard rot: candidate duplicate flows โ€” two flows between the same pair with similar names, queued for a person to confirm, flows whose mechanism tagged value was empty, elements with no relationships at all. The weekly modelling session opened with that report, and clearing it rarely took ten minutes โ€” the point of nightly checks being precisely that no week's debt is ever large. An estate model that is wrong in small known ways stays trusted; one that is wrong in unknown ways stops being consulted, quietly, and nobody tells you.

Tying applications to the machines they run on

A migration is a physical event, so the logical picture is not enough. We modelled the hosting layer as nodes โ€” virtualisation clusters, database servers, storage โ€” within a grouping per data centre, and connected each application to what it runs on. This is the least glamorous modelling in the engagement and it pays for itself twice. Once during wave planning: two applications sharing a database server are, for migration purposes, more entangled than their interface diagram admits, and shared infrastructure surfaced a dozen such couplings. And once during the estate's endgame: the hosting contract's cost tail is driven by what remains powered on, so "what is still on this cluster" is a question the model needed to answer month by month.

We stopped the technology modelling at the level migration decisions needed โ€” clusters, database platforms, storage classes, the handful of specialised broadcast appliances โ€” rather than descending into full infrastructure inventory. The rule of thumb we applied, here as in most enterprise architecture work: model to the level of the decision you are supporting, and not one level deeper, because every level below the decision is maintenance without a customer.

How the repository itself was set up

A word on plumbing, because migration models have more readers than authors and the setup should reflect that. The repository ran on the organisation's existing Pro Cloud Server, with a package structure that separated the stable from the volatile: a reference package for the reconciled application and infrastructure inventory, a package per domain for dependency views, a transitions package for plateaus, gaps and work packages, and a decisions package that recorded disposition rationale and the programme's structural choices. Package-level security kept editing rights narrow โ€” the two programme architects and ourselves โ€” while every programme member could read, and the wider organisation reached the monthly views as exports.

Each agreed plateau was captured at the moment of its approval by baselining the reference, domain and transitions packages together โ€” EA baselines are taken per package, so the plateau snapshot is a set of baselines made in one ceremony, which gave the programme something migrations rarely have: the ability to answer "what did the plan say in January" from the tool rather than from an argument. Restoring context for a disputed disposition took minutes on the two occasions it was needed, and the second occasion visibly shortened the dispute. The discipline costs one baseline per milestone and repays itself the first time anyone says "but we agreed".

A disposition for every application

With the landscape in place, each application received a migration disposition โ€” the familiar vocabulary: rehost, replatform, refactor, repurchase, retire, retain โ€” held as a tagged value with an agreed value list, plus a rationale note and a status recording whether the disposition was proposed, agreed or executed. Disposition workshops ran per domain with the application owners, IT and the incoming migration partner at the table, working through the portfolio application by application with the model's dependency picture projected behind the discussion.

The dependency picture changed decisions in the room. A departmental tool pencilled in for retirement turned out to feed the advertising system's nightly reconciliation; its retirement moved behind a replacement interface. An ageing rights management system everyone wanted to refactor was serving four downstream consumers that could not absorb change on the programme's timeline; it became a rehost now, refactor later, and the later got a plateau of its own. Around a tenth of the portfolio ended in retire, which is on the low side of what these exercises usually find, and mostly reflected an estate that had already been through a consolidation a few years earlier.

For communication, EA's diagram legends earned their keep: legends keyed to the disposition tagged value colour every application on a view automatically, so a domain's migration picture is one diagram with no manual colouring to go stale. The same views, filtered by criticality or data sensitivity, answered the security team's questions about what was moving when.

Modelling the destination, not only the exit

Exit programmes obsess over what leaves and forget to describe where it lands, so we modelled the destination early. The cloud landing zone appeared in the repository as its own technology grouping โ€” network segments, the container platform, the managed database services, the storage tiers โ€” at the same deliberate altitude as the data centre modelling: enough structure to place applications and reason about their neighbours, no pretence of being an infrastructure-as-code inventory. Each application with a rehost or replatform disposition gained a planned deployment link to its target, so every wave view could show both ends of every move.

For the handful of genuinely contentious systems, the model carried competing options side by side. The media asset management platform had two credible targets โ€” a lift of the existing stack onto cloud infrastructure, or a rebuild of its transcode pipeline onto the managed container platform with object storage behind it โ€” and each existed as a small candidate view with its cost and risk notes attached. The steering committee chose the lift, with the rebuild parked as a post-migration plateau; the unchosen option remains in the model with the decision that parked it, which is precisely the kind of institutional memory that evaporates when options live in slide decks. Modelling both cost two diagrams and an afternoon; the debate it structured had been running unstructured for a year.

The target modelling also gave the security and compliance teams their entry point. Data sensitivity and data residency tagged values, carried by every application since the inventory phase, could be read against target storage and network placement, and the two systems whose archival obligations required in-country storage were flagged by a model search rather than by late-stage audit โ€” the cheap kind of discovery, made early because the destination was modelled at all.

Waves follow dependencies, not org charts

The programme's first draft wave plan, made before the model existed, had grouped applications by department โ€” newsroom first, advertising second โ€” because departments are how organisations think. The model made the problem with that visible within a day: the departments share systems and data flows so heavily that no department is a closed set, and a wave that moves one side of a chatty interface pays for it in latency, VPN complexity and double-running costs until the other side follows.

We re-cut the waves from the dependency graph instead, using a script over the automation API to pull the flow and deployment relationships into a clustering analysis and propose groupings that keep heavily connected applications in the same wave โ€” the mechanics are the standard fare of our automation API work. The proposals were then adjusted by hand in workshops for the constraints no algorithm sees: the broadcast calendar, contract renewal dates on software licences, the fact that nobody migrates the playout chain in the same quarter as a major sporting event. A standing model search lists every flow that crosses a wave boundary in the wrong direction โ€” an application scheduled early depending on one scheduled late โ€” and that list, reviewed monthly, is the plan's immune system. It started at around forty items; each was either resequenced, given an interim interface, or accepted with eyes open.

WaveContentWhy this order
PilotLow-risk corporate toolsProve the pipeline, train the teams
Wave 1Digital publishing clusterCloud-friendly stack, few ties to broadcast
Wave 2Media asset management and its satellitesLargest data gravity, longest lead time
Wave 3Broadcast-adjacent and specialised systemsHardest dependencies, a few appliances moving to a small co-location facility rather than cloud

Transition states as plateaus

To make the journey visible rather than implied, we modelled it with ArchiMate's implementation and migration concepts: a baseline plateau, one plateau per wave completion, and a target plateau, with each plateau linked to the applications and technology still on-site at that point โ€” membership derived from the wave tagged values by script rather than curated by hand, and gaps recording what changes between adjacent states. Work packages tie the waves to the programme's delivery plan, so the model and the programme plan describe the same reality at different altitudes.

Figure 2: The migration roadmap as plateaus โ€” from baseline through pilot and two transition states to the target, with the work packages that move the estate between them
Figure 2: The migration roadmap as plateaus โ€” from baseline through pilot and two transition states to the target, with the work packages that move the estate between them

The plateaus turned out to be the concept the non-architects took to fastest. A disposition list is abstract; "this is the estate on the day the media asset management wave completes, and these eleven systems are still in the data centre" is a picture a finance director can interrogate. The cost conversation about the hosting tail โ€” what remains, until when, at what monthly rate โ€” ran off the plateau views for the rest of the engagement.

A model the steering committee actually used

Monthly, a small set of generated views went to the steering committee: the disposition-coloured landscape per domain, the current plateau against plan, the cross-wave dependency exceptions, and a one-page list of what had changed in the model since the last meeting. Generating rather than drawing them mattered โ€” the committee learned quickly that the pictures were the database, not an artist's impression of it, and started asking questions of the model in meetings, which is the moment an architecture repository stops being the architecture team's diary and becomes programme infrastructure.

The same discipline governed change. When a disposition changed, it changed in the model with a note and a date, and the next month's pack reflected it automatically. Two versions of the migration plan never got the chance to develop, because no other artefact carried dispositions or wave assignments.

One habit we imposed on ourselves deserves passing on: every question the committee asked that the pack could not answer was written down, and the next month's pack answered it or said why not. Half a dozen views in the final pack exist because a finance director or a programme manager asked a question the architects had not thought to anticipate. Steering packs decay into ritual when they stop being shaped by their readers; this one stayed alive because its readers were, in a small way, its editors.

What the model could not know

Honesty requires a section on the model's blind spot. The dependency picture was built from human knowledge โ€” workshops, interviews, existing documentation โ€” and human knowledge of a twenty-year-old estate has holes. The pilot wave found two of them: undocumented scheduled file transfers, one feeding a compliance archive, discovered when the source system moved and the transfers stopped. Neither caused lasting damage; both were embarrassing in exactly the way that teaches.

The response was procedural, not heroic. Each wave now opens with a period of network observation on the systems due to move โ€” long enough to catch the weekly and monthly scheduled jobs โ€” and ends with a reconciliation step: the migration partner's observation, actual connections seen on the wire before and after the move, is compared against the modelled flows, and the differences are triaged into the model. The model gets truer wave by wave. We are wary of promising more than that: a repository built from interviews is a map of what people know, and treating it as a map of reality is how maps get people lost. The model's job is to make the programme's knowledge explicit, inspectable and improvable โ€” which it did โ€” not to be omniscient, which it cannot.

A dependency model built from workshops is complete enough to plan with and never complete enough to trust blindly. Budget for the reconciliation step after every wave; the model earns accuracy the same way the programme earns confidence โ€” incrementally.

Where the programme stands

At the time of writing, the pilot and the first wave are complete and the media asset management wave is in execution. The hosting exit date still holds on the current plan, with the contingency the plateau views make visible rather than hidden in a spreadsheet. We would rather report that plainly than describe the migration as finished; it is not, and the model's usefulness in the remaining waves is precisely the thing the earlier waves were building evidence for.

What has already finished is the knowledge transfer. The programme's own architects maintain the model day to day โ€” dispositions, flows, plateau content โ€” with our involvement reduced to a monthly model health review: orphan checks, tagged value completeness, the cross-wave exception list. The scripts and searches ship with a runbook, and the reconciliation procedure is the migration partner's contractual step rather than our personal habit. An instrument the client cannot operate without its maker is an instrument half-delivered.

What we would do differently

We would bring the migration partner into the disposition workshops from the first session rather than the third; their early absence meant a handful of replatform decisions were revisited once their engineers saw them, which cost calendar time the programme did not have. We would also start the file transfer hunt earlier and more systematically โ€” grepping scheduler configurations and transfer service logs in week one rather than trusting workshops to surface them โ€” since that is where both pilot surprises came from, and mechanical sources find what memory does not.

And we would set firmer expectations about the CMDB. A quiet assumption existed that the model would be synchronised back into the CMDB continuously; we deferred it to focus on the migration itself, and the deferral was right, but leaving the assumption unspoken cost an awkward meeting in month five. The eventual agreement โ€” model leads during the programme, one reconciliation into the CMDB at each plateau โ€” should have been written down in week one, next to the other one-hour decisions that become one-week arguments when postponed.

If your organisation is heading for a hosting exit or a cloud programme with an estate nobody can fully describe, the dependency model is the part that cannot start too early โ€” you can reach us through our contact page.

This case study describes a representative engagement pattern. Organisational details are illustrative and do not identify a specific client.