A migration, not a synchronisation
The first decision is the one that shapes everything else: is Git going to remain part of the workflow, or is this a one-way move?
Continuous bidirectional synchronisation between a Git repository and a database repository sounds appealing and is a bad idea. You now have two systems that can both accept writes, with different concurrency models, and you have to reconcile them. Every hard problem the database repository was meant to solve returns, wearing a hat.
A one-time migration is a bounded, comprehensible operation. Afterwards the database is the source of truth and the Git repository becomes a read-only archive.
The shape of a safe migration
The step that earns its keep is the dry run. Scanning and discovery are mechanical; producing a manifest of what would be imported, before anything is written, is what turns a risky bulk operation into a reviewable one.
A useful manifest lists, per file: the source repository and path, the commit it was read from, whether it parsed, its element and relationship counts, and any duplicate detected against something already in the target. Reviewing that list takes an architect an hour and prevents the class of problem that is expensive afterwards.
What you will find
Migrations of real estates surface things nobody knew were there. Expect some of these:
- Models that do not parse. Usually the result of a bad merge that was committed and never opened again.
- Duplicates. The same model in two repositories, subtly diverged, both apparently current.
- Abandoned experiments. Branches and directories nobody has touched in three years.
- Models with one element. Someone created a file and never came back.
The temptation is to migrate everything and sort it out later. Resist it: the target repository's value comes from people trusting what is in it, and seeding it with four hundred models of which sixty matter destroys that on day one. Migrate what is current and owned; archive the rest where it already is.
What does not come with you
Two things, and being clear about them up front prevents disappointment.
Git history. The commit log is a history of files. The intermediate states were never model states anyone approved, and most of them will not parse cleanly out of context. You can record where a model came from — repository, path, commit SHA — which is genuinely useful provenance. You cannot meaningfully turn a thousand commits into a thousand revisions.
Branch structure. If teams were using branches for parallel exploration, that workflow does not survive a move to a lock-based repository. This needs to be discussed before migration, not discovered after.
Keep the Git repositories, read-only, for as long as your retention policy requires. Migration does not have to mean deletion, and leaving the archive in place removes most of the anxiety from the decision.
Sequencing
Migrate one team first, ideally one that is enthusiastic and whose models are in reasonable shape. Let them work in the new repository for a few weeks. The problems that surface will be about workflow rather than data, and they are much cheaper to fix with one team than with all of them.
The second team goes far more smoothly, and by the third the migration is routine. Attempting the whole estate in one weekend produces a support queue nobody can work through and a first impression you do not get back.
Deciding what is current
The hardest part of a migration is not technical. It is establishing which of four hundred discovered models anyone still cares about, and that question cannot be answered by the migration tool.
Three heuristics get you most of the way, and they can be applied from the manifest before anything is imported:
- Last commit date. Anything untouched for two years is a candidate for the archive rather than the repository.
- Size. Models with fewer than a dozen elements are usually experiments.
- Named owner. If nobody will claim it, it is not current — and asking is a fast way to find out.
The third is the decisive one and the only one that requires people. A short exercise where each team confirms which models are theirs typically removes half the estate and produces the ownership metadata you were going to have to collect anyway.
Provenance is worth recording
Record, per imported model, the source repository, the path within it, and the commit SHA it was read from.
This costs three columns and answers the question that arrives six months later: "is this the same model that was in the old repo, or did someone change it during migration?" Without provenance that is unanswerable, and the doubt undermines confidence in the whole migration.
Running the dry run properly
Every migration tool offers a dry run and most teams treat it as a formality — run it, see no errors, proceed. The dry run is where the migration is actually decided, and reading its output properly takes longer than the migration itself.
What to look for, in the order it matters:
- Files that will not parse. Usually the result of an unresolved merge conflict committed months ago. These are not migration failures; they are models that have been broken in the repository since then, and someone has been working around them.
- Duplicate identifiers across models. Two models that were copied from each other share element IDs. Merging them into one repository will either collide or silently duplicate, and which one happens depends on the tool.
- Models with no recent commits. Anything untouched for a year is a candidate for not migrating at all. Migrating it costs nothing technically and costs a great deal in the search index, where it will surface alongside current content.
- Branches that are not merged. Work someone did and never integrated. Ask before discarding; the answer is sometimes that the branch is the real current state.
The output of a good dry run is not a green tick. It is a list of decisions for a person to make, and if the list is empty the dry run was not looking hard enough.
What to do with the Git history
The commit history is the thing people most want to bring across and the thing that translates worst. A repository stores revisions of a model; Git stores revisions of a file, and the two are not the same shape. A commit that touched four models is one Git object and four repository revisions, and a commit message written for a file diff rarely says anything useful about any of the four.
Three positions, and the middle one is right for most estates:
- Import nothing. Start clean, keep the Git repository read-only as an archive. Fast, and it discards the one thing auditors sometimes ask for.
- Import a provenance record, not the history. Each migrated model carries the source repository, path, commit hash and date it came from. One property, permanently answers "where did this come from", and points at the archive for anyone who needs the detail.
- Replay every commit as a revision. Technically possible, expensive, and it produces a revision history whose messages are about files. Worth it only where a regulator has asked for continuity of history, and worth confirming they have before committing to it.
Keeping the old repository around
The instinct after a successful migration is to delete the GitHub repository, and it is worth resisting for longer than feels necessary. Set it read-only instead, and leave it that way for a year.
The reason is not sentiment. It is that migrations surface problems on a delay: someone opens a model in month four, finds a view they remember as more complete, and the only way to settle whether something was lost is to look at the source. A read-only archive makes that a five-minute check. A deleted repository makes it an argument.
Read-only matters as much as retained. A repository that still accepts pushes will receive them, from a script nobody remembered or an architect who had it cloned. Then there are two systems of record again, which is the condition the migration existed to end.
Telling people it has moved
The technical migration completes in a weekend. The human one takes a quarter, and it fails in a specific way: a handful of people keep using the old workflow because nobody told them in a form they noticed, and their work diverges quietly.
What works is unglamorous. Tell people before, not after. Name a date when the old repository goes read-only, and hold it, because a date that slips teaches everyone the next one will too. Give each architect a single instruction — install this, open this — rather than a migration guide, because a guide gets bookmarked and the instruction gets done.
Then watch for the tell: models appearing in the repository with conventions from the old estate, or an architect asking where the branch went. Both mean someone is working from an old mental model, and both are cheap to correct in week two and expensive in month six.
Deduplicating on the way in
A Git-based estate accumulates copies. Someone forked a model to try something, someone else cloned it into a project folder, a third person exported and re-imported it and the identifiers changed. By the time a migration happens there are typically three or four versions of the important models and no obvious way to tell which is current.
Migrating all of them is the default and it is a mistake that compounds: from that point the repository contains four systems of record and search returns all four, so every reader has to decide which to trust and they will each decide differently.
The comparison that resolves it is usually mechanical. Sort candidate models by last commit date, count elements in each, and look at whether one is a strict superset of the others. In most estates one model is both the newest and the largest, and that is the answer. Where it is not — where a newer model is smaller — someone deleted content deliberately or accidentally, and that is worth ten minutes with the person who did it.
Whatever is not migrated should be listed, with a reason, in the same place as the provenance record. "We chose not to bring this" is a defensible position. "We do not know why that is missing" is not, and they are indistinguishable a year later without the list.
The rollback nobody plans
Migration plans describe the forward path in detail and treat rollback as a formality, because the source repository still exists and reverting seems to mean pointing people back at it. That is true for the first day and stops being true immediately afterwards.
The moment an architect publishes into the new repository, the two systems have diverged and rolling back means losing that work or manually exporting it. Within a week there will be a dozen such changes spread across several people, and nobody will have a list.
The practical answer is to make the rollback window explicit and short. For the first working week, the new repository is in use but everything published to it is also exported nightly to a file the old process could ingest. That is a few lines of scripting and it buys a genuine ability to reverse. After that week, rollback is not a plan and should not be presented as one — the honest position is that you fix forward.
What improves once the migration lands
Migrations are justified by what stops hurting, and it is worth being concrete about which pains disappear immediately and which need further work, because teams that expect all of it on day one are disappointed by a successful migration.
Immediate: merge conflicts stop existing, because concurrent editing is now handled by a lease rather than by a three-way merge on XML. "Which file is current" stops being a question. Access can be scoped to a project rather than to a repository, so a supplier can be given one domain instead of everything.
Needs further work: the estate is not more accurate on Monday than it was on Friday. Duplicate elements survive migration unless someone removes them. Ownership is as unpopulated as it was. And nothing is published to anyone outside the architecture team until a publication path is built, which is a separate project.
Setting that expectation before the migration is what stops the question three months later about why the repository has not fixed the architecture practice.
The week after
Migration guides end at the cutover; estates live in the week after it, and that week has a shape worth knowing in advance so its events read as normal rather than as omens.
Expect three kinds of traffic. First, the login stragglers — the architect back from leave who missed every announcement and opens the old repository, finds it read-only, and needs the two-minute orientation everyone else got last Tuesday. This is why the read-only banner on the old estate carries a link and a name, not just a lock. Second, the workflow friction reports: the lease expiring during someone's long modelling session, the publish comment they resent writing, the keyboard shortcut that moved. Triage these honestly — some are settings to tune, some are habits adjusting, and the difference matters because tuning the platform to recreate file-era behaviour defeats the migration. Third, and most valuable, the discoveries: within days someone searches across the whole estate for the first time and finds the duplicate integration two teams built in parallel, or the reference to a platform that retired last year. Publicise these finds; they are the migration paying for itself in public, in its first week.
Then hold the retrospective while memory is fresh — twenty minutes, three questions: what broke, what surprised, what should the next migration (there is always a next one — the sister team, the acquired company's models) do differently. The notes from that meeting are the cheapest consulting the organisation will ever produce for its future self, and writing them is the last act of the migration proper. After that, the estate is simply how things work — which was the destination all along.
Done in this order, the migration is a fortnight of care rather than a quarter of risk — and the estate that lands on the other side keeps everything Git taught the team about history and discipline, minus the parts XML was never going to support.