Deploying an Architecture Repository in a Regulated Bank

The deployment is the easy part

A repository server, a management portal, a PostgreSQL database, behind whatever reverse proxy the organisation already uses. Containerised, usually against a managed database service. In a functioning platform team this is a day or two of work.

Figure 1: The deployment shape, and the two things that catch people
Figure 1: The deployment shape, and the two things that catch people

What takes months is everything around it, and it is entirely predictable — which means it can be prepared for.

The two technical things that catch people

Corporate certificate authorities

In most banks, TLS traffic is inspected, which means certificates are issued by an internal CA. Browsers trust it because it is pushed by policy. Your application's runtime does not, unless someone told it to.

The symptom is that the platform cannot reach the identity provider's discovery endpoint, with a certificate validation error that looks like a configuration mistake. Plan for a documented way to add CA certificates to the container's trust store, and test it early.

Egress proxies

The server usually cannot reach anything directly. If it needs the IdP, a GitHub Enterprise instance for migration, or a licence endpoint, each of those needs a proxy configuration and a firewall rule, and each rule is a ticket with a lead time.

Enumerate every outbound connection the platform makes, with destination and purpose, before the first review meeting. Being asked "what does it talk to" and answering precisely changes the tone of the whole engagement.

What the review will ask for

Largely the same list everywhere, and having it ready is the difference between six weeks and six months:

ArtefactWhat it is for
Threat modelShows you thought about attack paths before they did
Software bill of materialsCycloneDX or SPDX; expect a vulnerability scan against it
Data classificationWhat the repository holds and how sensitive it is
Authentication designFlow diagram, token lifetimes, session revocation
Authorization modelRoles, scoping, and how effective permissions are derived
Logging and auditWhat is recorded, retained how long, and who can read it
Backup and recoveryRPO, RTO, and evidence of a tested restore
Penetration testOften required before production for anything internet-adjacent

The data classification conversation

This one is worth anticipating because the answer is not obvious and the wrong answer is expensive.

An architecture repository usually contains no customer data. It does contain a complete description of the bank's application landscape, its integration points, its security controls and — by omission — its gaps. Security teams frequently classify that higher than the architecture team expects, and the classification drives encryption, access review cadence and hosting constraints.

Have the conversation early. Discovering after deployment that the repository is classified in a band that forbids the hosting model you chose is a rebuild.

Sequencing that works

  1. Non-production first, with real SSO against the test IdP. Identity integration is where surprises live.
  2. One team, real models, for several weeks. This produces the operational questions the review will ask anyway.
  3. The formal review, with the artefact list above already prepared.
  4. Production with a small scope, expanded once backup and restore have been demonstrated, not merely configured.

The temptation is to run the review in parallel with the pilot to save time. In practice the pilot answers half the review's questions, so running it first makes the review shorter — and the answers are evidence rather than intentions.

The questions that come from outside security

Security review is the one people prepare for. Three other functions will also have questions, and being ready shortens the whole process.

FunctionWhat they askPrepare
Data protectionDoes it hold personal data?Owner fields, audit actors, and your retention position
ProcurementSupport, escrow, exitLicence terms and what happens to the data if you leave
OperationsWho runs this at 3am?Runbook, health endpoints, escalation path
Records managementHow long is this kept?Retention for revisions, baselines and audit

The data protection conversation is the one most often skipped and it is genuinely relevant: an architecture repository holds names in ownership fields and identities in every audit event, which is personal data even though nobody thinks of the system that way.

Sizing

A practical note, since it comes up in every deployment discussion and the answers are undramatic.

An architecture repository is small by enterprise standards. Models are measured in megabytes, revisions accumulate slowly, and concurrent users number in the tens rather than the thousands. The database is almost always over-provisioned if sized by habit.

What does need attention is publish concurrency: model loads and publishes are bursty and hold connections for the length of a transaction. Size the connection pool against simultaneous publishes rather than against total users, and the rest takes care of itself.

The evidence pack, assembled before it is asked for

Every regulated deployment ends in the same request: show us the evidence. Assembling it reactively takes weeks and produces gaps. Assembled during the build it takes an afternoon, because everything in it is a by-product of work already being done.

ArtefactWhere it comes from
Software bill of materialsThe build pipeline, in CycloneDX or SPDX. Generated per release, not written.
Threat model with review datesA document, but a short one, and the review dates are what the reviewer checks.
Backup and restore test recordThe restore rehearsal. An untested backup is not evidence of anything, and reviewers know to ask when it was last exercised.
Access review recordA periodic export of who holds which grant, signed off by someone accountable. Quarterly is the usual expectation.
Audit log retention statementOne paragraph naming what is logged, for how long, and who can read it.
Data classification decisionWritten down with its rationale. The rationale matters more than the level, because the level will be challenged.

Change management, and the friction nobody budgets for

The technical deployment takes days. Getting a change record through a bank's process takes longer, and the delay is not incompetence — it is the process working. Budgeting for it changes the plan.

Three specifics that consistently surprise people on their first regulated deployment:

  • Standard versus normal change. The first deployment is a normal change with a CAB slot. Getting subsequent upgrades classified as standard changes is worth doing early, because it converts every future release from a three-week process into a ticket.
  • Back-out plan as a hard requirement. For a database with schema migrations, "restore the backup" is often the honest answer, and saying so plainly goes down better than an implausible reversible-migration story.
  • Freeze periods. Most banks freeze around quarter and year end, and a deployment scheduled into one will not move the freeze. Find the calendar before promising a date.

Where the repository sits in the disaster recovery tier

This question arrives late and reshapes the design, so it is worth settling early. An architecture repository is almost never tier one — no customer transaction fails if it is down for a day — and teams sometimes argue for a higher tier out of pride, which buys expensive infrastructure and a recovery obligation nobody wanted.

The honest position is usually tier three: recovery within a business day, with a recovery point of a few hours. That is achievable with nightly database backups plus transaction log shipping, and it is defensible in a review.

The exception worth arguing for is the audit trail. If the repository is the system of record for architecture decisions in a regulated process, losing four hours of audit events is a different kind of problem from losing four hours of modelling work. Splitting the recovery objective — tighter for the audit log than for the model — is unusual but it is the argument that holds up.

Third-party risk assessment, from the vendor's side of the table

If the repository is a commercial product, it goes through vendor risk assessment, and that process asks questions about the supplier rather than the software. Knowing which ones arrive shortens the cycle considerably.

  1. Where is data processed and stored? For an on-premises deployment the answer is "your data centre", which ends most of the questionnaire. Say it early.
  2. What happens if the supplier fails? Source code escrow, or an open data format, or both. A repository whose contents can be exported to a standard exchange format answers this without escrow.
  3. Sub-processors. An on-premises product with no telemetry has none. If it phones home for licence checks, that is a sub-processor and it belongs on the list.
  4. Support access. Whether supplier staff can reach the data, under what controls, and whether it is logged. "No standing access, break-glass only, logged" is the answer that passes.

Who owns it after go-live

The question that derails more regulated deployments than any technical issue is who runs the thing on the Monday after launch. Architecture teams assume IT operations will take it. Operations assume the architecture team built it and can run it. Neither assumption is written down, and the gap surfaces at the first patching cycle.

The split that works divides on the boundary between the platform and its content. Operations own the database, the host, the backups, the patching and the monitoring — everything they already do for twenty other systems, with this one added to the list. The architecture team owns the permission model, the metamodel, the publication schedule and the content itself.

What needs stating explicitly is the middle: user provisioning and access reviews. It looks operational and is actually a governance decision, because deciding which architect may edit which domain requires knowing what the domains are. Leaving it with operations produces the over-broad default grant that shows up in every audit.

Write this down before go-live, in whatever form the organisation recognises — a service description, a RACI, an operating model page. Not because the document will be read, but because writing it forces the conversation to happen while there is still time to change the answer.

The first audit, and what it actually examines

Internal audit will look at the repository within the first year, and the examination is narrower and more procedural than most teams expect. They are not assessing whether the architecture is good. They are assessing whether the controls described in the deployment documentation are the controls that operate.

That means the findings come from gaps between the written process and the observed one. The access review that was supposed to be quarterly and last happened in March. The backup restore test that the runbook requires annually and has never been performed. The audit log retention set to ninety days when the policy says two years. None of these is an architecture problem and all of them are findings.

The corollary is that promising less in the documentation is often the better strategy. A stated quarterly access review that happens quarterly is a clean finding; a stated monthly review that happens twice a year is an exception with a remediation plan. Describe the cadence you will actually sustain.

The one thing worth over-preparing is the ability to answer "who could see what, on this date". It is the question that distinguishes a repository with a real audit trail from one with a log file, and it is asked in almost every review.

Cost, and the line item nobody forecasts

The licence and the infrastructure are easy to forecast and are not where the budget goes wrong. Three costs consistently arrive unbudgeted in regulated environments.

The security review itself. Penetration testing, architecture review, third-party risk assessment — each has a real cost and, in banks that charge internally, a real charge. For a first deployment this can approach the software cost.

Migration of existing content. Every estate has years of models in files, and moving them is never a straight import. The models disagree with each other, use conflicting conventions, and contain duplicates that only become visible once everything is in one place. Budget this as a project, not a task.

The first year of running it. Someone has to own the permission model, curate the metamodel, and answer questions. This is a fraction of a person and it is a real fraction. Deployments that budget zero for it produce a repository that works technically and decays organisationally, which is the more expensive failure because it takes two years to become visible.

What the security team ends up liking

A closing observation from the other side of the table, because the review relationship does not end at go-live and it pays to know where it lands. Twelve months after deployment, the security functions that scrutinised the repository hardest are reliably its heaviest users — and the conversion follows a pattern worth anticipating.

It starts with an incident or an assessment where someone needs the blast radius of a platform in minutes, and the repository answers a question that used to take a week of emails. Then the threat modellers discover the current-state application landscape is maintained by someone else, for free, with owners attached. By the second audit cycle, the same reviewers who challenged the deployment are citing it: the access recertification runs off the repository's grant report, the threat model references its catalogues, and the annual review of the review — the meta-layer regulated banks excel at — points to the repository as evidence that the architecture control operates.

This matters for how the original review is conducted. The team across the table are not gatekeepers to be survived; they are the future power users, and every concession made grudgingly — the audit export, the scoped access, the classification exercise — is infrastructure they will later depend on. Practices that understand this build the review's demands properly instead of minimally, and collect the dividend for years: a security function that defends the repository's budget, because the repository quietly became one of its own controls.

The pattern across all of it: in a regulated deployment, the artefact under review is never really the software — it is the practice's own operational maturity, with the repository as the occasion. Prepare the evidence, respect the sequence, budget the friction, and the review becomes what it should have been all along: two organisations agreeing, carefully, that a useful thing can be run safely.