A Publication Pipeline, End to End

Six steps

A publication that runs unattended every week needs to be a pipeline rather than a script, mostly because of what has to happen when something goes wrong. Six steps cover it.

Figure 1: The pipeline, and the two points at which it is allowed to abort
Figure 1: The pipeline, and the two points at which it is allowed to abort

1. Validate

Structural checks against the repository before anything expensive happens. Unresolved references, orphaned views, broken geometry. Cheap, fast, and able to stop the run before you have spent twenty minutes on an extraction that was never going to produce a good portal.

2. Extract

Read the model into a snapshot. Keep this separate from generation, and keep the snapshot: it makes the next step re-runnable without touching the repository, which turns template work from a six-minute cycle into a two-second one.

3. Generate

Render pages, diagrams, search index, catalogues. Pure function of the snapshot plus configuration. No repository access, no network, fully testable.

4. Verify

Sanity-check the output before it replaces anything. Page count within range of the last run, index present and non-trivial, no diagram files of zero length. This step exists to catch the half-successful generation.

5. Deploy

Replace the live portal, ideally atomically. Everything in the earlier steps was reversible; this is the one that is visible to readers.

6. Notify

Record what happened and tell someone. Success as well as failure — see below.

Only two steps may abort

Validation and verification. Everything else either succeeds or is a genuine error that should surface as one.

The reason for being strict about this is that partial success is the enemy. A run that extracts 80% of the model and publishes it produces a portal that looks complete and is not — and no reader can tell. Better to fail the run, leave last week's portal in place, and let someone look at it.

"The previous portal stays in place" is the correct failure behaviour. A week-old portal is a known quantity. A portal missing a fifth of its content is worse than no portal at all, because it will be trusted.

Notify on success too

The instinct is to alert only on failure, on the reasonable grounds that nobody wants a weekly email saying everything is fine.

The problem is the failure mode where the job does not run at all — the scheduled task was disabled, the server was rebuilt, the credential expired in a way that prevented startup. Silence then means both "fine" and "nothing happened", and you cannot distinguish them.

The cheapest fix is not an email; it is putting the generation timestamp on every page. A stalled schedule becomes visible to every reader immediately, which is a far better detector than any monitoring you would have set up.

Keep configuration out of the code

Scope, branding, perspectives, target and schedule are configuration. They change on a different clock from the pipeline and are usually changed by different people.

The practical test: can someone change the portal's accent colour, or add a package to the excluded list, without editing a script and without you? If not, every small request routes through one person, and that person becomes the reason the portal stops evolving.

Where the time actually goes

When a pipeline like this is slow, it is almost never slow evenly. In every installation we have measured, one step dominates and the other five are rounding errors, and the dominant step is nearly always extraction. Reading a few thousand elements out of a repository over an automation interface is chatty work: each element, each connector, each tagged value is its own round trip unless you go out of your way to batch them, and most extraction code starts life as a loop that does not.

This matters for the design in a specific way. If extraction takes twenty minutes and everything else takes forty seconds, then any workflow that re-runs the whole pipeline to see a small change is unusable, and people will start editing the generated output by hand — which is the beginning of the end, because the next scheduled run silently erases their work. The cure is structural rather than heroic: keep the snapshot, and make every step after extraction re-runnable from it. Template tweaks, branding changes, a new catalogue — all of these become a two-second cycle against the cached snapshot, and the twenty-minute extraction runs once a week when it is genuinely needed.

If extraction itself is the problem, measure before optimising. In our experience the first factor of ten comes from batching queries rather than walking the object model one element at a time, and the second comes from extracting only the packages in scope instead of the whole repository. Beyond that you are into choosing a different extraction path altogether, which is a bigger decision than a performance tweak and deserves its own analysis.

The snapshot is the pivot of the design

It is worth being explicit about why step 2 produces a file rather than feeding step 3 directly, because the snapshot quietly solves four problems that have nothing to do with each other.

Figure 2: One extraction, four consumers: generation, diffing, replay and audit evidence
Figure 2: One extraction, four consumers: generation, diffing, replay and audit evidence

The first is the re-runnability described above. The second is diffing: two snapshots can be compared field by field, and the comparison is the cheapest change log you will ever build. "What changed since last week's portal" is a question every reader eventually asks, and answering it from snapshots costs a few dozen lines of code, where answering it from the repository means reconstructing history a modelling tool may not even keep.

The third is replay. When someone disputes what the portal said in March — and in a governance context, someone eventually will — you regenerate March's portal from March's snapshot and settle the question in minutes. This is the property that turns a publication pipeline into audit evidence rather than just a website, and it costs nothing more than not deleting old snapshots.

The fourth is testing. A generation step that is a pure function of snapshot plus configuration can be tested with a fixture snapshot in a way that code talking to a live repository never can. Our own test suites run against a small, deliberately ugly fixture model — orphaned views, duplicate names, a diagram with no elements — because the fixture is where you encode every model pathology you have ever met, and the pipeline that survives the fixture survives the client.

On format: anything structured and stable will do, but flat JSON files have a virtue that is easy to underrate, which is that they diff readably in tools everyone already has. A SQLite snapshot is more pleasant to query; a JSON snapshot is more pleasant to argue about. We have shipped both, and for portals whose snapshots double as evidence, JSON wins.

Verification rules that earn their keep

Step 4 is easy to describe and easy to get wrong in both directions. Too lax, and it waves through the half-successful generation it exists to catch. Too strict, and it fails the run every second week for reasons nobody cares about, and within a month someone adds a flag to skip it — at which point you have a pipeline with a verification step in the documentation only.

The rules that have earned their keep across our installations are few and blunt. Page count within a tolerance of the last run — twenty per cent covers reorganisations without letting a gutted portal through. Every diagram file non-empty, because a zero-byte SVG is the classic signature of a renderer that crashed mid-run. The search index present, parseable, and within a similar tolerance of the page count, because a portal whose search silently indexes nothing looks fine until the first query. And a sample of internal links resolved against the generated output — a sample, not all of them, because link checking is the step that otherwise grows to dominate the runtime.

Two disciplines keep the step honest. First, every rule that fires must name the artefact and the number: "index has 14 entries, floor is 380" is actionable at eight in the morning, "verification failed" is not. Second, thresholds ratchet rather than being edited ad hoc. When the portal legitimately doubles in size, the tolerance recentres on the new run; when someone wants to lower a floor, that is a change with a reviewer, not an edit made under deadline pressure. The failure mode this prevents has a name in operations — normalisation of deviance — and publication pipelines are as prone to it as anything else that runs unattended.

What verification is not is a model quality gate. Naming conventions, missing descriptions, unapproved lifecycle states — those belong in step 1, against the repository, where the message can route to the architect who owns the package. By step 4 the model is out of the picture; the only question is whether the generator did its job. Keeping the two apart keeps both sets of messages short, and there is a separate discussion of what belongs in the model-side gate.

Rolling back a bad publication

Eventually something wrong reaches readers anyway. Verification checks shape, not truth: a model change that is structurally valid and factually wrong — a system marked retired that is very much in production — sails through every gate you can build, gets published, and gets noticed by the system's owner within the hour. What happens next depends entirely on decisions made before it happened.

If deployments are atomic and outputs are kept, the answer is boring: put the previous output back, fix the model, let the next run republish the truth. Keeping the last handful of generated portals costs a few hundred megabytes and turns "roll back" from an incident into a file copy. If the portal is deployed as a static site, this is one directory swap or one revert of a deploy, which is most of the argument for static deployment in one sentence.

The subtler question is whether to roll back at all. A portal that is wrong about one system and current about six hundred others may serve readers better than last week's portal, which is wrong about this week in a hundred small ways nobody has enumerated. The honest answer is that it depends on the blast radius of the error, and the person who should decide is the practice owner, not the pipeline. What the pipeline owes them is the ability to execute either decision in five minutes — republish from the previous snapshot, or hotfix the model and trigger an off-schedule run. A pipeline that can only go forward, on schedule, leaves them managing an incident with a tool that has one gear.

One habit makes all of this easier: treat an off-schedule run as normal rather than exceptional. If triggering the pipeline manually requires remembering an incantation, finding the right machine and holding your breath, it will be put off until the scheduled slot — and the window during which the portal is known to be wrong stretches from hours to days. The manual trigger should be the same pipeline, same gates, same notifications, differing only in when it starts. Runs that skip the gates because "it's urgent" are how the second incident gets published while everyone is watching the first.

What good looks like after six months

Nobody thinks about it. The portal is current because it republishes on a schedule, the schedule is visible on the page, failures are loud, and content changes happen in the repository where they belong.

If your publication still requires a person on the day it runs, it is a task rather than a pipeline, and it will stop the first time that person is on holiday.

Starting from the script you already have

Almost nobody builds this from a blank page. The usual starting point is a script that has grown for two years: it connects, extracts, generates and copies, all in one file, and it mostly works. The good news is that the six-step shape can be reached by refactoring rather than rewriting, and in a particular order.

Split extraction from generation first, with the snapshot as the seam — it is the cut that pays for every other one, because the moment generation runs from a file, you can iterate on it without touching the repository. Add verification second; it is twenty lines and it starts catching things the first month. Make the deploy atomic third. Validation, notification and externalised configuration can each wait until the pain arrives, and it will announce itself clearly: a bad model wasting a run, a silent failure nobody noticed, a colleague asking you to change a colour.

What does not survive the refactor is the habit of editing the script for content decisions. The day scope changes ship as configuration instead of code edits is the day the pipeline stops having a single human dependency — which was the point of the exercise, and is worth stating as the acceptance test before you begin.

Running it in CI instead of a scheduler

A Windows scheduled task is the obvious home for a publication pipeline and not the only one. If the extraction path allows it, running the pipeline as a CI job brings several things a scheduler does not.

  • Logs someone already watches. A failed CI job is visible in a place people look; a failed scheduled task is visible in Event Viewer.
  • Configuration in version control. Changes to scope or branding get reviewed like any other change.
  • Reproducibility. The same pipeline can be run against a cached snapshot to reproduce a past publication.
  • Secrets management that already exists. Rather than a credential in a config file next to the task.

The constraint is the extraction. If it needs a modelling tool installed, you need an agent with that tool, which is a Windows machine somebody maintains — at which point a scheduled task on that machine may genuinely be simpler. The decision follows the extraction path.

What to do when it fails

Failures cluster into three kinds, and the right response differs:

FailureResponseUrgency
Validation found model errorsRoute to the architect who owns the packageThis week
Extraction could not connectCheck the repository and the credentialToday
Verification rejected the outputInvestigate before publishing anythingToday, and do not force it

The third is the one where the temptation to override is strongest, because the portal is now a week stale and someone is asking. Resist it: verification rejected the output for a reason, and publishing a portal that failed its own sanity check is how a portal loses its audience permanently.