Skip to content
Registry StackDocsv0.39.0

Upgrade and retire a deployment

View as Markdown

Use this runbook when you move a deployment of Base Registry Engine (BReg), Registry Casework, Evidence, and Registry Scheduling to a new release, or take it out of service. Each product’s own guide carries its upgrade command; this page gives the order across products, the version rule that order enforces, and what each product must leave behind when it is retired.

Every Registry Stack artifact of a release shares one version: the runtimes, their adopter tooling, and the Rust, Node.js, and Python clients. Run the same release on every side. There is no supported version-skew window, between two runtimes or between a client and a runtime, and a rolling upgrade that mixes releases is not supported:

  • Casework and BReg are checked. Casework compares BReg’s Registry-Engine-Version header with its own release on every contract read and refuses any other release; see Upgrade Casework and BReg in lock-step.
  • Adopter tooling is checked. evidencectl refuses an evidence or bregctl binary of another version.
  • Clients are not checked, and break silently. Clients send no version, and runtimes accept any caller. The clients decode responses strictly, so a member one release adds can fail an older client’s decoding, and a member one release requires can be missing from an older runtime’s answer. Upgrade each client to the same release as the runtime it calls.
  • BReg activation is not rolling. Once bregctl apply activates a successor, every breg process still running the previous package fails readiness until it restarts onto the successor. Plan the restart as part of the activation, not after it.

Before 1.0, a release reads only the state its immediate predecessor wrote, and there is no reverse path. If you run an older release, upgrade one release at a time, and finish this runbook for each release before you start the next. See Compatibility direction.

BReg goes first, because Casework and Evidence both read it. Casework follows because it refuses a BReg of another release. Evidence and its wallet-facing front end come next. Scheduling reads none of them and upgrades on its own, before the clients. Each release’s notes list, per product, the project and runtime-file changes its upgrade needs; make them in the step that upgrades that product. Rehearse the whole sequence on a restored copy first, as Run a disaster-recovery drill describes.

  1. Back up every database. Take the BReg and Casework pg_dumps, and the Scheduling one when it runs, and copy those products’ packages and runtime files. These backups are the only way back; see Roll back.
  2. Stop evidence-oid4vci, when it runs. Its outstanding offers end here; see Operational limits.
  3. Pause and drain Evidence traffic for every question that reads BReg, or pause the whole instance when you cannot drain one question at a time. Keep the previous Evidence candidate.
  4. Upgrade BReg. Make the project and runtime-file changes the release notes list, rebuild the deployed project with no model change using the new bregctl test and bregctl package --test-receipt FILE --output BUILD (a receipt from the earlier release is not accepted), point package.root at BUILD/package (and package.expectedDigest, when set, at the package digest it reports, packageDigest with --format json), check it against the live database with bregctl plan --runtime-config FILE --package BUILD/package naming the same directory, apply it with bregctl apply --runtime-config FILE --package BUILD/package, run bregctl verify --runtime-config FILE, and restart every breg process on the new binary; bregctl status --runtime-config FILE confirms what the database activated. A model change follows as its own successor; see Activate the successor. When an apply stops before it finishes, rerun it with the same database.roles: a retry under other roles is refused as apply.resume.roles_differ, naming the roles the activation started with. A new package applied under database.roles other than the ones the active activation serves with is refused as apply.successor.roles_differ before maintenance; apply the active package under the new roles first, then the new package. A role change, the active package applied under other database.roles, cannot be assessed by bregctl migration reconcile: when one stops before it finishes, fix the cause and rerun the same apply, which resumes it. Casework reports each BReg source as unavailable from here until step 5 finishes.
  5. Upgrade Casework. Settle pending attempts first, since an upgrade can strand them. When step 4 changed the registryRevision a BReg source serves, check each such source with caseworkctl check PROJECT --against-breg-package DIR --source-id ID, repin it with caseworkctl source add BREG_PROJECT --project PROJECT --source-id ID --apply, package the Casework project again, and point the runtime file’s package.root at the new package (and package.expectedDigest, when set, at its digest); see Check a BReg source’s pinned revision. Stop every earlier casework process before the apply: one left running holds the audit writer lock the new runtime needs, and can strand a source attempt when the apply changes a source’s binding generation. Run caseworkctl plan --runtime-config FILE and caseworkctl apply --runtime-config FILE with the new binaries, then start casework and run caseworkctl doctor --runtime-config FILE; see Plan, apply, and serve.
  6. Upgrade Evidence. Re-import each BReg source whose export the release notes say changed, then package a fresh candidate with the new evidencectl package, run evidence check --runtime-config FILE --require-runtime-dependencies with the new binary (add --without-audit-lock while the earlier instance still runs), stop the earlier evidence, start the new one, and verify a fresh synthetic assertion. Then start the new evidence-oid4vci.
  7. Upgrade Scheduling, when it runs. Stop the earlier scheduling runtime, or keep its destination bindings until its hook deliveries drain. Run schedulingctl plan --runtime-config FILE and schedulingctl apply --runtime-config FILE with the new binary, then start scheduling; schedulingctl status --runtime-config FILE confirms what the database activated.
  8. Resume traffic, and upgrade clients. Roll out every application that embeds a Registry Stack client at the same release before it calls the upgraded runtimes.

No product migrates backwards. A Casework binary refuses a database whose schema is newer than it supports, and a BReg package applies forward only. Evidence refuses a runtime file carrying keys its release does not know. To return to the previous release:

  • BReg. To return to the previous release, stop every breg and retire the upgraded database, restore the step 1 pg_dump into a fresh database, and point the previous runtime file copied in step 1 at it. The restored copy carries the claim of the database it was dumped from, so adopt it with the previous release’s bregctl instance-claim adopt --runtime-config FILE --acknowledge-original-retired (see After a logical restore for why), then start the previous breg with the previous package and that runtime file; the new release’s tools cannot read the earlier package or runtime file. To undo only a model change, roll forward to a successor that reverts it, and when the data itself must go back, restore the pre-activation backup; see Roll back by rolling forward.
  • Casework. Restore the pre-apply backup and run the previous binary, package, and runtime file copied in step 1; see Restore Registry Casework for what that restore loses.
  • Evidence. Restart the previous binary with the previous bundle and runtime file.
  • Scheduling. Restore the pre-apply database backup, then run the previous binary, package, and runtime file copied in step 1. Always restore first: an earlier Scheduling runtime may not detect a schema newer than its own.

Because the release rule holds in both directions, rolling back one product means rolling back every product that must match it: Casework refuses a BReg that went back without it.

Retire the products in the reverse of the upgrade order, so nothing still running depends on what you have removed. Retiring a product leaves records behind: signatures others still verify, decisions others may challenge, and audit entries your retention policy holds. None of the products has a decommission command, so the steps below are yours.

Scheduling reads no other product and no other product reads it, so it can retire at any point in this sequence.

  1. Stop routing booking requests to Scheduling, and tell the callers holding appointments how those appointments will be honoured.
  2. Stop scheduling. Reminder and observer-hook deliveries are outbox work its own workers send, so every delivery still pending stops with it and is not sent.
  3. Scheduling has no export command. Keep a final pg_dump for as long as your policy holds its appointment records.
  4. Ship the final audit file, its sealed files, and the schedulingctl sibling of audit.path.
  1. Stop evidence-oid4vci, then stop routing requests to Evidence.
  2. Keep the published public keys available. An assertion Evidence issued stays valid for up to maximumAssertionValiditySeconds (at most one year), and the JWKS is served only by the running process, so publish the key set elsewhere or give it to each relying party, and keep it until that period and the verifier clock skew have passed since the last issuance.
  3. Ship the final audit files after the process stops, and keep every audit master whose pseudonyms an investigation may still need; see Rotate the audit master.
  4. Retire the signing key in Transit only after step 2’s period ends; see Retire the old version.
  5. Retire or rebind the Evidence provider in every BReg project whose governed actions call it, since an action that resolves against a retired provider fails.
  1. Stop staff access. Settle every pending and uncertain attempt, as Settle an uncertain source attempt describes, so no change is left half-applied at a source.
  2. Finish or close the reviews Casework hosts for BReg. A review left open makes BReg wait for a result that never comes; close it from BReg with bregctl review-recovery close.
  3. Casework has no export command. Keep a final pg_dump for as long as your policy holds its decisions and accountability records, or erase source-backed payload copies first with caseworkctl retention erase.
  4. Stop the runtime, ship the final audit file, its sealed files, and the caseworkctl companion file, then revoke Casework’s client credentials at each BReg source’s identity provider.
  1. Stop every consumer, then close the import authorities still open with bregctl import-authority close.
  2. Export what must outlive the registry. bregctl data export exports one entity through one access profile, checkpointed; a final pg_dump is the only complete copy. See Export with a checkpoint.
  3. Erase what your policy does not let you keep, before that final backup: see Erase retained history and Erase expired Evidence uses.
  4. Stop breg. Pending webhook deliveries stop with it and are not sent; check bregctl webhook list first if receivers must see them.
  5. Ship the final audit files and the bregctl companion file.
  6. Keep the field-encryption key for as long as you keep any backup that holds sealed values, and the audit hash key for as long as you keep the audit archive. Destroy each only when the last record that needs it is gone.