Versioned archive. You are viewing v0.38.0. For the latest released guidance, use Latest release. Report archive issues on GitHub.
Use this runbook when you operate Base Registry Engine (BReg), Registry Casework, and Evidence together and need a recovery plan that covers all three. Each product restores differently: BReg restores a database and then proves it is the one instance serving the registry, Casework restores a database and then catches up from the BReg sources it reads, and Evidence holds no database at all. The procedures below link each product’s own guide for the details and state what a restore does not bring back. Continuous integration restores both BReg and Casework from a logical backup and checks these procedures end to end.
Know what a database backup leaves out
Section titled “Know what a database backup leaves out”A PostgreSQL backup holds rows. It does not hold the audit files, the secrets, or the key material each runtime reads at startup, and several of those cannot be recreated once lost. Back up these alongside every database backup, on the same schedule:
Base Registry Engine state
Section titled “Base Registry Engine state”- The PostgreSQL database, with
pg_dumpbefore every activation, as Back up and restore the database describes. - The package directory, the runtime file, and the secret files it references. A restore needs the package the backup was taken under.
- The field-encryption key when the project encrypts a field: the Transit key, or the local data-key file. The database holds only a wrapped key or an identifier, so losing the custodian’s key loses every sealed value, in every backup.
- The audit hash key named by
audit.hashKeyRef. A restore under another key writes pseudonyms that no longer match the ones already in the archive. - The attachment bucket, when
attachmentStoragebinds an S3-compatible backend. The database holds references; the bucket holds the bytes, so back it up at the same point.
Registry Casework state
Section titled “Registry Casework state”- The PostgreSQL database, with
pg_dumpbefore everycaseworkctl apply; name the dump with--backupso the activation ledger records it. - The installed package directory, the runtime file, and its secret files: the audit key
(
audit.hashKeyRef), the task-authority signing key (taskAuthority.signingKeyRef), the database URLs, and each source’s client credentials.
Evidence state
Section titled “Evidence state”- The governed bundle, the runtime file, and the public JWK files.
- The Transit signing key, which Transit keeps. Escrow it through your Transit backup; a local JWK signer is for local assurance only.
- The audit master and subject-binding secrets, including every previous audit master whose pseudonyms must stay recomputable.
Audit files
Section titled “Audit files”No database backup holds a product’s audit files. Ship them to append-only storage as the retention guides describe, so the archive, not the host, is the audit record: BReg, Casework, and Evidence. A restore never replaces the live audit directory.
Restore Base Registry Engine
Section titled “Restore Base Registry Engine”Stop every breg process that serves the registry before you restore. How you restore then
depends on how the copy was taken, because BReg tells a logical copy from its original but cannot
tell a physical one.
After a logical restore
Section titled “After a logical restore”A copy restored from pg_dump carries the instance claim of the database it was dumped from, so
it refuses to serve: breg exits at startup with the Registry database is not the instance its claim names, and bregctl doctor reports startup.instance_claim.mismatch. A runtime that is
already serving checks the claim on every readiness probe as well, and answers GET /ready with
503 once the claim no longer names its database.
Point the restored runtime file at the new database and change nothing else in it. Keep
identity.instanceId in particular: the webhook deliveries pending at the backup point were
captured under the event source it derives, so while any of them remains, bregctl doctor refuses
a changed value with startup.instance_id.pending_deliveries and breg refuses to start. Change
it, if you must, only after those deliveries drain.
Follow
Back up and restore the database through
verify and instance-claim status, then retire the original for good and adopt the copy:
bregctl --format json instance-claim adopt \ --runtime-config /etc/breg/runtime.yaml \ --acknowledge-original-retiredPass --acknowledge-original-retired only once no client can reach the original. Two databases
serving one registry both accept writes, and nothing can merge their histories afterwards. Until
the adopt, bregctl import-authority list shows every authority the backup held open as open,
including one you closed after the backup point. Adopting supersedes every import authority the
copy held open, so open again the ones you still need.
After a physical restore
Section titled “After a physical restore”A base backup, a point-in-time recovery, a volume snapshot, or a promoted replica keeps the original’s identity, so nothing refuses it. Two things are then your job before it serves. Keep the original stopped or fenced from clients. And end the import authorities that come back open: an authority you closed after the backup point is open again in the copy, and would admit an import you already ended. Adopt the copy once, as for a logical restore:
bregctl --format json instance-claim adopt \ --runtime-config /etc/breg/runtime.yaml \ --acknowledge-original-retiredThe claim already names the copy, so the adopt claims it again: it raises the claim’s epoch and
supersedes every open import authority in one transaction, and the audit entry records the event
reclaimed with the authorities it superseded. Open again the ones you still need.
Repeat what the backup undid
Section titled “Repeat what the backup undid”A restore rewinds every row to the backup point, including the ones you erased or delivered since. Once the copy serves, work through these:
- Erasures. Erasure does not reach backups, so a restore brings back erased history and
expired Evidence uses. Repeat each
history eraserequest you ran after the backup point, and runbregctl evidence-retention erase-expired --runtime-config /etc/breg/runtime.yaml --before <now>so Evidence assertions past their 24-hour window are erased again. See What cannot be undone. - Webhooks. Deliveries still pending at the backup point are sent again once the copy serves,
under the same
Idempotency-Keythey carried the first time, so a receiver that deduplicates by that key accepts each one once. Their attempt count is rewound with the rest of the row, so the audit records the same attempt number for each of them twice. Deliveries completed before the backup point are not sent again, and delivered events are not recalled. Inspect what is pending withbregctl webhook list; see Inspect and recover deliveries. - The audit is not rewound. The runtime audit file and its companion keep every entry the
original wrote after the backup point, including the commits and deliveries the restore lost, and
the adopted copy appends to the same files. A delivery pending at the backup point is therefore
recorded twice, once by each database. The adoption entry in the companion file (operation
breg.instance_claim.adopt, eventadopted) marks where the copy’s history resumes; record the backup time and the adoption time in the change record so the lost window can be read from the archive. Never restore an older audit file over the live one. - Reviews the authority lost. When a review environment such as Casework is restored from an
older backup, BReg reports
result-unknown-to-authorityfor each review it submitted after that backup. While the review environment is still unreachable, the code isresult-lookup-uncertaininstead, whichresubmitrefuses; it changes once the restored environment answers. Recover each withbregctl review-recovery resubmitorclose.
Restore Registry Casework
Section titled “Restore Registry Casework”Casework is not the system of record for source-backed work: the BReg source is. A restored Casework database therefore catches up with BReg on its own, while the coordination state it alone holds returns to the backup point. Restore the latest dump into a database provisioned as Provision PostgreSQL describes, keep the package and the audit key the backup was taken under, and start the runtime. Its schema must match the binary, and its activation ledger must name the package the runtime loads; see Plan, apply, and serve.
Check the ledger before you start the runtime. On the freshly provisioned database,
caseworkctl status reports no active package and caseworkctl doctor refuses with
casework.doctor.check-failed, because no migration has been applied yet. Once pg_restore has
run, caseworkctl status reports the same active package and history it reported at the backup
point, caseworkctl plan reports changesPending: false with the only refusal
casework.activation.already-active and databaseIdCheck: matches, and caseworkctl doctor
passes. Pass each command the restored runtime file with --runtime-config.
What comes back and what does not:
- Source-backed work items are replayed from BReg. Each source’s reconciliation pass rediscovers every submitted change request and re-reads every work item the database holds as active, so a request applied or cancelled after the backup point settles, and a request submitted after it appears as fresh work. That fresh work item has a new identifier, not the one the lost item had, so its audit entries carry another item pseudonym.
- Everything Casework alone records is at the backup point. Claims, private drafts, attempts, item history, clocks, directory changes, absences, task grants, and every unified review (requests, tasks, decisions, and results) are as the backup left them. Later work on them is lost.
- The audit file is not rewound. The live audit file and its shipped archive still hold the entries for the work the restore lost, and the restored runtime appends after them. An item that settled in the lost window settles again after the restore, so its completion is recorded twice under the same item pseudonym. Keep these entries; never restore an older audit file over them.
Reconcile the gap before staff resume work:
- Read the audit archive for the lost window. Entries between the backup time and the restore time record each lost operation: its event, outcome, and item revision, with the work item, team, queue, and principal as keyed pseudonyms and no payload values. No command turns a pseudonym back into a work item or a person, so use the archive to count and order what was lost, ask the staff who worked in that window to redo their claims, decisions, and directory changes from the restored state, and check the redone work against those counts. Record the backup time and the restore time in the change record.
- Settle attempts the restore left pending. A pending attempt whose lease has expired may or may not have reached BReg. Check the change request in BReg, then mark it uncertain and settle it as applied or not applied.
- Check restored review tasks against BReg before anyone decides them. A review decided after
the backup point comes back as the backup left it, open or held by the reviewer who had claimed
it, while BReg may already hold its result and may already have applied the request. No command
reconciles the two, so read the BReg request’s
data.request.reviewprojection first. Reviews BReg submitted after the backup point are unknown to the restored Casework; recover them from the BReg side as Repeat what the backup undid describes.
A database restored under a different package refuses to start until caseworkctl apply
activates the configured package, and that apply refuses a package that would strand pinned work;
see
Refuse a package that strands pinned work.
Hand over the Casework audit key
Section titled “Hand over the Casework audit key”The Casework audit key only pseudonymizes identifiers. Entries are not chained or signed, so a rotation has no chain to bridge and nothing to re-sign. Its pseudonyms carry no key version, though, so entries written under the old key and the new one look alike in one file. Hand the key over this way:
- Write the new key into a fresh owner-only secret file, as
Keep the audit file shows, and point
audit.hashKeyRefat it. - Give
audit.patha fresh file name in the same change, so the file boundary is the key boundary. The writer no longer manages the old path, so ship its active file and every sealed file once the runtime has stopped. - Restart the runtime and run
caseworkctl doctor --runtime-config /etc/registry-casework/runtime.yaml. Record the rotation time, both file names, and who holds each key in the change record. - Keep the old key under the same controls for as long as you may need to recompute pseudonyms in the entries written under it.
Continuity is a property of the archive, not of the files: prove it by showing that the shipped files cover the old path through its last sealed and final active file, the new path from its first entry, and no gap in time between the two.
Restore Evidence
Section titled “Restore Evidence”Evidence keeps no application database, so a restore is a redeploy. Install the same release, bundle, runtime file, public JWK files, and secrets, then check them before serving:
evidence check --runtime-config /etc/evidence/runtime.yaml --require-runtime-dependenciesRate limits and caches start empty. Restore audit history into your archive, never into the live
audit directory. When evidence-oid4vci delivers credentials, start it after Evidence is ready:
its offers live in memory, so a restart invalidates every outstanding offer and wallets must ask
for a fresh one. See
Operational limits.
Run a disaster-recovery drill
Section titled “Run a disaster-recovery drill”A drill proves the backups restore, the key material is recoverable, and your team can do it within the recovery time you promise. Run it at least before going live and after every change to the backup tooling.
An adopted BReg copy delivers the webhooks pending in its backup, a restored Casework reads and acts on the sources it names, and Evidence reads the sources in its bundle. Before starting anything, point every event destination, Casework source, Evidence source, and OpenID Connect client in the drill’s runtime files at non-production systems, and keep the drill network unable to reach production.
- Restore the latest BReg and Casework backups onto drill hosts, and redeploy Evidence, using only what your backup store and secret escrow hold. A secret you have to fetch from a production host is a gap in the backup set.
- Run
bregctl verifyandbregctl instance-claim status, and confirm the status reports that the claim names another database, then adopt the copy. Passing--acknowledge-original-retiredis sound here only because the drill network cannot reach production or its clients. - Run
bregctl doctor,caseworkctl doctor, andevidence check --require-runtime-dependencies, start the runtimes, and wait forGET /readyon each. - Read a known record through BReg, a known work item through Casework, and a synthetic assertion through Evidence. Read one encrypted field, when the project encrypts one, to prove the field-encryption key restored.
- Open the shipped audit archive for the period before the backup, and confirm every file is there and ends with a complete entry.
- Record the backup time, the restore time, the measured recovery time, and every step that failed, and destroy the drill hosts and their copies of the secrets.