Skip to content
Registry StackDocsv0.38.0

Back up, restore, and drill recovery

View as Markdown

Use this runbook when you operate Base Registry Engine (BReg), Registry Casework, and Evidence together and need a recovery plan that covers all three. Each product restores differently: BReg restores a database and then proves it is the one instance serving the registry, Casework restores a database and then catches up from the BReg sources it reads, and Evidence holds no database at all. The procedures below link each product’s own guide for the details and state what a restore does not bring back. Continuous integration restores both BReg and Casework from a logical backup and checks these procedures end to end.

A PostgreSQL backup holds rows. It does not hold the audit files, the secrets, or the key material each runtime reads at startup, and several of those cannot be recreated once lost. Back up these alongside every database backup, on the same schedule:

  • The PostgreSQL database, with pg_dump before every activation, as Back up and restore the database describes.
  • The package directory, the runtime file, and the secret files it references. A restore needs the package the backup was taken under.
  • The field-encryption key when the project encrypts a field: the Transit key, or the local data-key file. The database holds only a wrapped key or an identifier, so losing the custodian’s key loses every sealed value, in every backup.
  • The audit hash key named by audit.hashKeyRef. A restore under another key writes pseudonyms that no longer match the ones already in the archive.
  • The attachment bucket, when attachmentStorage binds an S3-compatible backend. The database holds references; the bucket holds the bytes, so back it up at the same point.
  • The PostgreSQL database, with pg_dump before every caseworkctl apply; name the dump with --backup so the activation ledger records it.
  • The installed package directory, the runtime file, and its secret files: the audit key (audit.hashKeyRef), the task-authority signing key (taskAuthority.signingKeyRef), the database URLs, and each source’s client credentials.
  • The governed bundle, the runtime file, and the public JWK files.
  • The Transit signing key, which Transit keeps. Escrow it through your Transit backup; a local JWK signer is for local assurance only.
  • The audit master and subject-binding secrets, including every previous audit master whose pseudonyms must stay recomputable.

No database backup holds a product’s audit files. Ship them to append-only storage as the retention guides describe, so the archive, not the host, is the audit record: BReg, Casework, and Evidence. A restore never replaces the live audit directory.

Stop every breg process that serves the registry before you restore. How you restore then depends on how the copy was taken, because BReg tells a logical copy from its original but cannot tell a physical one.

A copy restored from pg_dump carries the instance claim of the database it was dumped from, so it refuses to serve: breg exits at startup with the Registry database is not the instance its claim names, and bregctl doctor reports startup.instance_claim.mismatch. A runtime that is already serving checks the claim on every readiness probe as well, and answers GET /ready with 503 once the claim no longer names its database.

Point the restored runtime file at the new database and change nothing else in it. Keep identity.instanceId in particular: the webhook deliveries pending at the backup point were captured under the event source it derives, so while any of them remains, bregctl doctor refuses a changed value with startup.instance_id.pending_deliveries and breg refuses to start. Change it, if you must, only after those deliveries drain.

Follow Back up and restore the database through verify and instance-claim status, then retire the original for good and adopt the copy:

Terminal window
bregctl --format json instance-claim adopt \
--runtime-config /etc/breg/runtime.yaml \
--acknowledge-original-retired

Pass --acknowledge-original-retired only once no client can reach the original. Two databases serving one registry both accept writes, and nothing can merge their histories afterwards. Until the adopt, bregctl import-authority list shows every authority the backup held open as open, including one you closed after the backup point. Adopting supersedes every import authority the copy held open, so open again the ones you still need.

A base backup, a point-in-time recovery, a volume snapshot, or a promoted replica keeps the original’s identity, so nothing refuses it. Two things are then your job before it serves. Keep the original stopped or fenced from clients. And end the import authorities that come back open: an authority you closed after the backup point is open again in the copy, and would admit an import you already ended. Adopt the copy once, as for a logical restore:

Terminal window
bregctl --format json instance-claim adopt \
--runtime-config /etc/breg/runtime.yaml \
--acknowledge-original-retired

The claim already names the copy, so the adopt claims it again: it raises the claim’s epoch and supersedes every open import authority in one transaction, and the audit entry records the event reclaimed with the authorities it superseded. Open again the ones you still need.

A restore rewinds every row to the backup point, including the ones you erased or delivered since. Once the copy serves, work through these:

  • Erasures. Erasure does not reach backups, so a restore brings back erased history and expired Evidence uses. Repeat each history erase request you ran after the backup point, and run bregctl evidence-retention erase-expired --runtime-config /etc/breg/runtime.yaml --before <now> so Evidence assertions past their 24-hour window are erased again. See What cannot be undone.
  • Webhooks. Deliveries still pending at the backup point are sent again once the copy serves, under the same Idempotency-Key they carried the first time, so a receiver that deduplicates by that key accepts each one once. Their attempt count is rewound with the rest of the row, so the audit records the same attempt number for each of them twice. Deliveries completed before the backup point are not sent again, and delivered events are not recalled. Inspect what is pending with bregctl webhook list; see Inspect and recover deliveries.
  • The audit is not rewound. The runtime audit file and its companion keep every entry the original wrote after the backup point, including the commits and deliveries the restore lost, and the adopted copy appends to the same files. A delivery pending at the backup point is therefore recorded twice, once by each database. The adoption entry in the companion file (operation breg.instance_claim.adopt, event adopted) marks where the copy’s history resumes; record the backup time and the adoption time in the change record so the lost window can be read from the archive. Never restore an older audit file over the live one.
  • Reviews the authority lost. When a review environment such as Casework is restored from an older backup, BReg reports result-unknown-to-authority for each review it submitted after that backup. While the review environment is still unreachable, the code is result-lookup-uncertain instead, which resubmit refuses; it changes once the restored environment answers. Recover each with bregctl review-recovery resubmit or close.

Casework is not the system of record for source-backed work: the BReg source is. A restored Casework database therefore catches up with BReg on its own, while the coordination state it alone holds returns to the backup point. Restore the latest dump into a database provisioned as Provision PostgreSQL describes, keep the package and the audit key the backup was taken under, and start the runtime. Its schema must match the binary, and its activation ledger must name the package the runtime loads; see Plan, apply, and serve.

Check the ledger before you start the runtime. On the freshly provisioned database, caseworkctl status reports no active package and caseworkctl doctor refuses with casework.doctor.check-failed, because no migration has been applied yet. Once pg_restore has run, caseworkctl status reports the same active package and history it reported at the backup point, caseworkctl plan reports changesPending: false with the only refusal casework.activation.already-active and databaseIdCheck: matches, and caseworkctl doctor passes. Pass each command the restored runtime file with --runtime-config.

What comes back and what does not:

  • Source-backed work items are replayed from BReg. Each source’s reconciliation pass rediscovers every submitted change request and re-reads every work item the database holds as active, so a request applied or cancelled after the backup point settles, and a request submitted after it appears as fresh work. That fresh work item has a new identifier, not the one the lost item had, so its audit entries carry another item pseudonym.
  • Everything Casework alone records is at the backup point. Claims, private drafts, attempts, item history, clocks, directory changes, absences, task grants, and every unified review (requests, tasks, decisions, and results) are as the backup left them. Later work on them is lost.
  • The audit file is not rewound. The live audit file and its shipped archive still hold the entries for the work the restore lost, and the restored runtime appends after them. An item that settled in the lost window settles again after the restore, so its completion is recorded twice under the same item pseudonym. Keep these entries; never restore an older audit file over them.

Reconcile the gap before staff resume work:

  1. Read the audit archive for the lost window. Entries between the backup time and the restore time record each lost operation: its event, outcome, and item revision, with the work item, team, queue, and principal as keyed pseudonyms and no payload values. No command turns a pseudonym back into a work item or a person, so use the archive to count and order what was lost, ask the staff who worked in that window to redo their claims, decisions, and directory changes from the restored state, and check the redone work against those counts. Record the backup time and the restore time in the change record.
  2. Settle attempts the restore left pending. A pending attempt whose lease has expired may or may not have reached BReg. Check the change request in BReg, then mark it uncertain and settle it as applied or not applied.
  3. Check restored review tasks against BReg before anyone decides them. A review decided after the backup point comes back as the backup left it, open or held by the reviewer who had claimed it, while BReg may already hold its result and may already have applied the request. No command reconciles the two, so read the BReg request’s data.request.review projection first. Reviews BReg submitted after the backup point are unknown to the restored Casework; recover them from the BReg side as Repeat what the backup undid describes.

A database restored under a different package refuses to start until caseworkctl apply activates the configured package, and that apply refuses a package that would strand pinned work; see Refuse a package that strands pinned work.

The Casework audit key only pseudonymizes identifiers. Entries are not chained or signed, so a rotation has no chain to bridge and nothing to re-sign. Its pseudonyms carry no key version, though, so entries written under the old key and the new one look alike in one file. Hand the key over this way:

  1. Write the new key into a fresh owner-only secret file, as Keep the audit file shows, and point audit.hashKeyRef at it.
  2. Give audit.path a fresh file name in the same change, so the file boundary is the key boundary. The writer no longer manages the old path, so ship its active file and every sealed file once the runtime has stopped.
  3. Restart the runtime and run caseworkctl doctor --runtime-config /etc/registry-casework/runtime.yaml. Record the rotation time, both file names, and who holds each key in the change record.
  4. Keep the old key under the same controls for as long as you may need to recompute pseudonyms in the entries written under it.

Continuity is a property of the archive, not of the files: prove it by showing that the shipped files cover the old path through its last sealed and final active file, the new path from its first entry, and no gap in time between the two.

Evidence keeps no application database, so a restore is a redeploy. Install the same release, bundle, runtime file, public JWK files, and secrets, then check them before serving:

Terminal window
evidence check --runtime-config /etc/evidence/runtime.yaml --require-runtime-dependencies

Rate limits and caches start empty. Restore audit history into your archive, never into the live audit directory. When evidence-oid4vci delivers credentials, start it after Evidence is ready: its offers live in memory, so a restart invalidates every outstanding offer and wallets must ask for a fresh one. See Operational limits.

A drill proves the backups restore, the key material is recoverable, and your team can do it within the recovery time you promise. Run it at least before going live and after every change to the backup tooling.

  1. Restore the latest BReg and Casework backups onto drill hosts, and redeploy Evidence, using only what your backup store and secret escrow hold. A secret you have to fetch from a production host is a gap in the backup set.
  2. Run bregctl verify and bregctl instance-claim status, and confirm the status reports that the claim names another database, then adopt the copy. Passing --acknowledge-original-retired is sound here only because the drill network cannot reach production or its clients.
  3. Run bregctl doctor, caseworkctl doctor, and evidence check --require-runtime-dependencies, start the runtimes, and wait for GET /ready on each.
  4. Read a known record through BReg, a known work item through Casework, and a synthetic assertion through Evidence. Read one encrypted field, when the project encrypts one, to prove the field-encryption key restored.
  5. Open the shipped audit archive for the period before the backup, and confirm every file is there and ends with a complete entry.
  6. Record the backup time, the restore time, the measured recovery time, and every step that failed, and destroy the drill hosts and their copies of the secrets.