Two things need backing up: the Orchestrator state root (queues, vault, principals and journals), plus the queue database itself when the queue runs on SQLite. Each has its own tool. Each fails closed rather than producing a backup that is quietly wrong.
#Backing up the state root
ops/state-backup.mjs copies the whole state root.
node ops/state-backup.mjs backup --root orchestrator/state --target /backups
node ops/state-backup.mjs backup --root orchestrator/state --target /backups --exclude-key-material
The backup writes a stamped directory, tasksultan-state-<UTC timestamp>, into the target. It verifies the copy before reporting success. It fails closed on any unreadable or corrupt JSON in the source tree. The destination path is printed on stdout as the machine-readable result; notes go to stderr.
The key-material guard is the part to understand. The vault's wrapping key is what makes an encrypted vault readable. Putting the key next to the ciphertext it unwraps makes every backup plaintext at rest. So a backup refuses when the root holds key material, naming the files. --exclude-key-material copies everything else and leaves the key behind. Detection is by name and by content: a small file whose entire content is 32 bytes of key material is flagged whatever it is called. A backup tree contains principals and vault material, so treat it like a secret. Retention is yours.
#Restoring the state root
node ops/state-backup.mjs restore /backups/tasksultan-state-2026-10-03T09-15-00-000Z --dry-run
node ops/state-backup.mjs restore /backups/tasksultan-state-2026-10-03T09-15-00-000Z --force
node ops/state-backup.mjs restore /backups/tasksultan-state-2026-10-03T09-15-00-000Z --force --allow-key-material
A restore is a dry run unless you pass --force, so you see what would be written before anything is. Point it at the stamp directory a backup printed, not at the target directory that contains it: a target dir is refused by name, with the stamp to use named in the message.
Symmetry with backup is deliberate. A restore refuses to write key material back into the state root, because that re-creates the exposure the backup guard removes and makes the operator's next backup fail closed. --allow-key-material exists only for a deliberate legacy recovery. It warns you to move the key out afterwards.
A restore proves bytes. It says so. The success line reads "byte-verified"; the notes state what was not established. If the restored vault is encrypted and the backup carried no key, the restore says the vault will not open until the wrapping key is supplied. Verify a restored root before trusting it:
npx tsx orchestrator/cli.ts vault doctor --root orchestrator/state
#Backing up a SQLite queue
A live SQLite database is not a file you copy. Copying it while a writer holds it can capture a torn state. It also ignores the -wal file with committed frames that are not yet checkpointed. The supported way is a consistent snapshot taken by the database itself.
import { snapshotQueueDatabase, inspectQueueDatabase } from './orchestrator/sqlite-queue-backup';
snapshotQueueDatabase({ file: 'state/queue.db', destFile: 'state/queue-snapshot.db' });
const check = inspectQueueDatabase('state/queue-snapshot.db');
snapshotQueueDatabase uses VACUUM INTO, which writes a fully formed, transactionally consistent database while writers are running. It refuses to overwrite an existing target. The placement guard applies to the destination as well as the source, so a snapshot is not written into a synced folder where the same corruption hazard applies. inspectQueueDatabase runs SQLite's own integrity_check and reports the row counts of queue_items and transactions, so a snapshot can be shown to restore to the same state rather than merely to exist.
#Recovering interrupted work
A generation is all or nothing. If a generate run is interrupted, its transaction journal survives. The next start rolls it back and restores the files it touched. Interrupted generations are recovered automatically at the start of a generation run. Run-state collection (see state gc) is inspection first, refuses --force until the retention policy is ratified. It journals every removal. Failed runs are keeper-by-default.