Skip to content

DevOps

A PostgreSQL backup is only useful after a restore drill

How I test a Rechvix backup in a scratch database, check business invariants and document the recovery procedure instead of trusting a successful archive command.

By · Published · 6 min read

For Rechvix I wrote a backup runbook that does not end when pg_dump exits successfully. The useful question is whether an operator can rebuild a working system from the archive and whether the restored accounting and inventory state still makes sense.

Restore somewhere disposable

The drill uses a scratch PostgreSQL instance, separate from production. I restore the archive there, apply the documented configuration and verify that the application can read the recovered data. PostgreSQL's pg_dump creates the archive; pg_restore loads a custom-format archive into the target. Their exit status is a starting signal, not the final proof.

Check invariants, not only row counts

I compare row counts for the tables that matter and check the trial balance. If a financial system reports a balanced ledger before the backup and an unbalanced one afterward, the restore is not acceptable even if the SQL completed. Stock balances are a projection of append-only movements, so their agreement is another useful check.

A runbook must include the people and secrets

A database archive alone does not tell a new operator which role may run migrations, which role the application uses, where encryption keys come from or which external integration credentials must be restored. Rechvix deliberately separates its schema-owning migrator from a non-owning runtime role. The restore procedure has to recreate that boundary.

Writing the runbook exposed a gap: API keys had no scope for backup operations. That was worth recording. A restore drill is useful partly because it makes omissions in the surrounding system visible before an outage.

Make the next drill repeatable

The output I want is a small record: archive date, PostgreSQL version, restore target, commands used, checks performed, result and any manual steps. If a step depends on one person's memory, it belongs in the runbook. If a check can be automated, it should eventually run without them.

References

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello