Skip to content

Infrastructure

Recovering from self-hosted data loss: what to do in the first hour

A calm recovery sequence when a self-hosted app loses data: stop writing, preserve what is left, find the real backup, restore to a side copy, verify, and fix the cause.

By · Published · 3 min read

Short answer: stop anything that writes to the affected disk or database, make a copy of whatever remains, then look for backups in every place they might exist before you try a clever repair. Restore into a separate location first, check the result against what you know should be there, and only then switch over. Panic causes the second loss, which is usually worse than the first.

Step 1: Stop making it worse

  • Stop the application and any job that writes: cron tasks, sync clients, backup scripts that might overwrite good backups with empty ones.
  • Do not run docker system prune, reinstall, or "just rebuild it". Anything that recreates volumes can destroy what is left.
  • If a disk is failing, power it down. Every minute of use can cost recoverable data.
  • Write down what happened, in order, with times. You will forget details later.

Step 2: Preserve what remains

Before you change anything, copy. If it is a volume or directory, archive it somewhere else. If a disk may be failing, image it with ddrescue onto a healthy disk and work on the image.

# copy a docker volume to an archive without touching it
docker run --rm -v app_db_data:/data:ro -v "$PWD":/out debian \
  tar czf /out/app_db_data-$(date +%F).tgz -C /data .

Step 3: Find every possible backup

People usually have more copies than they remember.

  • Your own backup jobs: local, off-site, object storage. Check their logs. A job that has been failing for weeks is common.
  • Provider snapshots at your VPS or cloud host.
  • A database replica, standby or a second environment.
  • Application exports and scheduled reports. An invoice PDF is not a database, but it is evidence.
  • Old laptops and dev copies with a recent dump.
  • Email, chat attachments and printed documents for small, critical records.

Step 4: Restore to a side copy

Never restore on top of the damaged system. Bring up a clean instance on another machine or under another name, and load the backup into it.

createdb -h localhost restore_check
pg_restore -h localhost -d restore_check --no-owner latest.dump
psql -h localhost -d restore_check -c "SELECT count(*) FROM invoices;"

Step 5: Verify before you switch

  • Row counts for the main tables.
  • The latest record dates. This tells you how much was lost between the backup and the incident.
  • Business checks: does the ledger balance, do stock totals match the last stock count? The point is in my restore drill post.
  • Whether the application starts and can log in against the restored data, including encrypted fields. Lost encryption keys can make a perfect backup unreadable.

Step 6: Close the gap

Whatever happened between the backup time and the loss has to come from somewhere: paper notes, bank statements, payment provider records, emails, customers. Make a list, reconcile and re-enter. Tell the people affected early and plainly.

Step 7: Find the cause and fix it

Data loss usually comes from a small set of causes.

  • An uninstall, redeploy or down -v that removed a volume. Some app stores delete app data on uninstall.
  • Bind mounts or volumes that were not where you thought.
  • A disk that filled up and corrupted a database.
  • A failing disk with no redundancy.
  • A backup that ran but could not be restored, or one that backed up the wrong path.
  • An operator command run on the wrong server.

Write a short note: what happened, what was lost, how it was recovered, what changes now. Then change something, not just promise to be careful.

What to change afterwards

  • Automated backups with alerting on failure and on suspiciously small output.
  • At least one off-site, versioned or immutable copy that a mistaken command cannot overwrite.
  • A scheduled restore test.
  • Names and prompts that make it hard to run destructive commands on production.
  • Documentation of where data actually lives. See the self-hosting checklist.

The best recovery is the boring one: you have a backup, you have tested it, you restore it, and the day continues.

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello