Insights / Technical notes

Monitoring restic backups on R2: from successful storage to verified restoration

Track snapshot freshness, repository integrity and restoration separately, and identify the application recovery steps that remain untested.

  • Cloudflare R2
  • restic
  • Backup
Monitoring restic backups on R2: from successful storage to verified restoration
Table of contents
  1. Fix the snapshot ID and one recovery goal
  2. Define the source and success
  3. Monitor freshness separately
  4. Check integrity and extraction
  5. Retain data through restic
  6. Remaining recovery checks
  7. October 6, 2026 update: stale locks and exclusive restore checks

A completed backup job does not establish that the required data can be recovered. This generalized internal case evaluates storage, integrity, restoration and service recovery separately. It does not describe a completed migration from another backup service.

Fix the snapshot ID and one recovery goal

Record the snapshot ID and restore needed files to an empty isolated destination. For settings, check references and permissions; for a database, test loading it in isolation. Record elapsed time. Repeat the same checks periodically to compare freshness and usable recovery scope.

restic:Restoring to an isolated destination

Define the source and success

The design uses restic encryption and deduplication with R2’s S3-compatible API. Check required operations against the compatibility table. For live databases, design consistent capture using dumps or application quiescence where appropriate.

Monitor freshness separately

Track the latest successful snapshot, overdue runs, failures, and integrity/restore results independently. Starting a job is not success. Also distinguish a failed notification from a failed backup.

Check integrity and extraction

Default restic check and checks that read stored data have different scopes. --read-data reads all data; record the scope of any partial check. Combine the repository checks with restoration into an isolated location and comparisons of contents or hashes. Extracting files does not verify application startup.

Retain data through restic

Choose snapshots with restic’s retention policy, then use forget, prune and check, inspecting what will remain before executing. Blanket age-based deletion in R2 can remove shared data needed by retained snapshots. Follow the retention documentation; object cleanup for image delivery is a different use case.

Remaining recovery checks

Scheduled operation, monitoring/notifications, retention and selected data extraction/integrity checks were performed. Startup of every application, recovery of configuration and dependencies, and retrieving separately held credentials still need end-to-end validation. Recovery time and acceptable data loss also require measurement. This case does not claim complete disaster recovery or proven cost savings.

Check evidence in stages from snapshot freshness to full recovery Retrieving data does not prove that applications or credentials can be recovered.
  1. Successful snapshot and freshness Check the timestamp of an actually successful snapshot and its delay from schedule. A job starting alone is not success.
  2. Integrity and isolated restore Record the check scope, run restore --verify in a separate location, and compare target files and hashes before reviewing retention and prune.
  3. Full disaster recovery Application startup, dependent data, and recovery of separately stored credentials remain unverified. Recovery time and acceptable data loss are unmeasured.

October 6, 2026 update: stale locks and exclusive restore checks

Further changes use ordinary restic unlock after obtaining exclusive operation access on the same host, handling stale locks only. They do not use --remove-all to remove active locks. A bounded --retry-lock handles contention, and restore checks and prune do not restart indefinitely on failure. Host-local exclusivity alone does not prove absence of competing hosts.

Restore with restore --verify into an isolated temporary directory, check required target files and consistency, then record the verified snapshot and configuration together. Verify restoration before reviewing retained snapshots and pruning. Judge backup freshness from a successfully created snapshot, not job startup or retries.

Target-data extraction remains distinct from full recovery including application startup and credentials. See monitoring and investigation for maintenance windows and acquisition failures.