Commit Graph

1 Commits

Author SHA1 Message Date
Teknium 0d00ebef79 fix(state): bound the state.db repair loop and stop dead-backup accumulation (#86747)
A corruption class the repair strategies cannot heal (b-tree page
damage) failed repair_state_db_schema on every process start, forever:
_claim_repair_attempt's in-memory set only bounds one process, so each
restart re-ran the full surgery AND took a fresh ~900MB forensic backup
of the same damaged bytes — 105 attempts / 89GB of dead
state.db.malformed-backup-* files over 11 days in the reporting install.

Three bounded behaviors, all sidecar-file based (no schema changes):

1. Persistent attempt ledger (<db>.repair-attempts.json): after 3 failed
   repair passes against the same file fingerprint (size + mtime_ns),
   repair_state_db_schema refuses with a terminal, actionable error
   (restore a backup / `sqlite3 state.db ".recover"` / delete the ledger
   to force a retry) instead of re-running surgery. Success clears the
   ledger; a replaced or restored file re-keys it and gets fresh
   attempts. Missing/corrupt ledger fails open (never blocks a first
   repair).

2. Backup dedupe: _backup_db_file reuses the newest existing forensic
   backup when it is byte-identical to the damaged file (size+mtime
   match, preserved by copy2) instead of copying another ~900MB.

3. Retention cap: only the 3 newest malformed-backup copies (plus
   sidecars) are kept; older ones are pruned after each new backup.
   Also fixes a same-second timestamp collision that silently
   overwrote an earlier forensic copy.

Tests cover ledger accumulation, terminal refusal (surgery not called,
no new backup), budget reset on file change, success-clears-ledger,
corrupt-ledger tolerance, dedupe, distinct-state backups, retention
prune incl. sidecars, and the end-to-end one-backup invariant.

Fixes #86747
2026-08-15 02:22:58 -07:00