A corruption class the repair strategies cannot heal (b-tree page
damage) failed repair_state_db_schema on every process start, forever:
_claim_repair_attempt's in-memory set only bounds one process, so each
restart re-ran the full surgery AND took a fresh ~900MB forensic backup
of the same damaged bytes — 105 attempts / 89GB of dead
state.db.malformed-backup-* files over 11 days in the reporting install.
Three bounded behaviors, all sidecar-file based (no schema changes):
1. Persistent attempt ledger (<db>.repair-attempts.json): after 3 failed
repair passes against the same file fingerprint (size + mtime_ns),
repair_state_db_schema refuses with a terminal, actionable error
(restore a backup / `sqlite3 state.db ".recover"` / delete the ledger
to force a retry) instead of re-running surgery. Success clears the
ledger; a replaced or restored file re-keys it and gets fresh
attempts. Missing/corrupt ledger fails open (never blocks a first
repair).
2. Backup dedupe: _backup_db_file reuses the newest existing forensic
backup when it is byte-identical to the damaged file (size+mtime
match, preserved by copy2) instead of copying another ~900MB.
3. Retention cap: only the 3 newest malformed-backup copies (plus
sidecars) are kept; older ones are pruned after each new backup.
Also fixes a same-second timestamp collision that silently
overwrote an earlier forensic copy.
Tests cover ledger accumulation, terminal refusal (surgery not called,
no new backup), budget reset on file change, success-clears-ledger,
corrupt-ledger tolerance, dedupe, distinct-state backups, retention
prune incl. sidecars, and the end-to-end one-backup invariant.
Fixes#86747