fix(state): use PASSIVE checkpoint for periodic WAL flush to prevent B-tree corruption

TRUNCATE checkpoint every 50 writes causes B-tree corruption on large
databases (65K+ pages) due to the exclusive-lock I/O pressure from
checkpointing thousands of frames at once.

Switch periodic _try_wal_checkpoint() to PASSIVE mode which does not
require an exclusive lock and cannot corrupt pages under I/O pressure.
Keep TRUNCATE in close() and pre-VACUUM paths where it is safe
(infrequent, controlled conditions).

Also replace silent `except Exception: pass` with logged warnings for
checkpoint failures so operators can detect early corruption signals.

Fixes #45383
This commit is contained in:
liuhao1024
2026-06-13 12:10:05 +08:00
committed by Teknium
parent 779c0dd80d
commit c2a3b9ce58
2 changed files with 158 additions and 21 deletions
+20 -21
View File
@@ -986,7 +986,7 @@ class SessionDB:
_WRITE_MAX_RETRIES = 15
_WRITE_RETRY_MIN_S = 0.020 # 20ms
_WRITE_RETRY_MAX_S = 0.150 # 150ms
# Attempt a PASSIVE WAL checkpoint every N successful writes.
# Attempt a WAL checkpoint every N successful writes (PASSIVE mode).
_CHECKPOINT_EVERY_N_WRITES = 50
# Merge fragmented FTS5 segments every N successful writes. The message
# triggers append one segment per insert; left unmaintained these grow
@@ -1301,35 +1301,34 @@ class SessionDB:
)
def _try_wal_checkpoint(self) -> None:
"""Best-effort TRUNCATE WAL checkpoint. Never raises.
"""Best-effort PASSIVE WAL checkpoint. Never raises.
Flushes committed WAL frames back into the main DB file and
truncates the WAL file to zero bytes. Keeps the WAL from
growing unbounded when many processes hold persistent
connections.
Flushes committed WAL frames back into the main DB file without
requiring an exclusive lock. PASSIVE is safe for frequent
periodic use because it does not block concurrent writers and
cannot corrupt B-tree pages under I/O pressure.
PASSIVE checkpoint was previously used here, but it never
truncates the WAL file — the file stays at its high-water
mark until an explicit TRUNCATE is called (which only
happened inside the infrequent vacuum()).
PASSIVE does not truncate the WAL file — it stays at its
high-water mark. WAL truncation happens in :meth:`close`
(TRUNCATE) and pre-VACUUM checkpoints, which run infrequently
under controlled conditions.
TRUNCATE may block writers briefly while checkpointing, but
_try_wal_checkpoint is called off the hot path (every 50
writes) and already runs under ``self._lock``, so the
additional hold time is negligible.
Previous TRUNCATE strategy caused B-tree corruption on large
databases (65K+ pages) due to the exclusive-lock I/O pressure
from checkpointing thousands of frames at once (issue #45383).
"""
try:
with self._lock:
result = self._conn.execute(
"PRAGMA wal_checkpoint(TRUNCATE)"
"PRAGMA wal_checkpoint(PASSIVE)"
).fetchone()
if result and result[1] > 0:
logger.debug(
"WAL checkpoint: %d/%d pages checkpointed",
result[2], result[1],
)
except Exception:
pass # Best effort — never fatal.
except Exception as exc:
logger.warning("WAL checkpoint (PASSIVE) failed: %s", exc)
def _try_optimize_fts(self) -> None:
"""Best-effort FTS5 segment merge. Never raises.
@@ -1357,8 +1356,8 @@ class SessionDB:
if self._conn:
try:
self._conn.execute("PRAGMA wal_checkpoint(TRUNCATE)")
except Exception:
pass
except Exception as exc:
logger.debug("WAL checkpoint (TRUNCATE) at close failed: %s", exc)
self._conn.close()
self._conn = None
@@ -7015,8 +7014,8 @@ class SessionDB:
# Best-effort WAL checkpoint first, then VACUUM.
try:
self._conn.execute("PRAGMA wal_checkpoint(TRUNCATE)")
except Exception:
pass
except Exception as exc:
logger.debug("WAL checkpoint (TRUNCATE) before VACUUM failed: %s", exc)
self._conn.execute("VACUUM")
return optimized