The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
* fix(desktop): treat a macOS ps: parent-marker mismatch as inconclusive, not death
`ps -o lstart=` is a TZ/locale-rendered wall-clock string. Electron caches it
once per app lifetime; the backend re-renders it per spawn. A timezone change
(travel, or the DST boundary) while the app stays open makes the SAME instant
differ byte-for-byte, and _is_serve_orphaned() treated that as proof the parent
died -- os._exit(0) ~5ms after HERMES_BACKEND_READY, silently, on every respawn
until a full quit.
Route mismatches through _parent_start_marker_mismatch_is_conclusive(): linux:/
win:/winms: machine markers stay conclusive (PID-reuse defence intact); any ps:
side degrades to the PID-liveness check the legacy Desktop path already uses.
Log one warning before os._exit(0) so the exit is no longer traceless.
Fixes#95693Fixes#93958
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
* fix(desktop): reject truncated ps: parent markers; blank marker env means absent
A ps: marker that got whitespace-split in env plumbing (`ps:Sat`) passed
_valid_parent_start_marker and armed a watchdog that could never match, so
the backend exited 0 right after HERMES_BACKEND_READY. Require a full
lstart value (>=4 tokens, a 4-digit year, a time). Treat empty
HERMES_PARENT_START_MARKER / HERMES_PARENT_NONCE as absent so a blank
inherited value degrades to PID-only tracking instead of disarming or
misfiring.
* chore: map contributor email for drewTuzson
* fix(desktop): log when the parent-death watchdog disarms on an unusable marker
Disarming is the fail-safe branch (the backend keeps serving) but it also
means this backend will never reap itself when the Desktop dies. Leave one
warning naming the rejected marker so that downgrade is not traceless.
* test(desktop): live macOS proof that a TZ-drifted ps: marker no longer kills the backend
Runs the real start_server in a subprocess against the host ps, with the
marker rendered under Europe/Paris and the backend under America/New_York.
Asserts READY + still alive after several watchdog polls, then that the
backend still exits once the stand-in parent is killed. Fails on main with
the reported symptom (exit 0, live parent); macos_only lane.
---------
Co-authored-by: ygd58 <buraysandro9@gmail.com>
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
Co-authored-by: Drew Tuzson <drew.tuzson@uqual.com>
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).