Files
hermes-agent/tests
teknium1 aa5817d9be fix(kanban): reap workers that outlive their finished run
A worker that called kanban_complete and then hung (e.g. holding deleted
state.db-wal/-shm inodes, which trips the DeletedWalGenerationError guard on
every later write) was unreachable by any command: the terminal transition
cleared tasks.worker_pid, _end_run cleared task_runs.worker_pid too, and every
reclaim sweep only looks at status='running' cards (#111791).

Keep the evidence and add the consumer: task_runs gains worker_started_at (the
spawn-time fingerprint tasks already carry), _set_worker_pid stamps it, and
_end_run leaves worker_pid / worker_started_at / claim_lock on the closed row.
reap_terminal_workers runs in the dispatcher's reclaim phase (every tick and
`hermes kanban dispatch --once`): a host-local pid on a closed run that is
still the fingerprinted process is terminated through the existing
_terminate_reclaimed_worker (SIGTERM, then SIGKILL after the poll window) and
recorded as a terminal_worker_reaped event; a pid that is gone or recycled
only has its evidence cleared; legacy rows without a fingerprint are never
signalled.

Slimmer redo of PR #111798 by @KoNit-K: same schema + retention shape, but the
reaper reuses _worker_alive / _terminate_reclaimed_worker(started_at=) instead
of a second start-time reader and a guarded-kill closure, scans every closed
run instead of a task-status allowlist, and clears dead evidence so rows are
not rescanned forever.

Fixes #111791
2026-09-15 18:35:32 -07:00
..
…