7abe9502ee
tasks.worker_pid outlives a reboot; afterwards the number can belong to any process. _pid_alive answered from bare existence, so reclaim_stale_claims kept extending the claim of a "live" stranger (stuck running task), and enforce_max_runtime / _terminate_reclaimed_worker SIGTERM'd then SIGKILL'd it. _set_worker_pid now records gateway.status.get_process_start_time(pid) as tasks.worker_started_at (additive column, NULL on legacy rows). _worker_alive (pid, started_at) is the liveness check every reader uses (reclaim, defer, reconcile, crash sweep, max-runtime, archive, reopen invalidation); a live pid whose fingerprint disagrees is a recycled PID: treated as dead, never signalled (termination reports pid_recycled). Legacy rows without a fingerprint keep the existence answer until their next spawn. Row hermes_cli/kanban_db_dispatch.py:458 (lane4_high) confirmed by tracing: the SIGKILL at :461 was gated only on _pid_alive.