b440a492b3
Follow-up to the salvaged restart-safe worker (#101877): - delivery_queue: a row still `pending` at the worker's wait timeout was marked `failed` and never drained, so any gateway outage longer than the 300s budget (e.g. a restart that runs `hermes update`) silently lost the delivery. Unclaimed rows are certainly unsent, not uncertain — leave them queued for the next gateway; only mid-send rows are fenced `unknown`. - delivery_queue: stop running the full-table prune UPDATE+COUNT inside every transaction (each `get_status` poll paid for it; terminalizing paths already prune explicitly); poll at 1s instead of 250ms. - delivery_queue/executions: use `hermes_state.apply_wal_with_fallback` (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied owner-liveness helpers in favour of the ones in cron.executions. - scheduler: the parent waited on the worker by re-opening the executions ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the process with a 1s timeout instead — the worker commits its terminal row before exiting — and reap stranded payload/ack files once terminal. - scheduler: skip the housekeeping drain until a worker has actually created deliveries.db, so non-systemd gateways never open it. - scheduler: set up hermes logging in the detached worker entrypoint; it runs with stdout/stderr on DEVNULL and previously logged nowhere. - tests: test_lost_fire_claim_stops_stale_delivery still mocked `mark_execution_running -> None`, which now means "ownership lost, return before run_job" — the test passed without ever reaching the path it names. Mocking `{}` restores it (mutation-checked).