Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).
- scheduler_provider.py: _profile_entry/_profile_cron_scope replace four hand-rolled
home-override+store blocks in the multiplex ticker; comments compacted to the WHY.
- incidents/executions/notepad/monitor/suggestions/blueprint_catalog/__init__: docstrings
and comments compacted; no code change (AST-identical).
- claim_job_for_fire returns the atomically claimed snapshot with a unique
fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
terminal writes by expected_fire_owner so a stale worker cannot record
over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
event (ownership loss OR external cancel) into run_job; the agent path
is interrupted cooperatively and script-based jobs (no_agent + pre-run
scripts) are hard-stopped with a process-tree kill (POSIX killpg
SIGTERM then SIGKILL for surviving group members; Windows
taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
the bare job ID, so a replacement run of the same job never consumes a
stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
completed one-shot retention (#80624), blocked_config preflight
(T1-26), and the advance_next_runs batch on top of current main.
Use database.journal_mode as the sole non-secret operator setting, preserve the vulnerable-SQLite safety gate and existing WAL databases, validate explicit DELETE results, document the active config path, and cover real SQLite openers with behavioral tests.
Add HERMES_JOURNAL_MODE env / database.journal_mode config for
virtiofs/NFS/SMB where WAL is not crash-safe. Route 5 bypass openers
through apply_wal_with_fallback so a single setting covers every .db
(#68545).
On vulnerable SQLite (e.g. 3.50.4), do not enable WAL for fresh/non-WAL
shared databases — prefer DELETE instead. Leave existing on-disk WAL
alone (no live downgrade under concurrent gateway/cron openers). Surface
Python/SQLite version details as a doctor warning (#69784).