hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Review fold-in on the salvage of #101877:
- `_deliver_result` routed to the durable queue whenever the worker's
`_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
delivering. A worker whose script dispatches another job in-process
(`hermes cron run <other>`) inherits that env and would have queued the
nested job's message under the outer execution id — `INSERT OR IGNORE`
then drops it silently. Match the marker against the delivering job's own
`execution_id`, as `run_one_job` already does. Regression test added
(mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
reported as a delivery error, so `mark_job_run` recorded
`last_status=delivery_failed` for a message the next gateway's drain goes
on to send, and nothing ever corrects the job record. Log and return
success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
of a second hardcoded terminal set.
Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).