hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
cron/scheduler.py no longer re-exports the split modules (scheduler_delivery /
_script / _prompt / _preflight); it imports only the 19 names it calls itself
(bottom-of-file, E402 kept for the import cycle). Dropped the shim-only
`import shutil` and the F401 note on windows_hide_flags (still used by
scheduler.py). Split modules now call same-module helpers directly, reach
sibling split modules via late-bound module refs (_delivery/_script/_preflight)
next to _sched, and import windows_hide_flags themselves; origin-resident names
(load_config, Path, _SCRIPT_TIMEOUT, heartbeat_run_claim, ...) still go through
_sched. Callers/tests import + patch the defining module.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Review fold-in on the salvage of #101877:
- `_deliver_result` routed to the durable queue whenever the worker's
`_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
delivering. A worker whose script dispatches another job in-process
(`hermes cron run <other>`) inherits that env and would have queued the
nested job's message under the outer execution id — `INSERT OR IGNORE`
then drops it silently. Match the marker against the delivering job's own
`execution_id`, as `run_one_job` already does. Regression test added
(mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
reported as a delivery error, so `mark_job_run` recorded
`last_status=delivery_failed` for a message the next gateway's drain goes
on to send, and nothing ever corrects the job record. Log and return
success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
of a second hardcoded terminal set.
Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).
cron/scheduler.py 6386 -> 3423. Four sibling modules hold the free-function clusters,
re-exported from cron.scheduler so every existing import/monkeypatch target keeps resolving;
origin-resident helpers are reached late-bound via `_sched` (same-object visibility, no
import cycle). Every moved node is AST-identical to the base. Source-inspection test for the
bash resolver repointed to cron/scheduler_script.py.
cron/scheduler.py (8535 -> 6353):
- _deliver_result split into per-target helpers: _resolve_target_transport (live/relay/standalone
transport + enablement), _deliver_via_live_adapter (_live_route_metadata for Telegram DM-topic
vs forum routing, _live_send_text with the cancel()-based timeout disambiguation,
_live_send_media, _seed_live_delivery_sessions), _standalone_send/_deliver_standalone. The
three interpreter-shutdown skip branches and the repeated log+append+continue pattern collapse
into one _note_target_error / one shutdown message; _TargetDelivery carries per-target state.
- Thread/channel session seeding unified into _seed_cron_session (was two near-identical
functions).
- run_job decomposed into _run_no_agent_job, _apply_monitor_gate, _load_cron_job_config
(_CronJobConfig), _resolve_job_runtime, _check_model_drift, _open_cron_session_db,
_run_agent_with_watchdog, _finalize_cron_session, plus one _run_doc_header and one _audit
closure for the success/failure paths.
- _run_one_job_body: ownership-lost bookkeeping, delivery composition and outcome classification
extracted; tick: _acquire_tick_lock/_release_tick_lock, _maybe_reap_dead_owners,
_sweep_stale_inflight_for_tick, _process_due_job, _submit_with_guard.
- _build_job_prompt: context_from injection and skill loading extracted; one
_prepend_context_block for the four fenced-data blocks.
- One _start_heartbeat_thread for the script- and fire-claim heartbeat threads.
- Dropped unreachable return in SharedRouteAdapters.get; boolean-return and nested-if shapes
collapsed.
- Comments/docstrings compacted by hand, rationale kept (fd-leak reason for the late SessionDB
close callback, title-persistence rules, no_agent classification gate, inactivity-vs-provider
timeout ordering, stale-claim force-release, interruption token keying).
Follow-up to the failure_deliver salvage (#100375):
- _preflight_check_delivery also checks the failure lane, so a typo'd
failure_deliver platform blocks at config-validation time instead of
surfacing only when a failure occurs — exactly when the notice must
not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
deliver (text normalization; empty clears the optional override
instead of coalescing), closing the one update path that could write
an unnormalized value into jobs.json.
4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
Review findings (Salt, NS-788):
B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.
S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.
S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.
Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.
- cron/scheduler.py: factor the preflight's primary-config route loader into
`_primary_profile_routes_for_current_home()` (one owner for both halves,
so route semantics cannot drift) and add `SharedRouteAdapters`, a
read-only view over the primary adapter map that resolves an adapter for
a (platform, target) ONLY when an enabled primary route with a
chat_id/thread_id maps that exact target to the current profile —
using the same `ProfileRoute.matches` predicate as inbound routing.
`_deliver_result` resolves the transport per target from it; everything
else (unmatched chat, disabled route, route for another profile, no
primary adapter, guild-only route) is a miss and never uses the primary
bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
fallback: with no matching route the view is falsy and delivers nothing.
Fixes#101113
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:
1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
the caller already has a running loop — the desktop dashboard shape) ran
the standalone sender on a fresh thread with NO profile ContextVars: the
home override and secret scope were gone, so the sender resolved the
process default's home/token (or, fail-closed under multiplex, raised
UnscopedSecretError). Wrap the submit in copy_context().run like the
session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.
2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
whose own gateway (with live adapters) is running; winning the tick-lock
race meant the adapter-less desktop ticker delivered standalone. The
multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
the desktop wires it to `_check_gateway_running(home)` so such profiles
are neither ticked nor heartbeated by the dashboard while their gateway
is alive (re-evaluated every cycle, no restart needed).
Fixes#100489
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
A cron job fired, the scheduler logged "delivered to telegram:<chat> via
live adapter", and nothing reached Telegram (#77763). The log line was not
evidence of a send:
* the silence-narration filter returns {"success": True, "delivered": False}
(a successful *drop*), and the dict-normalization branch read only
"success", so a filtered message counted as delivered;
* an empty payload (no text, no media) skipped the send entirely and still
fell into the "delivered" branch;
* the log line named the chat but not the lane, so a wrong-thread delivery
and a phantom one are indistinguishable after the fact.
_confirm_adapter_delivery now inspects both result shapes: an explicit
`delivered: False` is a rejection even with a truthy `success`, and a
success with no message_id and no raw_response is accepted but logged as
UNVERIFIED. The empty-payload case fails closed into the existing
standalone/warn handling, and the delivered log carries thread= and
message_id=.
Failing closed on the live lane is only half the fix on a native target:
the standalone fallback sent the same empty payload, and the Telegram
adapter returns SendResult(success=True) for empty content without an API
call — a phantom live delivery became a phantom standalone one. Both
_send_to_platform call sites now sit behind one skip guard, so "empty
payload fails closed" holds on every lane (#77763).
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
on every surface
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.
tick() now checks, BEFORE acquiring the tick lock:
skew detected (boot fingerprint != disk revision)
AND this process does not own the gateway runtime lock
AND that lock is held (a fresh gateway is alive)
-> raise CronTickYielded, skipping the tick entirely
Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
proceed; yielding is a certainty claim, never a guess
The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.
Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.
gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).
When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).
The #83182 fix moved the secret scope to span execute→deliver so the
delivery path resolves platform TOKENS (e.g. TELEGRAM_BOT_TOKEN) from the
job-owning profile's .env instead of the host process os.environ. But
cron delivery also resolves the DESTINATION via the legacy
<PLATFORM>_HOME_CHANNEL env mirror, and _env_home_target_chat_id /
_get_home_target_thread_id read it through raw os.getenv — NOT
get_secret. So under multiplex, a socrates cron job whose tick was won
by the default gateway process resolved DISCORD_HOME_CHANNEL from the
default profile's os.environ and landed in the default pax-hermes-chat
instead of socrates-hermes-chat.
Make the chat-id and thread-id resolution legs read through get_secret
when a profile secret scope is installed (mirroring the token leg), so
the owning profile's home channel wins. Non-multiplex / no-scope callers
keep reading os.environ as before.
Verified: with os.environ=default chat and the socrates scope active,
_env_home_target_chat_id('discord') returns the socrates chat; with no
scope it returns the os.environ legacy value.
Review follow-up: the comment records that the reset scopes delivery,
deferred-agent teardown, claim-loss handling and bookkeeping, and that an
earlier inner-finally placement left _deliver_result unscoped — so a
tidy-up does not silently reintroduce the bug.
Salvaged from #56876 (cron half only; the delegation half is superseded
by #98237). run_job's ephemeral AIAgent constructor passed api_key /
base_url / provider / api_mode from the resolved runtime but dropped
request_overrides, so cron jobs on custom providers silently lost
extra_body / extra_headers request settings.
Replace #96637's inline active_profile_homes() closure with #96508's
module-level _existing_profile_homes() filter (testable in isolation).
Widen _ensure_cron_dir from 3 to 12 mkdir sites across cron/ so every
directory creation fails closed for deleted named profiles, not just
the 3 originally protected. Add _is_named_profile_path() that checks
'profiles' in path parts (works for subdirs like cron/output/<job> and
scripts/ that the original parent.name heuristic couldn't reach).
Co-authored-by: misterdas <das7514@gmail.com>
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
extract it next to _get_script_timeout/_get_media_send_timeout. Using
load_config() (deep-merge) also removes the default-drift hazard flagged
in review: cron.session_db_timeout_seconds now resolves from
DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
10.0 staying in sync with it, and drops the distant-state coupling to
run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
new position the store init happens inside the guarded worker, not before
it).