cron/scheduler.py (8535 -> 6353):
- _deliver_result split into per-target helpers: _resolve_target_transport (live/relay/standalone
transport + enablement), _deliver_via_live_adapter (_live_route_metadata for Telegram DM-topic
vs forum routing, _live_send_text with the cancel()-based timeout disambiguation,
_live_send_media, _seed_live_delivery_sessions), _standalone_send/_deliver_standalone. The
three interpreter-shutdown skip branches and the repeated log+append+continue pattern collapse
into one _note_target_error / one shutdown message; _TargetDelivery carries per-target state.
- Thread/channel session seeding unified into _seed_cron_session (was two near-identical
functions).
- run_job decomposed into _run_no_agent_job, _apply_monitor_gate, _load_cron_job_config
(_CronJobConfig), _resolve_job_runtime, _check_model_drift, _open_cron_session_db,
_run_agent_with_watchdog, _finalize_cron_session, plus one _run_doc_header and one _audit
closure for the success/failure paths.
- _run_one_job_body: ownership-lost bookkeeping, delivery composition and outcome classification
extracted; tick: _acquire_tick_lock/_release_tick_lock, _maybe_reap_dead_owners,
_sweep_stale_inflight_for_tick, _process_due_job, _submit_with_guard.
- _build_job_prompt: context_from injection and skill loading extracted; one
_prepend_context_block for the four fenced-data blocks.
- One _start_heartbeat_thread for the script- and fire-claim heartbeat threads.
- Dropped unreachable return in SharedRouteAdapters.get; boolean-return and nested-if shapes
collapsed.
- Comments/docstrings compacted by hand, rationale kept (fd-leak reason for the late SessionDB
close callback, title-persistence rules, no_agent classification gate, inactivity-vs-provider
timeout ordering, stale-claim force-release, interruption token keying).
Follow-up to the failure_deliver salvage (#100375):
- _preflight_check_delivery also checks the failure lane, so a typo'd
failure_deliver platform blocks at config-validation time instead of
surfacing only when a failure occurs — exactly when the notice must
not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
deliver (text normalization; empty clears the optional override
instead of coalescing), closing the one update path that could write
an unnormalized value into jobs.json.
4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
Review findings (Salt, NS-788):
B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.
S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.
S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.
Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.
- cron/scheduler.py: factor the preflight's primary-config route loader into
`_primary_profile_routes_for_current_home()` (one owner for both halves,
so route semantics cannot drift) and add `SharedRouteAdapters`, a
read-only view over the primary adapter map that resolves an adapter for
a (platform, target) ONLY when an enabled primary route with a
chat_id/thread_id maps that exact target to the current profile —
using the same `ProfileRoute.matches` predicate as inbound routing.
`_deliver_result` resolves the transport per target from it; everything
else (unmatched chat, disabled route, route for another profile, no
primary adapter, guild-only route) is a miss and never uses the primary
bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
fallback: with no matching route the view is falsy and delivers nothing.
Fixes#101113
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:
1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
the caller already has a running loop — the desktop dashboard shape) ran
the standalone sender on a fresh thread with NO profile ContextVars: the
home override and secret scope were gone, so the sender resolved the
process default's home/token (or, fail-closed under multiplex, raised
UnscopedSecretError). Wrap the submit in copy_context().run like the
session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.
2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
whose own gateway (with live adapters) is running; winning the tick-lock
race meant the adapter-less desktop ticker delivered standalone. The
multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
the desktop wires it to `_check_gateway_running(home)` so such profiles
are neither ticked nor heartbeated by the dashboard while their gateway
is alive (re-evaluated every cycle, no restart needed).
Fixes#100489
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
A cron job fired, the scheduler logged "delivered to telegram:<chat> via
live adapter", and nothing reached Telegram (#77763). The log line was not
evidence of a send:
* the silence-narration filter returns {"success": True, "delivered": False}
(a successful *drop*), and the dict-normalization branch read only
"success", so a filtered message counted as delivered;
* an empty payload (no text, no media) skipped the send entirely and still
fell into the "delivered" branch;
* the log line named the chat but not the lane, so a wrong-thread delivery
and a phantom one are indistinguishable after the fact.
_confirm_adapter_delivery now inspects both result shapes: an explicit
`delivered: False` is a rejection even with a truthy `success`, and a
success with no message_id and no raw_response is accepted but logged as
UNVERIFIED. The empty-payload case fails closed into the existing
standalone/warn handling, and the delivered log carries thread= and
message_id=.
Failing closed on the live lane is only half the fix on a native target:
the standalone fallback sent the same empty payload, and the Telegram
adapter returns SendResult(success=True) for empty content without an API
call — a phantom live delivery became a phantom standalone one. Both
_send_to_platform call sites now sit behind one skip guard, so "empty
payload fails closed" holds on every lane (#77763).
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
on every surface
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.
tick() now checks, BEFORE acquiring the tick lock:
skew detected (boot fingerprint != disk revision)
AND this process does not own the gateway runtime lock
AND that lock is held (a fresh gateway is alive)
-> raise CronTickYielded, skipping the tick entirely
Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
proceed; yielding is a certainty claim, never a guess
The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.
Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.
gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).
When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).
The #83182 fix moved the secret scope to span execute→deliver so the
delivery path resolves platform TOKENS (e.g. TELEGRAM_BOT_TOKEN) from the
job-owning profile's .env instead of the host process os.environ. But
cron delivery also resolves the DESTINATION via the legacy
<PLATFORM>_HOME_CHANNEL env mirror, and _env_home_target_chat_id /
_get_home_target_thread_id read it through raw os.getenv — NOT
get_secret. So under multiplex, a socrates cron job whose tick was won
by the default gateway process resolved DISCORD_HOME_CHANNEL from the
default profile's os.environ and landed in the default pax-hermes-chat
instead of socrates-hermes-chat.
Make the chat-id and thread-id resolution legs read through get_secret
when a profile secret scope is installed (mirroring the token leg), so
the owning profile's home channel wins. Non-multiplex / no-scope callers
keep reading os.environ as before.
Verified: with os.environ=default chat and the socrates scope active,
_env_home_target_chat_id('discord') returns the socrates chat; with no
scope it returns the os.environ legacy value.
Review follow-up: the comment records that the reset scopes delivery,
deferred-agent teardown, claim-loss handling and bookkeeping, and that an
earlier inner-finally placement left _deliver_result unscoped — so a
tidy-up does not silently reintroduce the bug.
Salvaged from #56876 (cron half only; the delegation half is superseded
by #98237). run_job's ephemeral AIAgent constructor passed api_key /
base_url / provider / api_mode from the resolved runtime but dropped
request_overrides, so cron jobs on custom providers silently lost
extra_body / extra_headers request settings.
Replace #96637's inline active_profile_homes() closure with #96508's
module-level _existing_profile_homes() filter (testable in isolation).
Widen _ensure_cron_dir from 3 to 12 mkdir sites across cron/ so every
directory creation fails closed for deleted named profiles, not just
the 3 originally protected. Add _is_named_profile_path() that checks
'profiles' in path parts (works for subdirs like cron/output/<job> and
scripts/ that the original parent.name heuristic couldn't reach).
Co-authored-by: misterdas <das7514@gmail.com>
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
extract it next to _get_script_timeout/_get_media_send_timeout. Using
load_config() (deep-merge) also removes the default-drift hazard flagged
in review: cron.session_db_timeout_seconds now resolves from
DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
10.0 staying in sync with it, and drops the distant-state coupling to
run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
new position the store init happens inside the guarded worker, not before
it).
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.
Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.
Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
- _target_mirror_eligible accepts a precomputed origin_match so the sole
production caller stops re-resolving origin + re-running the origin
match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
precedence (standalone correctness: per-job False must beat raw global
True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
'not inchannel_continuable' and reword it for the widened seed scope.
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.
Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).
Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
- Make repair_explicit_computer_use_media_paths fail-open internally
(cosmetic repair must never abort delivery); drop the cron-only
try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
gateway.media_repair seam; pre-existing tests keep pinning the
gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
Follow-up to the salvaged fix from PR #94439:
- Extract the repair into gateway/media_repair.py (shared module) and
re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
background tasks (_run_background_task_inner) and cron job delivery
(cron/scheduler.py) — both call agent.run_conversation directly and
never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
content first instead of regex-scanning the raw string, which yielded a
doubled-backslash path artifact and rewrote the response to a path the
model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
the compression-fallback last-user slice (incl. no-user fail-closed).
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.
The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.
Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.
Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.
Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.
Fixes#88655
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.
Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.