A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.
The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.
Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.
Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.
Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.
Fixes#88655
cron_delivery_targets() now also lists machine-local bot-chat:<profile>
entries; the sibling test's exact set-equality assertion predates them.
Scope the platform assertions to gateway entries and pin that bot-chat
entries are always home_target_set.
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.
Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
seed_mock.assert_not_called() alone could pass for the wrong reason — a
harness failure before delivery also leaves the seed uncalled. Assert
the real adapter recorded exactly one live send to the origin chat, so
the test pins the D6 thread-fallback decision, not an accidental
no-delivery.
Unspecced MagicMock/AsyncMock adapters fabricate
supports_inchannel_continuable_for_platform as a truthy callable, so the
scheduler's duck-typed D6 gate silently took the relay accessor branch in
every in-channel test — the native scalar fallback the fixtures describe
was never exercised, and setting supports_inchannel_continuable=False on
a mock could not force thread mode. Pin the accessor to None on both
mock fixtures (matching a real native adapter, which never defines the
method), and add a fallback-boundary test with a real minimal adapter
class: scalar False -> in_channel fails safe to thread, flat seed never
fires.
The seed-key fix made the SESSION scoped, but the delivery leg still
dropped the scope: cron route_metadata carried only job_id (+thread),
DeliveryRouter stamps scope_id only for the configured HOME channel, and
the RelayAdapter's per-chat scope cache is cold after a gateway restart
(learned from inbound only). A scoped Slack origin that is not the home
chat therefore egressed with NO tenant discriminator, and the connector's
fail-closed guard could reject the brief before delivery — the
delivery-leg sibling of the seed-key scope gap.
Copy origin.scope_id into the live text and media routing metadata for
ORIGIN-MATCHING targets only (setdefault — never overrides router/home
stamping). Fan-out/broadcast targets are excluded by the origin gate: a
fan-out target's tenant is not the origin's, and a wrong scope is worse
than none.
Tests: restart-shaped positive (scoped non-home origin -> scope_id on
routed metadata, RED before this fix) and legacy negative (scope-less
origin stamps nothing).
The seed was decoupled from the mirror opt-in (in_channel is the
continuation surface regardless of attach_to_session), but the
thread-id-clearing gate above it still read mirror_this_target. With the
advertised default config (attach_to_session=false, cron.mirror_delivery
unset) and an origin carrying a real thread_id, the brief kept delivering
INTO the origin thread while the flat (thread_id=None) session got
seeded — brief and continuation surface in different places, so a plain
reply never saw it.
Flatten on the same gate as the seed: origin_target (with the existing
live_adapter_ready guard). Fan-out/broadcast targets are unaffected.
Test drives _deliver_result with a thread-carrying origin and default
knobs, asserting on the routed DeliveryTarget.thread_id — RED on the old
gate, GREEN now.
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.
Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.
Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
Live incident (Alice canary 2026-08-20, job 8e21a957b77b): the continuable
thread seed created its session with chat_type='thread', but a Slack DM
in-thread reply arrives chat_type='dm' and build_session_key routes DM
threads through the DM arm (...:dm:<chat>:<thread>). Seed row and reply row
never matched — the reply had no brief in context (continuation amnesia).
is_dm on _seed_cron_thread_session selects the seeded chat_type at both call
sites (opened-thread and the companion in_channel thread seed); channel
threads are unchanged. Sibling lane of the flat seed's is_dm fix. Tests pin
the key-equality contract: seeded key == the key the reply builds.
Two live failures from the Alice canary (2026-08-19, jobs 28a24afebd81 /
83b93f8be379), both leaving a continuable in_channel cron with amnesia:
1. Seed mirrored via origin heuristics and silently dropped the brief.
_seed_cron_channel_session created the flat session row, then
mirror_to_session RE-DISCOVERED the target via find_session_by_origin —
whose multi-candidate bail-out returns None on a populated chat (flat
session + N per-message thread sessions sharing one chat_id, mixed
user_ids). Receipt: 'in_channel seed did NOT land on slack:D0BJTDCSR7C'.
mirror_to_session now accepts an explicit session_id and both cron seeds
pass the exact row they just created; origin-scan remains the fallback
for callers that genuinely don't know the target.
2. The brief's OWN THREAD was never seeded. in_channel delivers flat, but
a flat Slack message still invites a thread reply (the natural mobile
affordance — exactly what the user did). That reply keys to
(chat, thread=<brief ts>), which no seed touched. The delivery's
message_id now anchors a companion _seed_cron_thread_session so BOTH
reply surfaces (plain channel message AND in-thread reply) continue the
job.
Also: thread-seed failures upgraded debug→WARNING (silent seed failure IS
the user-facing bug), and the thread seed reports landed/not-landed.
Regression tests drive both against the live failure shapes: exact-session
mirror asserted via session_id kwarg; thread companion asserted via the
SendResult message_id anchor. Clean-fixture blind spot noted: the E2E
harness used a fresh store with one row, which is why heuristic rediscovery
looked fine pre-production.
Live regression (Alice, 2026-08-19): a continuable cron with
cron_continuable_surface=in_channel delivered its brief flat, but the
flat-session seed was gated on mirror_this_target = mirror_enabled AND
origin-match. Without attach_to_session (and with cron.mirror_delivery
defaulting False) the seed never ran; the next plain reply resolved to a
blank (slack, chat, None) session and the agent had no idea about its own
delivery message — not continuable in channel OR thread (in_channel mode
correctly skips thread creation, so there was no thread session either).
in_channel IS the continuation surface, not a mirror nicety: gate the seed
on origin-match alone, and resolve origin_user_id for any origin-matching
target so the seeded key still carries the scheduling user on
per-user-isolated chats. attach_to_session remains the opt-in for the
separate default-surface mirror behavior.
Regression test drives the delivery path with attach_to_session=False and
asserts the seed fires with the right user_id (fails on pre-fix code).
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
Field report (enterprise, v0.20.0): cron jobs delivering text + PDF/image
attachments to Slack DMs deliver both on scheduled ticks but text-only on
manual `hermes cron run <job-id>`. Same box, same token, same scopes —
the divergence is process context and error visibility, not credentials.
Three defects, one bug class (attachment failures invisible + policy
divergence between the gateway process and standalone processes):
1. Standalone lane swallowed warnings: platform standalone senders
(Slack files_upload_v2, Discord, ...) report per-file upload failures
in result['warnings'] while returning success=True for the delivered
text leg. _deliver_result only read result['error'], so the run was
marked ok and the attachment vanished without a trace. Warnings now
surface into delivery_errors (and the job's last_error).
2. Live-adapter lane swallowed media failures: _send_media_via_adapter
logged failures at WARNING and returned None. It now returns per-file
error strings and _deliver_result records them — text-delivered-but-
attachment-failed is a visible partial failure on BOTH lanes.
3. Media-policy env bridge was gateway-only: gateway.strict /
media_delivery_allow_dirs / trust_recent_files were translated from
config.yaml to the env vars validate_media_delivery_path reads ONLY in
gateway startup. A CLI-process manual run filtered attachment paths
under a different policy — in strict/allowlisted deployments the exact
reported symptom (scheduled delivers, manual drops, silently). The
translation now lives in gateway/media_policy.apply_media_policy_env
(idempotent, env-wins, never raises); gateway startup delegates to it
and _deliver_result applies it before filtering. Attachments dropped
by the policy filter are also reported in the run status instead of
only a stderr WARNING.
On v0.20.0 specifically the failure was double-blind: the pre-9cf2cbd382
isinstance(resp, dict) gates meant upload failures were undetectable in
the sender AND unsurfaced by the scheduler. 9cf2cbd382 (in 2026.8.13)
fixed detection; this fixes visibility and policy parity.
8 new tests (tests/cron/test_media_delivery_parity.py): warnings→errors,
clean-delivery control, media-reaches-sender control, live-adapter
failure/dropped-path reporting, bridge helper semantics, strict+allowlist
end-to-end in a non-gateway process, and the .env-strict/config-allowlist
split that reproduces the field symptom. Mutation check: disabling the
warnings loop and the bridge fails exactly the 2 guarding tests.
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).
fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:
- No-op for the built-in ticker (its tick loop already self-heals
past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
external retry is de-duplicated) and runs fire_claimed in a daemon
thread, mirroring the webhook admission pattern, so housekeeping
never blocks for the length of an agent run. Provider re-arm logic
(Chronos NAS one-shots) runs exactly as for a normal fire.
Docs: cron.md section + cron.misfire_grace_minutes reference.
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.
Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.
9/9 tests green; ruff clean.
Follow-ups on the #87591 salvage:
- cron/scheduler.py: wrap the three clear_run_claim call sites in a
best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
and on the interpreter-shutdown path (or with a corrupt store) it could
itself raise, defeating the skip-cleanly purpose of these early exits.
A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
clear_run_claim unit contract (one-shot cleared / already-clear noop /
recurring never touched / unknown id), all three dispatch-failure paths
through a real tick() clear the claim, and a raising clear_run_claim
does not crash the tick. Mutation-verified: reverting the fix makes the
suite fail.
Follow-ups on the #87259 salvage:
- cron/scheduler.py: the ledger-terminal reconciliation now requires the
terminal execution row's claimed_at to be >= the in-memory claim's
registration time (_running_since). Without this, the latest terminal
row for a recurring job is usually the PREVIOUS run's outcome — a fresh
claim in the try_register_running_job -> create_execution window (or a
finished run whose worker finally block hasn't released yet) would be
force-released and the job double-dispatched. Unparseable/missing
claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
query under _running_lock — list() over a set concurrently mutated by
try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
two new race-guard tests (previous-run terminal row never releases a
fresh claim; missing claimed_at fails closed). Mutation-verified:
removing the ownership guard fails both.
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.
sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).
Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
Follow-ups on the #87261 salvage:
- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
Re-arming to `now` fired CRON jobs at times their expression excludes —
a weekday-only 9am job whose Friday run errored would fire on SATURDAY
(croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
— the next LEGAL occurrence — and only when that actually moves
next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
_jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
errored Friday re-arms to Monday (not Saturday), correctly-parked cron
value untouched, interval job still due immediately.
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.
cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.
Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
"\x00" in script_path raises TypeError when a caller passes a non-str
(e.g. a pathlib.Path, which is not iterable) — the guard itself would
crash the scheduler. All current call sites pass plain str, but the
guard must be crash-proof: str() first, then check. Adds a regression
test running a real script through _run_job_script with a Path argument,
which fails with TypeError on the pre-fix guard.
_windows_cron_python_invocation bypasses the uv venv launcher (to avoid
flashing a console window) and re-attaches the venv via PYTHONPATH — but
PYTHONPATH entries are plain sys.path additions and never get .pth
processing, so editable installs (pip install -e) were invisible to cron
script jobs (ModuleNotFoundError).
Bootstrap the script with site.addsitedir() on the venv site-packages,
then exec it as __main__ via runpy.run_path, preserving the script
directory on sys.path (python script.py semantics). Falls back to a
plain invocation when the venv layout is unresolvable.
Fixes#86721.
`hermes cron run <job_id>` (a one-shot CLI invocation) dispatches
manual runs via the same background-delegation path as an agent's
`cronjob(action='run')` tool call (tools/cronjob_tools.py's
_try_dispatch_background_run -> dispatch_async_delegation(role=
"cron_run", runner=_runner, ...)). The runner thread lives in the
calling process's shared daemon executor. When the one-shot process
exits right after printing "Triggered job: ...", the in-flight runner
dies mid-execution, leaving its cron/executions.db row permanently
stuck at status='claimed' -- every subsequent `hermes cron run` on the
same job then reports "Ran now: failed" because of the still-claimed
row.
cron/executions.py already has the exact self-heal this needs:
recover_interrupted_executions() correctly identifies and reclassifies
'claimed'/'running' rows whose owner process has provably exited
(_owner_is_live checks PID existence AND matches process start-time,
so a reused PID isn't mistaken for the original live owner) to
'unknown', unblocking the job for a fresh claim. But it was only ever
called once, at the long-lived scheduler ticker's own startup
(cron/scheduler.py:379's self.recover_interrupted()) -- a one-shot CLI
invocation has no equivalent "startup" moment of its own, so this
self-heal never ran for it.
Added a call to recover_interrupted_executions() at the top of
_try_dispatch_background_run, right after the async-delivery-supported
gate and before any claim attempt for the current job -- mirroring
exactly what the long-lived scheduler already does at its own
startup, just triggered per one-shot invocation instead of once at
daemon startup. Wrapped in try/except: pass (best-effort; a failure
here must not block the actual dispatch this function exists for).
Traced (but did not attempt to fix) the deeper "why does the runner
die with the process at all" question -- that's the harder problem
options 1/2 in the issue describe (route to the persistent scheduler,
or block the one-shot process until completion). This fix addresses
the more urgent, more clearly-scoped symptom: a stranded stale claim
permanently blocking ALL future manual runs of the affected job, which
is option 3 from the issue and the one with an existing, already-
correct implementation just needing to be wired into this call site.
Added 3 regression tests to a new file, following the established
real-subprocess dead-owner pattern already used in
tests/cron/test_execution_ledger.py (a genuinely-dead PID, not a
mock, matching the real-world failure mode exactly): a sanity test
confirming the stale claim sits unrecovered without the fix; a direct
test of recover_interrupted_executions() reaping such a claim; and a
unit test on _try_dispatch_background_run itself confirming recovery
is called before any claim attempt. Verified as a genuine regression
by reverting the fix and confirming the unit test fails with recovery
never having been called.
35/35 pass across the new test file plus tests/cron/test_execution_ledger.py
and tests/tools/test_cronjob_run_background.py (no regression).
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.
Two-part fix:
1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
channel stateless (scoped ContextVar set/reset around the call) before
invoking the cron API, so `async_delivery_supported()` gates off
`_try_dispatch_background_run` and the run executes synchronously to
completion in the CLI process — the same behavior `hermes -z` already
gets via declare_stateless_channel().
2. cron/scheduler.py: tick() now periodically invokes
recover_interrupted_executions() (previously only run at scheduler
startup), so execution rows whose exact owner process is provably dead
(pid + process start time check in _owner_is_live) are reaped to
'unknown' by the long-lived gateway ticker without a restart.
Throttled to once per 300s so idle 60s ticks don't pay a ledger
connection every cycle.
Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.
Fixes#86721
The salvaged regression from #86582 predates the claim_job_for_fire
owner-fencing that landed with #70638; mock the claim and heartbeat so
the healthy job actually runs through the fenced flow.
Agent crons resolve OAuth credentials before the agent loop. A short
macOS/WARP DNS blip raised httpx.ConnectError ([Errno 8] nodename nor
servname provided) from xai-oauth token refresh, and the scheduler only
walked fallback_providers on AuthError — so Daily Focus Kickoff died
even when XAI_API_KEY / Anthropic were healthy.
Treat ConnectError/DNS OSError (and cause-chain equivalents) like
AuthError when selecting the fallback chain. Keep provider+model atomic.
Regression test covers the ConnectError path.
- tests/cron/test_sessiondb_init_hang.py: add threading/time imports the
salvaged late-close regression tests rely on.
- tests/test_hermes_state.py: drop
test_close_closes_wal_read_connection_created_on_worker_thread — main
replaced per-thread WAL reader ownership with the pooled read-connection
design (permits + checkout/return), so cross-thread reader draining no
longer exists in the form the test asserted.
run_job() submits SessionDB() to a one-worker executor and abandons the
worker (shutdown(wait=False)) when init exceeds the cron timeout. If the
constructor later completes inside that abandoned worker, the Future's
result — an open SessionDB holding .db/WAL/SHM handles — was orphaned and
never closed, leaking descriptors until EMFILE. Attach a done-callback on
the timeout path that retrieves and closes any eventual late result.
Salvage note: the lazy-recall ownership half of #72822 (_owns_session_db
tracked on AIAgent, owned handle closed in close()) already landed on main;
this carries the remaining cron timeout-abandon half with its regression
test.
Pin that cron/subagent non-streaming calls receive a read timeout matching the stale budget, that an explicit timeout is left alone, and that force_close_tcp_sockets finds sockets on httpcore PoolRequest.connection and clears the socket timeout before shutdown without close().
- claim_job_for_fire returns the atomically claimed snapshot with a unique
fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
terminal writes by expected_fire_owner so a stale worker cannot record
over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
event (ownership loss OR external cancel) into run_job; the agent path
is interrupted cooperatively and script-based jobs (no_agent + pre-run
scripts) are hard-stopped with a process-tree kill (POSIX killpg
SIGTERM then SIGKILL for surviving group members; Windows
taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
the bare job ID, so a replacement run of the same job never consumes a
stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
completed one-shot retention (#80624), blocked_config preflight
(T1-26), and the advance_next_runs batch on top of current main.
Restores pass-through behavior for cron delivery and react/unreact that
was lost when d409f6748 routed them through resolve_send_target. Stored
cron job targets the channel directory doesn't recognize (e.g.
telegram:ops-room on a fresh install, photon group GUIDs) used to go to
the adapter verbatim; after d409f6748 they were silently dropped. Same
for react on platform-native ids.
Adds an opt-in pass_unresolved_references flag to resolve_send_target,
passed only by cron and react. Model-facing send tool stays strict.
Plugin platforms with a parser stay strict for all callers. The optional
validator still has the final say over passed-through ids.
Follow-up fixes on salvage:
- Update test_cron_relay_delivery_guards.py mock lambdas to accept **kw
(file added to main after PR branch point; lambdas didn't accept the
new keyword argument)
- Consolidate duplicated pass-through blocks into _pass_through_unresolved
local helper
Fixes#85128
Co-authored-by: Adolanium <Adolanium@users.noreply.github.com>
Port t_3778a491's in-flight stale-claim guard, absent from origin/main.
_submit_with_guard adds a job id to _running_job_ids before the future
that owns its release exists. Anything that hangs or dies between the
add and pool.submit (EAGAIN thread exhaustion on a substrate spike, or a
wedged SessionDB.__init__ on a stale sqlite flock) leaks the claim; every
later tick short-circuits with 'already running - skipping' silently - no
execution row, no last_error, no counter - until the gateway process
restarts. This wedged 4 recurring no_agent router/watchdog jobs (verdict-
router, wake-scanner, auto-review-router, blocked-task-notifier) for ~1h47m
on 2026-08-14 (t_20e23f84), cleared only by manual force-run.
- Record claim timestamp + pending-future sentinel in the same critical
section as the add; replace sentinel with the owning future after submit.
- sweep_stale_inflight() runs every tick (even idle) and force-releases
claims older than max(2*interval, 30m floor) with no live future: WARNING
cron.inflight.forced_release, get_inflight_guard_stats() counter, JSONL
record, and mark_job_run(success=False) so the wedge surfaces as last_error.
- Wrap the pre-future init (create_execution/copy_context) so an exception
there releases the claim immediately instead of leaking it.
- Finite-repeat jobs are released without mark_job_run so a forced release
never consumes a one-shot budget.
Scheduler-internal only: no provider/model routing, no credentials, no
spend, no guardrail weakening, no cron permission widening.
Tests: tests/cron/test_inflight_stale_guard.py (18), plus regression tests
for the recurring EAGAIN re-dispatch and the create_execution/pool-submit
leak paths. Full tests/cron/: 616 passed.