A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.
This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:
- int/float → int(raw) (floats truncated)
- numeric string ('120') → int(raw)
- 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
whitespace-tolerant) → sys.maxsize sentinel
- YAML None/null → default (90)
- bool/list/dict/garbage → default (with debug log)
All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.
The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.
Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
The seed-key fix made the SESSION scoped, but the delivery leg still
dropped the scope: cron route_metadata carried only job_id (+thread),
DeliveryRouter stamps scope_id only for the configured HOME channel, and
the RelayAdapter's per-chat scope cache is cold after a gateway restart
(learned from inbound only). A scoped Slack origin that is not the home
chat therefore egressed with NO tenant discriminator, and the connector's
fail-closed guard could reject the brief before delivery — the
delivery-leg sibling of the seed-key scope gap.
Copy origin.scope_id into the live text and media routing metadata for
ORIGIN-MATCHING targets only (setdefault — never overrides router/home
stamping). Fan-out/broadcast targets are excluded by the origin gate: a
fan-out target's tenant is not the origin's, and a wrong scope is worse
than none.
Tests: restart-shaped positive (scoped non-home origin -> scope_id on
routed metadata, RED before this fix) and legacy negative (scope-less
origin stamps nothing).
test_flat_key_wins_over_subblock asserted the OPPOSITE of its name (the
sub-block wins, matching _relay_slack_extra). Rename to what it proves.
Also note in _resolve_cron_surface_mode why its fallback differs from
_relay_slack_extra's all-or-nothing sub-dict: the flat key is the legacy
staging shape, and a flat knob applies to every fronted platform, gated
only by the per-platform D6 capability check.
_cron_mirror_delivery_enabled still promised 'cron deliveries live only
in the cron job's own session' as the unconditional default, but the
in_channel continuable surface now seeds the target session regardless
of attach_to_session/cron.mirror_delivery (the seed IS the continuation
feature, and in_channel is itself opt-in). State the carve-out where the
guarantee is documented.
RelayAdapter.supports_inchannel_continuable is a scalar adopted from the
PRIMARY identity's handshake descriptor, but one RelayAdapter fronts N
platforms and the connector advertises the bit per platform. Reading the
scalar for every logical platform both leaked a Slack-primary True onto
other fronted platforms (activating the flat surface their descriptor
never advertised) and suppressed a non-primary platform's advertised
True (forcing thread mode on capable Slack behind a Discord primary).
Add supports_inchannel_continuable_for_platform(platform): resolves the
platform's own negotiated descriptor via descriptor_for_platform (the
same Phase 1.5 seam max_message_length uses), scalar fallback only when
the per-platform descriptor is unavailable. The scheduler's D6 gate
prefers the query when the adapter provides it; native adapters keep
the class-attribute path byte-identically.
Tests: two-platform descriptor matrix (primary-True no-leak,
non-primary-True honored, unknown-platform scalar fallback).
The seed was decoupled from the mirror opt-in (in_channel is the
continuation surface regardless of attach_to_session), but the
thread-id-clearing gate above it still read mirror_this_target. With the
advertised default config (attach_to_session=false, cron.mirror_delivery
unset) and an origin carrying a real thread_id, the brief kept delivering
INTO the origin thread while the flat (thread_id=None) session got
seeded — brief and continuation surface in different places, so a plain
reply never saw it.
Flatten on the same gate as the seed: origin_target (with the existing
live_adapter_ready guard). Fan-out/broadcast targets are unaffected.
Test drives _deliver_result with a thread-carrying origin and default
knobs, asserting on the routed DeliveryTarget.thread_id — RED on the old
gate, GREEN now.
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.
Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.
Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
Live incident (Alice canary 2026-08-20, job 8e21a957b77b): the continuable
thread seed created its session with chat_type='thread', but a Slack DM
in-thread reply arrives chat_type='dm' and build_session_key routes DM
threads through the DM arm (...:dm:<chat>:<thread>). Seed row and reply row
never matched — the reply had no brief in context (continuation amnesia).
is_dm on _seed_cron_thread_session selects the seeded chat_type at both call
sites (opened-thread and the companion in_channel thread seed); channel
threads are unchanged. Sibling lane of the flat seed's is_dm fix. Tests pin
the key-equality contract: seeded key == the key the reply builds.
Continuability is explicit via attach_to_session or cron.mirror_delivery;
those knobs control transcript mirroring for the ordinary thread/default
surface. However, the in_channel surface must still receive its delivery
text to seed the continuation session. Previously mirror_text was populated
only when the optional mirror knob was enabled, so in_channel jobs created
with the default false settings passed an empty string to the seed helper,
which returned False. The live symptom was a delivered cron message with no
continuation context; Alice reproduced it three times (latest ef7bd2869d15).
Keep cleaned delivery text available for continuable surface seeding while
retaining mirror_enabled for the separate _maybe_mirror_cron_delivery path.
Targeted cron/in_channel regression suite: 1008 passed, 1 skipped.
Two live failures from the Alice canary (2026-08-19, jobs 28a24afebd81 /
83b93f8be379), both leaving a continuable in_channel cron with amnesia:
1. Seed mirrored via origin heuristics and silently dropped the brief.
_seed_cron_channel_session created the flat session row, then
mirror_to_session RE-DISCOVERED the target via find_session_by_origin —
whose multi-candidate bail-out returns None on a populated chat (flat
session + N per-message thread sessions sharing one chat_id, mixed
user_ids). Receipt: 'in_channel seed did NOT land on slack:D0BJTDCSR7C'.
mirror_to_session now accepts an explicit session_id and both cron seeds
pass the exact row they just created; origin-scan remains the fallback
for callers that genuinely don't know the target.
2. The brief's OWN THREAD was never seeded. in_channel delivers flat, but
a flat Slack message still invites a thread reply (the natural mobile
affordance — exactly what the user did). That reply keys to
(chat, thread=<brief ts>), which no seed touched. The delivery's
message_id now anchors a companion _seed_cron_thread_session so BOTH
reply surfaces (plain channel message AND in-thread reply) continue the
job.
Also: thread-seed failures upgraded debug→WARNING (silent seed failure IS
the user-facing bug), and the thread seed reports landed/not-landed.
Regression tests drive both against the live failure shapes: exact-session
mirror asserted via session_id kwarg; thread companion asserted via the
SendResult message_id anchor. Clean-fixture blind spot noted: the E2E
harness used a fresh store with one row, which is why heuristic rediscovery
looked fine pre-production.
Seed failure was logger.debug — invisible in production while being the
exact 'agent has no idea about its own brief' symptom. WARNING on: seed
exception (with reason), seed returning False, and in_channel delivery to
a non-origin target.
Live regression (Alice, 2026-08-19): a continuable cron with
cron_continuable_surface=in_channel delivered its brief flat, but the
flat-session seed was gated on mirror_this_target = mirror_enabled AND
origin-match. Without attach_to_session (and with cron.mirror_delivery
defaulting False) the seed never ran; the next plain reply resolved to a
blank (slack, chat, None) session and the agent had no idea about its own
delivery message — not continuable in channel OR thread (in_channel mode
correctly skips thread creation, so there was no thread session either).
in_channel IS the continuation surface, not a mirror nicety: gate the seed
on origin-match alone, and resolve origin_user_id for any origin-matching
target so the seeded key still carries the scheduling user on
per-user-isolated chats. attach_to_session remains the opt-in for the
separate default-surface mirror behavior.
Regression test drives the delivery path with attach_to_session=False and
asserts the seed fires with the right user_id (fails on pre-fix code).
Field report (enterprise side-by-side, 2026-08-18, finding 1 — the relay-only
blocker): on relay-fronted Slack, cron briefs always deliver into a dedicated
thread; the flat continuable surface (cron_continuable_surface: in_channel)
that native Slack supports is inert, so plain DM replies never continue the
job and the main conversation never sees the brief.
Three gaps closed:
- CapabilityDescriptor gains supports_inchannel_continuable (default False,
additive within contract_version 1; from_json ignores it from old
connectors, old gateways filter it as unknown). The connector advertises
it per platform at handshake.
- RelayAdapter maps the bit onto the adapter capability surface in both the
constructor and _apply_descriptor (renegotiation), so the scheduler's D6
fail-safe gate sees it exactly like native Slack's class attribute.
- _resolve_cron_surface_mode replaces the scheduler's inline flat-key read:
native keeps the shipped flat shape; the relay lane reads the same
per-logical-platform sub-block as the documented relay Slack knobs
(platforms.relay.extra.slack.cron_continuable_surface), sub-block wins,
scoped so a slack block cannot leak onto other fronted platforms.
The seed path needs no changes: RelayAdapter inherits set_session_store
(wired by the generic adapter boot loop) and _seed_cron_channel_session
keys the flat session off the logical platform_name.
12 new tests: descriptor default/from_json/legacy-absence, adapter mapping
constructor + renegotiation, and the surface-knob matrix (native flat key,
relay sub-block, per-platform scoping, precedence, defaults).
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):
- Promote the media-send timeout to the standard resolution pattern:
HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
cron.media_send_timeout_seconds in config.yaml, then 300s default
(mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
(environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
in delivery_errors (post-#88631 the reason reaches the run status, not
just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
The media delivery path used a hardcoded future.result(timeout=30).
Large attachments legitimately exceed it with no way to raise the limit.
Read HERMES_CRON_MEDIA_SEND_TIMEOUT, matching the existing
HERMES_CRON_SCRIPT_TIMEOUT / HERMES_CRON_TIMEOUT /
HERMES_CRON_SESSION_DB_TIMEOUT convention in the same module.
TimeoutError carries no message and str(TimeoutError()) is the empty
string, so the media-send warning rendered with nothing after the colon.
Fall back to the exception class name when str(e) is empty.
Field report (enterprise, v0.20.0): cron jobs delivering text + PDF/image
attachments to Slack DMs deliver both on scheduled ticks but text-only on
manual `hermes cron run <job-id>`. Same box, same token, same scopes —
the divergence is process context and error visibility, not credentials.
Three defects, one bug class (attachment failures invisible + policy
divergence between the gateway process and standalone processes):
1. Standalone lane swallowed warnings: platform standalone senders
(Slack files_upload_v2, Discord, ...) report per-file upload failures
in result['warnings'] while returning success=True for the delivered
text leg. _deliver_result only read result['error'], so the run was
marked ok and the attachment vanished without a trace. Warnings now
surface into delivery_errors (and the job's last_error).
2. Live-adapter lane swallowed media failures: _send_media_via_adapter
logged failures at WARNING and returned None. It now returns per-file
error strings and _deliver_result records them — text-delivered-but-
attachment-failed is a visible partial failure on BOTH lanes.
3. Media-policy env bridge was gateway-only: gateway.strict /
media_delivery_allow_dirs / trust_recent_files were translated from
config.yaml to the env vars validate_media_delivery_path reads ONLY in
gateway startup. A CLI-process manual run filtered attachment paths
under a different policy — in strict/allowlisted deployments the exact
reported symptom (scheduled delivers, manual drops, silently). The
translation now lives in gateway/media_policy.apply_media_policy_env
(idempotent, env-wins, never raises); gateway startup delegates to it
and _deliver_result applies it before filtering. Attachments dropped
by the policy filter are also reported in the run status instead of
only a stderr WARNING.
On v0.20.0 specifically the failure was double-blind: the pre-9cf2cbd382
isinstance(resp, dict) gates meant upload failures were undetectable in
the sender AND unsurfaced by the scheduler. 9cf2cbd382 (in 2026.8.13)
fixed detection; this fixes visibility and policy parity.
8 new tests (tests/cron/test_media_delivery_parity.py): warnings→errors,
clean-delivery control, media-reaches-sender control, live-adapter
failure/dropped-path reporting, bridge helper semantics, strict+allowlist
end-to-end in a non-gateway process, and the .env-strict/config-allowlist
split that reproduces the field symptom. Mutation check: disabling the
warnings loop and the bridge fails exactly the 2 guarding tests.
When an external scheduler (Chronos on hosted deployments) cannot
deliver a fire — dead loopback hop at fire time, retry budget exhausted
— the job's next_run_at stays parked in the past and nothing ever runs
it: external providers have no local tick loop, so the day is silently
lost even if the gateway heals minutes later (4 consecutive nightly
misses in the field).
fire_overdue_jobs() in cron/scheduler_provider.py, called from the
gateway housekeeping loop every 5 minutes:
- No-op for the built-in ticker (its tick loop already self-heals
past-due jobs) and when cron.misfire_grace_minutes <= 0.
- Waits out a grace window (default 10 min) so the external scheduler's
own retry backoff gets first right to deliver.
- Claims via the provider's claim_fire (store CAS — a concurrent late
external retry is de-duplicated) and runs fire_claimed in a daemon
thread, mirroring the webhook admission pattern, so housekeeping
never blocks for the length of an agent run. Provider re-arm logic
(Chronos NAS one-shots) runs exactly as for a normal fire.
Docs: cron.md section + cron.misfire_grace_minutes reference.
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
The PR's guard used `(job.get('schedule') or {}).get('kind')` which
crashes with AttributeError when schedule is a raw string (e.g.
'every 5m'), as happens in test_parallel_pool.py fixtures and any
job created via create_job(schedule='every 1h'). Use the
isinstance guard pattern already used at lines 5181 and 5325.
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.
Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.
9/9 tests green; ruff clean.
Follow-ups on the #87591 salvage:
- cron/scheduler.py: wrap the three clear_run_claim call sites in a
best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
and on the interpreter-shutdown path (or with a corrupt store) it could
itself raise, defeating the skip-cleanly purpose of these early exits.
A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
clear_run_claim unit contract (one-shot cleared / already-clear noop /
recurring never touched / unknown id), all three dispatch-failure paths
through a real tick() clear the claim, and a raising clear_run_claim
does not crash the tick. Mutation-verified: reverting the fix makes the
suite fail.
get_due_jobs() stamps a run_claim on one-shot jobs before returning
them as due, and mark_job_run() clears it on successful completion.
When dispatch itself fails (interpreter shutdown, executor submit
error, execution-creation error) the job never reaches mark_job_run
and the stale claim blocks re-dispatch until the TTL expires
(default 30 min).
Add clear_run_claim() to jobs.py and call it on every early-exit
path in _submit_with_guard so the job stays due and fires on the
next healthy tick — matching the existing scheduler comment's
promise.
Fixes#86522
/simplify-code findings on the salvage stack:
- efficiency HIGH: latest_executions() ran a SQLite connect + DDL + query
every tick for the whole duration of ANY running job, even when every
claim had a live future and the result was never consulted. Two-phase
now: snapshot (job_id, future) under _running_lock, query the ledger
only for claims whose future is missing/pending/done — the healthy
steady state pays zero DB work per tick.
- quality: inline ("completed", "failed", "unknown") tuple duplicated
cron/executions._TERMINAL_STATES (drift risk) — import the constant.
- reuse: hand-rolled naive-timestamp normalization in _row_belongs_to_claim
duplicated cron.jobs._ensure_aware's legacy-naive policy — reuse it.
- quality: dropped the tautological 'if fut is None or pending or done'
re-check (control only reaches it after the live-future continue) and
collapsed the two copy-pasted release blocks into one with a computed
reason.
24/24 tests green; mutation check re-verified on the final stack
(defeating the ownership guard fails exactly the 2 race-guard tests).
Follow-ups on the #87259 salvage:
- cron/scheduler.py: the ledger-terminal reconciliation now requires the
terminal execution row's claimed_at to be >= the in-memory claim's
registration time (_running_since). Without this, the latest terminal
row for a recurring job is usually the PREVIOUS run's outcome — a fresh
claim in the try_register_running_job -> create_execution window (or a
finished run whose worker finally block hasn't released yet) would be
force-released and the job double-dispatched. Unparseable/missing
claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
query under _running_lock — list() over a set concurrently mutated by
try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
two new race-guard tests (previous-run terminal row never releases a
fresh claim; missing claimed_at fails closed). Mutation-verified:
removing the ownership guard fails both.
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.
sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).
Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
loops (_start and _start_multiplex) along with duplicated function-local
imports — extracted _note_tick_failure() next to _backoff_wait_seconds
so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
_is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
_is_fd_exhaustion_text() out and use it from both.
11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
Follow-ups on the #87796 salvage:
- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
lock-failure raise site — the ticker loop's except handler already runs
reclamation once per failed tick, so the raise-site call doubled the
gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
duplicated verbatim in start() and _start_multiplex() into a module-level
_backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
cleanly (exit 1) instead of dumping a traceback — tick() raising on real
lock-acquisition failures is new behavior from this fix.
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.
- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
/simplify-code findings on the salvage stack:
- reuse HIGH: _compute_grace_seconds duplicated the exact croniter
two-fire period measurement _schedule_cadence_seconds implements
(interval minutes*60 branch included) — grace is now derived from the
shared helper, so cadence is measured in exactly one place (and grace
computations now benefit from the per-expr cache too).
- efficiency: _cron_cadence_cache was unbounded in principle (deleted/
edited exprs never evicted) — hard 256-entry bound with full clear;
rebuild cost is two croniter evals per live expr.
80 recovery/rearm/jobs tests + 94 scheduler tests green; ruff clean.
Follow-ups on the #87261 salvage:
- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
Re-arming to `now` fired CRON jobs at times their expression excludes —
a weekday-only 9am job whose Friday run errored would fire on SATURDAY
(croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
— the next LEGAL occurrence — and only when that actually moves
next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
_jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
errored Friday re-arms to Monday (not Saturday), correctly-parked cron
value untouched, interval job still due immediately.
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.
cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.
Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
When check_gateway_lifecycle refuses a cron script that lives on a
FileProvider path, the generic error implied the job contained a dangerous
gateway lifecycle command. Surface the real reason instead — the script
lives on a cloud-synced path whose evicted placeholder could hang the
preflight scan — while staying fail-closed. Regression test asserts the
cron-script scan path blocks without opening the file and that the message
names the cloud-synced path rather than a lifecycle command.
Widen #88052 per review:
- The walk-level short-circuit only protected _contains_unsafe_gateway_action;
the sibling caller _read_script_for_scanning still opened cloud-resident cron
scripts and could hang preflight. Move the check into _read_referenced_script,
the shared choke point, so every caller fails closed without opening.
- Generalize _is_apple_file_provider_path -> _is_cloud_placeholder_path: detect
~/Library/CloudStorage (Dropbox/OneDrive/Google Drive third-party FileProvider
domains) alongside iCloud's Library/Mobile Documents.
- Regression tests: CloudStorage lexical path blocked without open; the choke
point itself refuses cloud paths with os.open forbidden.
"\x00" in script_path raises TypeError when a caller passes a non-str
(e.g. a pathlib.Path, which is not iterable) — the guard itself would
crash the scheduler. All current call sites pass plain str, but the
guard must be crash-proof: str() first, then check. Adds a regression
test running a real script through _run_job_script with a Path argument,
which fails with TypeError on the pre-fix guard.
_run_job_script wrapped only expanduser() in its ingestion try/except.
On Linux an unexpandable NUL-bearing value raises inside that call, so it
landed on the clean fail-with-report path; on Windows expanduser() never
expands '~user' (and so never raises), and the NUL surfaces later as an
uncaught ValueError from resolve()/exists() — crashing the scheduler.
Align with cron.lifecycle_guard._expand_candidate_path, which already
documents this as the whole-class fix (the per-syscall catching produced
#76762, #77703, #77780, #78256): reject '\\x00' eagerly at the ingestion
boundary so both platforms fail identically.
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.
Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.
Fixes#86661
- WARN when the venv site-packages layout is unresolvable and the script
falls back to plain PYTHONPATH execution, so 'editable installs
invisible' failures are diagnosable.
- Docstring: note that runpy does not set __package__/__spec__ the way a
direct python script.py invocation does.
_windows_cron_python_invocation bypasses the uv venv launcher (to avoid
flashing a console window) and re-attaches the venv via PYTHONPATH — but
PYTHONPATH entries are plain sys.path additions and never get .pth
processing, so editable installs (pip install -e) were invisible to cron
script jobs (ModuleNotFoundError).
Bootstrap the script with site.addsitedir() on the venv site-packages,
then exec it as __main__ via runpy.run_path, preserving the script
directory on sys.path (python script.py semantics). Falls back to a
plain invocation when the venv layout is unresolvable.
Docker Desktop writes fpath=(~/.docker/completions ...) into .zshrc.
The referenced-script walk then opened that directory, saw a non-regular
file, and fail-closed — blocking source ~/.zshrc on every terminal
command. Directories are not scripts; devices stay fail-closed.
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.
Two-part fix:
1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
channel stateless (scoped ContextVar set/reset around the call) before
invoking the cron API, so `async_delivery_supported()` gates off
`_try_dispatch_background_run` and the run executes synchronously to
completion in the CLI process — the same behavior `hermes -z` already
gets via declare_stateless_channel().
2. cron/scheduler.py: tick() now periodically invokes
recover_interrupted_executions() (previously only run at scheduler
startup), so execution rows whose exact owner process is provably dead
(pid + process start time check in _owner_is_live) are reaped to
'unknown' by the long-lived gateway ticker without a restart.
Throttled to once per 300s so idle 60s ticks don't pay a ledger
connection every cycle.
Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.
Fixes#86721
The errno literal set {8, 7, 11, 51, 60, 61, 65} mixed macOS getaddrinfo
constants with errno values: on Linux socket.EAI_NONAME is -2 and
EAI_AGAIN is -3, so genuine DNS failures were missed while unrelated
OSErrors carrying errno 8/11 (ENOEXEC/EAGAIN semantics differ) could be
misclassified as transient. Classify socket.gaierror against the EAI_*
constants and plain OSError against named errno constants instead.
Addresses the platform-portability review on #83977.
Agent crons resolve OAuth credentials before the agent loop. A short
macOS/WARP DNS blip raised httpx.ConnectError ([Errno 8] nodename nor
servname provided) from xai-oauth token refresh, and the scheduler only
walked fallback_providers on AuthError — so Daily Focus Kickoff died
even when XAI_API_KEY / Anthropic were healthy.
Treat ConnectError/DNS OSError (and cause-chain equivalents) like
AuthError when selecting the fallback chain. Keep provider+model atomic.
Regression test covers the ConnectError path.