Review findings (Salt, NS-788):
B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.
S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.
S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.
Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
_record_timezone_migration_catchup was a line-for-line clone of
_record_persisted_error_recovery (counter bump, bounded recent list,
best-effort jsonl append). Extract _append_telemetry_record and route
both through it; one shared history cap replaces the two per-counter
constants. Also correct the "distinct from catch_up_occurrences" comment:
a migrated row that is also past its grace window increments both.
No behavior change; both recorders write the same entries to the same
files.
Upgrading from a UTC-scheduling build to one that honours the profile
timezone (Europe/Brussels) left daily cron jobs sitting in jobs.json with
pre-migration instants — e.g. next_run_at "2026-09-02T04:00:00+00:00" for
expr "0 4 * * *". _ensure_aware normalizes that to 06:00+02, which the
expression excludes, so the stale-expression guard (#93049) read it as a
direct jobs.json edit, logged exactly that, and re-anchored to tomorrow
without firing. The due occurrence disappeared with no error anywhere.
The guard only asked "is the stored instant an occurrence of the current
expr?", never "why not?" — and the two possible answers demand opposite
actions. Add _classify_stale_cron_next_run, which distinguishes them by
whether normalization itself moved the wall clock:
* expr_edit — wall clock unchanged (or the stored wall clock is
not an occurrence either): the instant is genuinely
excluded by the current expression. Re-anchor
without firing, exactly as before.
* timezone_migration — the stored value's own wall clock IS a legal
occurrence and it only left the lattice because
_ensure_aware converted it to a different offset.
Fall through and fire the overdue run once.
Because every value written by this build carries the configured offset, a
real expr edit leaves the wall clock untouched and can never be reclassified
as a migration, so the #93049 protection is intact. At-most-once is
unchanged: the fire flows through the normal due path and the usual
advance_next_run / mark_job_run re-anchor rewrites next_run_at in the
current offset, so the legacy instant is never read again. Future local
wall-clock occurrences are untouched — not-yet-due rows never reach the
guard, and the #28934 offset-repair branch still runs first for a
still-future stored wall clock.
The migration case is classified explicitly rather than retried broadly: it
logs cron.timezone_migration.catch_up with the stored and normalized
instants plus both offsets, and increments a probe-visible counter
(get_timezone_migration_catchup_stats, timezone_migration_catchups.jsonl)
kept separate from catch_up_occurrences so an operator can tell "the upgrade
backlog is draining" from "runs are missing their grace window".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.
Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.
Fixes#83993
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).
- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
scheduled_at, dispatched_at, lateness_seconds, and kind
(on_time / late / catch_up, classified against the ticker tolerance and
the schedule's catch-up grace window). Manual triggers and one-shots are
not stamped (no scheduled instant to be late against / retired beyond
grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
catch-up, in both the built-in ticker and external provider paths.
CLI surface only — no new tools, no policy engine.
Addresses the visibility half of #99879.
The create-path coercion (salvaged from #78928) left update_job and the
cronjob tool's update handler comparing/storing raw strings: repeat=
'forever' via update raised TypeError in the tool path and stored the raw
string via update_job, breaking the next mark_job_run ('str' has no
.get). Extract normalize_repeat_value as the shared chokepoint (shape
from #77366 by @andrexibiza, with garbage-rejection semantics) and route
create_job, update_job, and the tool update handler through it.
Completed counters are preserved across repeat updates.
Class: #66824#64520#7142#71987#95706, update half of #77366.
Contract bug (2026-08-04): the cronjob tool schema documents '30m' as
'(every 30 minutes)' — recurring — but parse_schedule returned
kind='once' for bare durations, silently creating a one-shot job for a
recurring request (agent passed '30m' for 'every 30 min', job ran once
and died). Bare durations ('30m','2h','1d') now parse as recurring
intervals matching the documented contract; explicit one-shot by
duration is 'in 30m'/'in 2h' (fires once that far from now). ISO
timestamps stay one-shot.
Also fixes the sibling repeat-coercion class (#66824/#64520/#7142):
repeat='forever'/'once'/'N' strings now coerce in create_job instead of
raising "'<=' not supported between instances of 'str' and 'int'".
Tool description rewritten to teach the corrected contract and steer
relative requests to 'in Nm' (no more hand-computed ISO timestamps).
Supersedes the doc-only direction of #53739 while keeping its goal
(relative one-shots must be expressible) via the 'in X' form.
Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
Widens _natural_every_to_cron to consume comma/'and'-separated weekday
lists ('Monday, Wednesday at 9am' -> '0 9 * * 1,3') and applies the same
helper to schedules without the 'every' prefix, matching the exact forms
the Desktop dialog advertises in the #51975 repro.
`parse_schedule()` detected cron expressions with a digit-only field pattern
(`^[\d\*\-,/]+$`), so any field using named months or weekdays — `MON`, `JAN`,
and common ranges/lists like `MON-FRI` or `MON,WED,FRI` — failed detection and
fell through to a confusing "Invalid schedule" error, even though croniter
supports them and they're standard cron.
Allow letters in the field pattern so these route to croniter for validation.
Truly-invalid expressions (`0 9 * * FUNDAY`, `99 9 * * MON`) are still rejected
there with a clear "Invalid cron expression" message; duration/interval/ISO
parsing is unchanged.
Adds tests for named weekdays/months (incl. ranges and lists) and that an
invalid named field is still rejected.
The Desktop dialog's advertised form in #51975 omits the 'every' prefix.
Reuse _natural_every_to_cron() on the bare schedule so 'weekdays at 9am',
'monday at 9:30', and 'daily at 7am' parse to cron expressions. Also
switch the salvaged branch's HAS_CRONITER check to _ensure_croniter()
(lazy-import refactor landed after the PR was cut).
parse_schedule's "every " branch passed everything after the prefix
straight to parse_duration(), so documented natural-language schedules
like "every monday 9am" and "every day at 9am" (AGENTS.md, SKILL.md,
cron docs) were rejected with "Invalid duration". Convert weekday and
daily/weekday/weekend phrases to cron expressions before the duration
fallback; "every 30m"/"every 2h" interval parsing is unchanged.
Check both resolved and unresolved path parts so a symlinked named
profile (e.g. profiles/dev -> /mnt/data/dev) is still detected.
Also use _ensure_cron_dir for output_dir in ensure_dirs() for consistency.
Simplify-code Phase 2 finding (medium severity).
Replace #96637's inline active_profile_homes() closure with #96508's
module-level _existing_profile_homes() filter (testable in isolation).
Widen _ensure_cron_dir from 3 to 12 mkdir sites across cron/ so every
directory creation fails closed for deleted named profiles, not just
the 3 originally protected. Add _is_named_profile_path() that checks
'profiles' in path parts (works for subdirs like cron/output/<job> and
scripts/ that the original parent.name heuristic couldn't reach).
Co-authored-by: misterdas <das7514@gmail.com>
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:
* state=completed: a one-shot that genuinely has no more occurrences,
ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
fails to produce a next occurrence (e.g. the croniter package is
missing at runtime). _mark_job_run_locked's own comment is explicit:
"Recurring jobs must NEVER be silently disabled" (issue #16265) — the
job is left enabled=True specifically so it keeps being a live,
recoverable job once the underlying issue resolves.
Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:
* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
its own is_terminal_job() check) never runs, because the check itself
skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
and claim steps the scheduler's own dispatch loop calls immediately
after get_due_jobs() for anything that DOES make it into the due list —
both also refuse the job, so even a manually-recovered next_run_at
would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
broken recurring job through the normal path.
The only way out was deleting the job and recreating it.
Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.
Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.
New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.
Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
Review follow-up on the #94033 salvage: guard the equality gate against a
future 'helpful' datetime normalization — any rewrite of next_run_at must
invalidate the marker, and normalizing would weaken that.
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.
- Grace gate: a once-kind job whose next_run_dt is more than
ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
record with a diagnostic file so it stops being scanned and the miss is
operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
process: skip this scan but KEEP the record so its mark_job_run can land
(avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
grace) so an explicitly re-run stale one-shot fires.
Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
Harden the id-keyed-map flatten with an id-preserving merge:
{**value, "id": value.get("id") or key} — an inline "id" wins,
otherwise the map key is adopted (external tools often key by id and
omit the inline copy; plain list(values) would emit id-less records
that collide or get dropped downstream). Non-dict junk values are
skipped with a warning instead of crashing the load. The self-heal
rewrite persists the id-merged, junk-free records.
Tests: key adopted when no inline id (and inline id wins over a
differing key), non-dict junk skipped with warning + list_jobs
survives + self-heal persists only valid records, all-junk map
flattens to [].
Layer on the load-boundary flatten: when load_jobs() encounters an
ID-keyed jobs map ({"jobs": {"<job_id>": {...}, ...}} — written by
external tools or hand edits, never by save_jobs()), it now not only
flattens to the list contract but persists the canonical
{"jobs": [...]} form back to disk via the existing auto-repair path
(save_jobs), so the store self-heals and subsequent reads are
idempotent.
Note: _peek_jobs_unlocked() intentionally does NOT tolerate the dict
shape — it returns None so the save path never shrink-merges against
an unrepaired baseline. The flatten + repair live only at the
load_jobs() boundary.
Regression tests cover the flatten, the reported list_jobs() traceback
path, idempotent on-disk repair, and the empty-map edge case.
Salvaged from PR #92994.
Co-authored-by: a-yeyang <88581400+a-yeyang@users.noreply.github.com>
get_due_jobs() fires purely off the stored next_run_at <= now, with no
check that the stored instant is still an occurrence of the schedule's
current expression. A direct jobs.json edit that narrows schedule.expr
(e.g. daily "0 7 * * *" -> weekdays "0 7 * * 1-5") keeps the stored
next_run_at computed under the old expression, so the job fires on days
the new expression excludes. The within-grace fire and the catch-up
"run once now" path both inherit the wrong instant.
Add a best-effort stale-schedule guard on the fire path: when the stored
next_run_at is not an occurrence of the current cron expression,
re-anchor it via compute_next_run() from the current expression and skip
the fire. Non-cron kinds, missing expr, croniter unavailability, and
malformed input all report a match so the fire path keeps its existing
semantics. Recomputation uses the current expression, so the re-anchor
converges and cannot defer a valid job forever.
Fixes#93049
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
get_due_jobs() stamps a run_claim on one-shot jobs before returning
them as due, and mark_job_run() clears it on successful completion.
When dispatch itself fails (interpreter shutdown, executor submit
error, execution-creation error) the job never reaches mark_job_run
and the stale claim blocks re-dispatch until the TTL expires
(default 30 min).
Add clear_run_claim() to jobs.py and call it on every early-exit
path in _submit_with_guard so the job stays due and fires on the
next healthy tick — matching the existing scheduler comment's
promise.
Fixes#86522
/simplify-code findings on the salvage stack:
- reuse HIGH: _compute_grace_seconds duplicated the exact croniter
two-fire period measurement _schedule_cadence_seconds implements
(interval minutes*60 branch included) — grace is now derived from the
shared helper, so cadence is measured in exactly one place (and grace
computations now benefit from the per-expr cache too).
- efficiency: _cron_cadence_cache was unbounded in principle (deleted/
edited exprs never evicted) — hard 256-entry bound with full clear;
rebuild cost is two croniter evals per live expr.
80 recovery/rearm/jobs tests + 94 scheduler tests green; ruff clean.
Follow-ups on the #87261 salvage:
- cron/jobs.py: the persisted-error re-arm now respects schedule legality.
Re-arming to `now` fired CRON jobs at times their expression excludes —
a weekday-only 9am job whose Friday run errored would fire on SATURDAY
(croniter measures a 24h cadence on Saturday, so 27h > cadence+grace and
the guard tripped). Cron jobs re-arm to compute_next_run(schedule, now)
— the next LEGAL occurrence — and only when that actually moves
next_run_at earlier; interval jobs (the 2026-08-14 incident class) keep
the immediate now re-arm, which is always legal for intervals.
- cron/jobs.py: cache _schedule_cadence_seconds' croniter measurement per
expr (mirrors scheduler.py's _cron_interval_cache) — it runs inside
_jobs_lock on every tick for every stale-errored job.
- tests/cron/test_persisted_error_rearm_legality.py (new): weekday job
errored Friday re-arms to Monday (not Saturday), correctly-parked cron
value untouched, interval job still due immediately.
The 2026-08-14 incident (t_20e23f84): 4 recurring no_agent interval jobs
EAGAIN-failed at 12:50 and recorded ZERO executions for ~1h47m, surviving a
gateway restart, cleared only by operator `cron resume` / force-run. The
in-memory stale-claim sweep (t_3778a491, already on origin/main) heals a
leaked `_running_job_ids` claim in-process, but a recurring job whose
PERSISTED state shows last_status=error and whose next_run_at was re-armed
into the future by mark_job_run is invisible to that sweep: it is not in the
running set and not due, so it just sits — the restart-surviving half.
cron/jobs.py::_get_due_jobs_locked now re-arms such a recurring job to
next_run_at=now when all hold: persisted last_status==error, last_run_at older
than cadence+grace (so it is a real wedge, not a normal transient-error retry),
next_run_at in the future, and not running in this process. The scheduler then
re-dispatches it on the next tick without force-run/resume. Logs
cron.persisted_error.recovered, bumps a probe-visible counter, appends a JSONL
row. Within-cadence errors are never force-re-armed.
Tests: tests/cron/test_recurring_persisted_error_recovery.py (clean behavioral
RED on unfixed main / GREEN here; 2 consecutive auto-fires; within-cadence not
re-armed). Full tests/cron/: 713 passed, 1 skipped.
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.
- cron/jobs.py: persist a failure_streak counter in mark_job_run —
incremented on agent failure, reset on success; delivery failures don't
count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
delivered failure summary once a recurring job's streak reaches
cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.
Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
- claim_job_for_fire returns the atomically claimed snapshot with a unique
fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
terminal writes by expected_fire_owner so a stale worker cannot record
over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
event (ownership loss OR external cancel) into run_job; the agent path
is interrupted cooperatively and script-based jobs (no_agent + pre-run
scripts) are hard-stopped with a process-tree kill (POSIX killpg
SIGTERM then SIGKILL for surviving group members; Windows
taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
the bare job ID, so a replacement run of the same job never consumes a
stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
completed one-shot retention (#80624), blocked_config preflight
(T1-26), and the advance_next_runs batch on top of current main.
A fleet-wide inference config change previously produced one 'Skipped to
prevent unintended spend' alert per unpinned job per tick — 40 jobs meant
40 alerts every tick until each was re-pinned (Coatue field report,
2026-08-11). The #44585 guard now reuses the #73506 alert-once shape the
preflight path already established: a persisted drift_alerted bit on the
job record, a [drift_skip:silent] marker on repeat ticks that suppresses
delivery, and the bit clears on the next successful run so a future drift
re-alerts. Only the drift branch consults the bit — every other failure
keeps alerting per tick.
The alert text also now says it is sent once, so operators know the job
stays skipped silently until pinned or restored.
4c2961c51 added referenced_skill_names() so the curator never archives a
skill a cron job depends on — paused jobs and infrequent schedules would
otherwise age their skills out and the next run fails to load them.
62972060c then taught the scheduler that jobs may store ABSOLUTE skill
paths, normalizing them through normalize_skill_lookup_name before
skill_view. The protection set kept returning the raw string, so it now
holds a full path while the curator matches it against bare skill names.
Those jobs silently lost their protection: the skill is archived, and the
next fire logs a warning and runs the job without its instructions.
Canonicalize each reference the same way the scheduler resolves it, with
a deferred import and a verbatim fallback so a resolver failure can never
drop a name (referenced_skill_names has exactly one caller, the curator's
protection lookup, so nothing else sees the change).
Four fix-forwards from the adversarial post-merge audit of the Aug 7
unreviewed merge batch:
- estop (#81148): is_engaged() now fails SAFE (engaged) on stat errors;
the gateway estop gate lets recognized slash commands and replies owned
by in-flight work (update prompts, clarify, slash-confirm, tool
approvals, running sessions) through instead of consuming them; new
gateway /pause [reason|off] command gives messaging-only operators an
in-band engage/resume path (busy_policy=dispatch so it works mid-run).
- cron monitor mode (#81138): execution-mode invariants (monitor x
no_agent, monitor_script x monitor_url, no_agent-requires-script) now
have ONE owner (_validate_job_mode_invariants) called from BOTH
create_job and update_job, so the create-time invariant can no longer
be silently violated through the update door.
- cron notepad (#81139): remove_job now clears the job's notepad rows
(clear_notepad was dead code -> orphaned KV state forever); clear is
best-effort and no-ops without creating notepad.db.
- delegation batch gate (#81141): template-marker regex narrowed to
multi-word placeholder shapes only (<feature name>, {file_path}) so
generics (Vec<T>), HTML tags, JSON snippets, glob braces and f-string
style no longer reject legitimate batches; duplicate-goal rejection
removed (best-of-N fan-outs are legitimate).
pause_job already sets enabled=false atomically with state/paused_at, but
get_due_jobs only checked enabled — so a contradictory record
(enabled=true + paused_at/state=paused) still fired. That was the 07-30
outage failure mode: list looked frozen, fleet kept merging.
- is_job_runnable / effective_job_state: pause markers gate fire; display
derives from the scheduler-honoured enabled flag so half-paused never
renders as [paused]
- get_due_jobs self-heals enabled=false + logs error on contradiction
- claim_job_for_fire uses is_job_runnable (paused_at counts too)
- list/format paths use effective_job_state
- behavioural tests: pause blocks due fire; half-pause self-disables
Validate a job's configuration BEFORE any agent machinery is constructed:
- missing provider API key (AuthError from a read-only
resolve_runtime_provider probe; skipped when a fallback_providers chain
is configured, since auth-fallback may rescue the run)
- attached skill not ready (skill_view readiness_status=setup_needed —
missing required env vars / commands / credential files)
- delivery platform unknown or unconnected (deliver=local/origin/all are
never checked; gateway-config load failures fail open)
On a failing check run_job returns a [blocked_config]-marked error without
constructing AIAgent/MCP/etc, so a misconfigured job never burns an LLM
call. run_one_job records last_status='blocked_config' and delivers the
alert exactly ONCE across ticks (persisted preflight_alerted bit — the
alert-once shape from the #73506 dead-pin auto-pause); the next healthy
run clears the marker so a future break re-alerts. Every preflight check
fails open: only an affirmative misconfiguration verdict blocks.
Config: cron.preflight (default true); `cron.preflight: false` restores
the old fail-during-run behavior. Documented in the cron user guide and
config defaults.
mark_job_run gains an optional status= override (unblocked call shape
unchanged) and drops preflight_alerted on any successful run.
Tests: tests/cron/test_preflight_config.py (blocked_config + no agent +
single alert across two ticks, healthy job unaffected, recovery clears
dedup, fallback-chain rescue, opt-out restores old behavior, skill
readiness miss, unknown delivery platform, deliver=local never loads
gateway config). Full tests/cron/ + cronjob tool suite green (525 tests).
Ported from: paperclipai/paperclip execution-semantics §5 (MIT);
in-repo precedent: #27948, #73506