Review findings (Salt, NS-788):
B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.
S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.
S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.
Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.
Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).
Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.
Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.
An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.
Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.
Co-authored-by: webtecnica <webtecnica@gmail.com>
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.
Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.
Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
The create-path coercion (salvaged from #78928) left update_job and the
cronjob tool's update handler comparing/storing raw strings: repeat=
'forever' via update raised TypeError in the tool path and stored the raw
string via update_job, breaking the next mark_job_run ('str' has no
.get). Extract normalize_repeat_value as the shared chokepoint (shape
from #77366 by @andrexibiza, with garbage-rejection semantics) and route
create_job, update_job, and the tool update handler through it.
Completed counters are preserved across repeat updates.
Class: #66824#64520#7142#71987#95706, update half of #77366.
Contract bug (2026-08-04): the cronjob tool schema documents '30m' as
'(every 30 minutes)' — recurring — but parse_schedule returned
kind='once' for bare durations, silently creating a one-shot job for a
recurring request (agent passed '30m' for 'every 30 min', job ran once
and died). Bare durations ('30m','2h','1d') now parse as recurring
intervals matching the documented contract; explicit one-shot by
duration is 'in 30m'/'in 2h' (fires once that far from now). ISO
timestamps stay one-shot.
Also fixes the sibling repeat-coercion class (#66824/#64520/#7142):
repeat='forever'/'once'/'N' strings now coerce in create_job instead of
raising "'<=' not supported between instances of 'str' and 'int'".
Tool description rewritten to teach the corrected contract and steer
relative requests to 'in Nm' (no more hand-computed ISO timestamps).
Supersedes the doc-only direction of #53739 while keeping its goal
(relative one-shots must be expressible) via the 'in X' form.
Signed-off-by: andrexibiza <84248988+andrexibiza@users.noreply.github.com>
parse_schedule's "every " branch passed everything after the prefix
straight to parse_duration(), so documented natural-language schedules
like "every monday 9am" and "every day at 9am" (AGENTS.md, SKILL.md,
cron docs) were rejected with "Invalid duration". Convert weekday and
daily/weekday/weekend phrases to cron expressions before the duration
fallback; "every 30m"/"every 2h" interval parsing is unchanged.
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
The public schema and job store already support per-job
attach_to_session, but the registry adapter dropped the argument.
Create silently omitted the field; update reported "No updates provided."
Fixes#84802
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
Simplify-pass follow-ups on the #87033 fix:
- _gateway_liveness_notice(plural=) authors both wording variants at one
site; removes the exact-substring .replace() that would silently no-op
if the create-path text is ever edited.
- Collapse the operator-precedence-trap conditional in list to a plain
'if jobs' — an empty list has nothing inert and now skips the probe.
- Fix docstring/code mismatch (builder returns gateway_running: True on
the happy path) and drop the dead try/except in
_warn_if_gateway_not_running (the helper never raises).
Follow-ups for the salvaged #93098:
- Move the tri-state liveness heuristic into hermes_cli.cron
(_builtin_gateway_liveness) so the CLI warning and the cronjob tool
share one implementation instead of two drifting copies.
- Surface gateway_running/warning on the list action too — an agent
inspecting jobs in a gateway-less environment has the same silent-
inert-job failure mode (#87033) as create. Empty lists stay quiet.
The builtin cron ticker only runs inside the gateway process. The CLI
surfaces this ('hermes cron list' / 'hermes cron status' both warn when
no gateway is running), but the model-facing cronjob tool returned a
clean success on create even with no gateway running - so the agent
confidently told the user a recurring task was scheduled while the job
could never fire.
Mirror the CLI's liveness heuristic in the tool's create path and attach
a tri-state gateway_running field to the result:
- true -> gateway running (or a non-builtin scheduler provider owns
firing, e.g. Chronos, which is exempt by design)
- false -> explicit warning telling the model the job is saved but will
NOT fire until the gateway starts, so it can relay that to
the user instead of reporting unqualified success
- null -> probe failed; claim neither way
Fixes#87033
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
The CLI (hermes cron create/edit) routes through cronjob(); removing the
parameter outright broke that lane (CI slices 6/9). The parameter is back
on the function, but CRONJOB_SCHEMA and the registry handler still omit
it — same pattern as the intentional model/provider/base_url omission.
New test proves a hallucinated reasoning_effort arg through the model
dispatch is dropped.
Standing policy: models do not make model-configuration decisions (the
only exception is user-defined profile selection in Bot Mode/kanban).
The per-job reasoning pin stays fully functional via
`hermes cron create/edit --reasoning-effort` and the job store; the
cronjob tool still SURFACES the pin in listings but cannot set it.
A schema-absence test pins the policy.
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.
- cron/jobs.py: new optional job field, validated at the storage choke
point against the canonical grammar via the shared
hermes_constants.parse_reasoning_effort (spelling-only; capability
clamping stays owned by the provider transports at send time, same as
config-set effort). Empty string clears on update; invalid values
raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
> agent.reasoning_overrides > agent.reasoning_effort at fire time,
after the auth-fallback model swap (the pin is model-independent by
design). A stored value that no longer parses warns and falls back to
config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
(create and update), conditional key in _format_job, schema documents
grammar/precedence/transport clamping/clear semantics. Agent-settable,
unlike model/provider pins: it cannot redirect spend to a different
model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.
Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.
Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.
Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).
Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
last_fire_error ({at, detail}) on the job record; mark_job_run clears
it on the next successful run so it always describes current
auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
job on the gateway-unreachable path, best-effort (never disturbs the
503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
"Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
provider is active but the api_server adapter is not running (the
fire path is dead-on-arrival; most common cause is API_SERVER_KEY
missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
CLI parity for the continuity toggle:
- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section
E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
Per review: expose run-to-run continuity as a boolean `continuity` flag on
cronjob create/update instead of asking users to know the reserved
context_from='self' value. The flag translates to the 'self' entry in
context_from internally (create: appends/omits; update: adds or removes
'self' while preserving other upstream refs). Schema documents the flag and
steers context_from back to job-id chaining only. Docs updated; 7 new tests.
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.
- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
(can't be validated against the store — the job doesn't exist yet at
create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
Follow-up to the salvaged #86862: surface reclaim counts at warning level
(mirrors the scheduler tick's reap handling from #86853) and keep a debug
trace when the best-effort recovery itself fails, instead of a bare pass.
Fixes#86721.
`hermes cron run <job_id>` (a one-shot CLI invocation) dispatches
manual runs via the same background-delegation path as an agent's
`cronjob(action='run')` tool call (tools/cronjob_tools.py's
_try_dispatch_background_run -> dispatch_async_delegation(role=
"cron_run", runner=_runner, ...)). The runner thread lives in the
calling process's shared daemon executor. When the one-shot process
exits right after printing "Triggered job: ...", the in-flight runner
dies mid-execution, leaving its cron/executions.db row permanently
stuck at status='claimed' -- every subsequent `hermes cron run` on the
same job then reports "Ran now: failed" because of the still-claimed
row.
cron/executions.py already has the exact self-heal this needs:
recover_interrupted_executions() correctly identifies and reclassifies
'claimed'/'running' rows whose owner process has provably exited
(_owner_is_live checks PID existence AND matches process start-time,
so a reused PID isn't mistaken for the original live owner) to
'unknown', unblocking the job for a fresh claim. But it was only ever
called once, at the long-lived scheduler ticker's own startup
(cron/scheduler.py:379's self.recover_interrupted()) -- a one-shot CLI
invocation has no equivalent "startup" moment of its own, so this
self-heal never ran for it.
Added a call to recover_interrupted_executions() at the top of
_try_dispatch_background_run, right after the async-delivery-supported
gate and before any claim attempt for the current job -- mirroring
exactly what the long-lived scheduler already does at its own
startup, just triggered per one-shot invocation instead of once at
daemon startup. Wrapped in try/except: pass (best-effort; a failure
here must not block the actual dispatch this function exists for).
Traced (but did not attempt to fix) the deeper "why does the runner
die with the process at all" question -- that's the harder problem
options 1/2 in the issue describe (route to the persistent scheduler,
or block the one-shot process until completion). This fix addresses
the more urgent, more clearly-scoped symptom: a stranded stale claim
permanently blocking ALL future manual runs of the affected job, which
is option 3 from the issue and the one with an existing, already-
correct implementation just needing to be wired into this call site.
Added 3 regression tests to a new file, following the established
real-subprocess dead-owner pattern already used in
tests/cron/test_execution_ledger.py (a genuinely-dead PID, not a
mock, matching the real-world failure mode exactly): a sanity test
confirming the stale claim sits unrecovered without the fix; a direct
test of recover_interrupted_executions() reaping such a claim; and a
unit test on _try_dispatch_background_run itself confirming recovery
is called before any claim attempt. Verified as a genuine regression
by reverting the fix and confirming the unit test fails with recovery
never having been called.
35/35 pass across the new test file plus tests/cron/test_execution_ledger.py
and tests/tools/test_cronjob_run_background.py (no regression).
- The gateway api_server fire webhook acknowledges 202 only after a
durable claim + execution row exist (admission failure stays retryable
as 503; a live claim answers 200 duplicate), then dispatches the
claimed snapshot with the live runner adapters (delivery parity with
the built-in ticker, including relay-fronted and E2EE platforms).
- Legacy single-phase providers (a documented fire_due override without
split hooks) keep being driven through their own hook. Capability
detection now credits claim_fire AND fire_claimed overrides, so
Chronos is correctly classified split-aware (its re-arm lives in
fire_claimed; the redundant fire_due passthrough override is removed).
- Multi-profile dashboards fail closed for external providers: an
unscoped reconcile would disarm other profiles' armed one-shots in the
shared NAS registry.
- Manual runs (cronjob run) carry the owner-bearing claimed snapshot
through every entry point, composing with upstream's manual-run
heartbeat (#76502) and background dispatch.
Note: current main moved the dashboard NAS webhook to a pure
forward-to-gateway design (the gateway owns execution and live
adapters), so the dashboard-side claim/tracking machinery from earlier
revisions of this PR is dropped; the durable admission contract lives in
the gateway webhook path.
Bug 1: relay-fronted Slack in thread-per-message mode stamps each top-level
message's own id as source.thread_id (session KEYING, native thread_ts
parity). Cron origin capture persisted that stamp as durable routing, so
every delivery landed inside the ephemeral thread spawned around the
creation message instead of the top-level conversation. Fix at the source:
_origin_from_env drops a Slack thread id equal to the creation message's
own id (genuine in-thread creations keep theirs). Fire-time repair for
already-persisted jobs: deliver=origin and the explicit-target Slack
re-attach treat an origin thread as stale when the origin chat is the
configured Slack home chat — top-level (or the home target's configured
thread) wins; non-home working threads are preserved.
Bug 2: _preflight_check_delivery and cron_delivery_targets validated
deliver prefixes against get_connected_platforms(), which only sees
natively configured platforms — a relay-only deployment ({relay}) rejected
'slack:CHAT' with 'no gateway credentials configured' although fire-time
routing (resolve_delivery_transport + RelayAdapter.fronts_platform)
delivers it. New gateway.relay.relay_fronted_platforms() (env-derived from
GATEWAY_RELAY_PLATFORMS — the same source that seeds the live adapter's
identity set, so validation and routing cannot disagree) is unioned into
the connected set when the relay is connected. Native topologies keep the
strict credential check unchanged.
Review caught a real gap: action='update' also accepts deliver, and the
tool description explicitly steers agents toward update-over-create — so
a cron-context agent updating a job to deliver='origin' would recreate
exactly the dangling literal-origin shape the create-path resolution
prevents (stored 'origin' on an origin-less job → fire-time home-channel
guessing or silent drop).
Wrap the update site in the same resolver. Semantics follow the create
precedent: in cron context, 'origin' means 'my run's target', resolved
concretely at mutation time; outside cron context updates are
byte-identical to before.
A job created from within a cron run must never store the literal
'origin' delivery target: the creating session is ephemeral, so by fire
time there is no origin to resolve and the scheduler falls back to
guessing a home channel. With agent scheduling enabled
(cron.allow_agent_scheduling), a scheduled agent creating follow-up jobs
would silently produce exactly that dangling shape.
Resolve at create time instead, in cron context only: 'origin' elements
(and an omitted deliver) are replaced with the creating run's concrete
target from the per-run HERMES_CRON_AUTO_DELIVER_* contextvars —
platform:chat_id[:thread_id], or 'local' when the creating run has no
concrete target. Explicit values ('local', 'all', platform:chat_id
targets) pass through verbatim, including inside comma lists. Chat and
CLI creates are byte-identical to before: the resolver is a no-op
outside cron-context sessions (HERMES_CRON_SESSION unset).
Cron-spawned agents have the cronjob toolset unconditionally denied, so
scheduled agents cannot create, tune, or remove jobs even when an
operator wants exactly that (reconciler-style jobs that manage a team's
cron table, follow-up one-shots scheduled from within scheduled work).
The denial is loop-prevention policy, not a security boundary: an agent
with the terminal toolset can already shell out to the CLI, so the
workaround exists but skips every limit and accounting layer.
Add cron.allow_agent_scheduling (config.yaml, default false — byte-exact
current behavior). When enabled, only 'cronjob' leaves the cron-context
denylist; 'messaging' and 'clarify' remain denied as interactivity
constraints, and the user-level agent.disabled_toolsets layering is
unchanged, so a user denylist entry still beats the gate. The cronjob
tool description now states the real policy and the quota bounds instead
of a blanket prohibition.
Two claim-failure diagnostic paths (cronjob_tools.py:629,921) still used
the old inline 'not enabled or state==paused' check. After get_job()
normalizes via effective_job_state, a half-paused record has
state='scheduled' and enabled=True, so the inline check returned False —
mislabeling the job as 'already being fired' instead of 'paused/disabled'.
Also hoists effective_job_state/is_job_runnable to the top-level import in
cronjob_tools.py (was function-local) and updates console_engine.py's
_format_job to use effective_job_state instead of the old inline
state-or-enabled derivation — a fourth display path the original PR missed.
Follow-up to PR #81287.
pause_job already sets enabled=false atomically with state/paused_at, but
get_due_jobs only checked enabled — so a contradictory record
(enabled=true + paused_at/state=paused) still fired. That was the 07-30
outage failure mode: list looked frozen, fleet kept merging.
- is_job_runnable / effective_job_state: pause markers gate fire; display
derives from the scheduler-honoured enabled flag so half-paused never
renders as [paused]
- get_due_jobs self-heals enabled=false + logs error on contradiction
- claim_job_for_fire uses is_job_runnable (paused_at counts too)
- list/format paths use effective_job_state
- behavioural tests: pause blocks due fire; half-pause self-disables
Add monitor-mode cron jobs: a cheap monitor source (monitor_script or
monitor_url) runs on every tick BEFORE any agent machinery is built.
Its output is hashed as exact bytes and compared to the hash stored
from the last agent-triggering tick:
- unchanged -> agent run suppressed entirely (no LLM, no delivery);
the tick is recorded as a silent no_change run visible in the
executions ledger doc
- changed -> a MONITOR CHANGE DETECTED block (capped unified diff of
previous vs current output + the new output) is injected into the
prompt via the existing extra_prompt seam, then a normal agent run
- first run -> always runs the agent with a baseline block
- source failure -> delivered as an ERROR alert, never treated as a
change; the stored hash is untouched so recovery to prior output
still suppresses
Implementation:
- cron/monitor.py (new): hash/diff/URL-fetch/state persistence.
monitor_script reuses _run_job_script (same ~/.hermes/scripts/
containment + interpreter rules); monitor_url is a bounded GET
(30s, 256KB, http/https only). Output is exact bytes by design —
scripts should emit stable output (documented).
- cron/jobs.py: additive job fields monitor_script / monitor_url /
monitor_state {last_output_hash, last_changed_at}. JSON job records
need no migration. create-time validation: sources are mutually
exclusive and incompatible with no_agent=True.
- cron/scheduler.py: one tight monitor gate in run_job between the
no_agent short-circuit and the LLM path (outside sibling-lane
regions). State persists in jobs.json + a per-job snapshot file, so
suppression survives scheduler restarts.
- tools/cronjob_tools.py: additive optional monitor_script/monitor_url
params on the cronjob tool (create + update, empty string clears),
path containment validated at the API boundary, surfaced in
_format_job.
- hermes_cli: --monitor-script/--monitor-url on `hermes cron create`
and `hermes cron edit`; `hermes cron list` shows the monitor source
and last-changed time.
Tests (tests/cron/test_monitor_kind.py, TDD): unchanged suppresses,
changed injects diff, first run always runs, hash persists across
module reload (restart), script failure is error-not-change with hash
untouched, create/update validation, tool wiring + path-escape reject.
Inspired by: ChatGPT Work monitor tasks (idea-level, docs-only);
enabler: #80774
Follow-up to the salvaged registration contract:
- share one _raise_if_cron_registration_error() helper for the two
byte-identical dashboard 424 except-blocks (web_server + cron router,
via the existing late() seam)
- add endpoint-level 424 coverage for /api/cron/blueprints/instantiate
(previously only the sync worker was tested)
- give chat/CLI surfaces a human-facing user_message() (job name, no
exception class name) and add a recovery hint (pause/resume or update
re-registers via provider reconcile) to the model/REST message
- consolidate five inline provider test doubles into one ABC-subclassing
make_cron_provider conftest factory; the web_server test double now
subclasses CronScheduler so an ABC rename fails loudly
- narrow the wrapper facade to keyword-only (**kwargs) and route the
tool's partial-failure return through tool_error()
Salvaged from PR #57342 by @liuhao1024 (with the injection-scan half
from PR #57360 by @ghedeselmabot): cronjob(action='run', prompt=...)
silently discarded the prompt argument — per-run context never
reached the spawned cron session.
The prompt is now threaded as extra_prompt through the whole chain
(cronjob run action → _try_dispatch_background_run/_execute_job_now →
_run_claimed_job → run_one_job → run_job → _build_job_prompt) and
appended to the stored prompt under a '## Run Context' header for
that single fire only — never persisted to the job definition. It
passes the same strict _scan_cron_prompt injection scan as stored
prompts before firing, and works identically on the background and
sync fallback paths.
Test fakes across tests/cron/ updated to accept the new kwargs
(sibling-test blast radius from the signature change).
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
Salvaged from PR #53395 by @izumi0uu: the fire claim's 300s TTL is
routinely outlived by real cron jobs, so claim_job_for_fire alone
cannot stop a manual cronjob(action='run') from double-firing a job
the ticker (or another manual run) is still executing.
Extract the ticker's _submit_with_guard running-set check into shared
module-level helpers (try_register_running_job / release_running_job)
and register manual runs through the same set — one dedupe owner, no
drift. Manual runs also become visible to get_running_job_ids (the
gateway shutdown drain, #60432) and mark_running_jobs_interrupted,
which previously could not see them.
The background dispatch path pre-checks the running set so a mid-run
job reports 'already running' in the tool response immediately
instead of as a delayed error completion event; the authoritative
atomic check remains in _run_claimed_job on the worker.
Co-authored-by: izumi0uu <izumi0uu@gmail.com>