cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:
* state=completed: a one-shot that genuinely has no more occurrences,
ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
fails to produce a next occurrence (e.g. the croniter package is
missing at runtime). _mark_job_run_locked's own comment is explicit:
"Recurring jobs must NEVER be silently disabled" (issue #16265) — the
job is left enabled=True specifically so it keeps being a live,
recoverable job once the underlying issue resolves.
Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:
* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
its own is_terminal_job() check) never runs, because the check itself
skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
and claim steps the scheduler's own dispatch loop calls immediately
after get_due_jobs() for anything that DOES make it into the due list —
both also refuse the job, so even a manually-recovered next_run_at
would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
broken recurring job through the normal path.
The only way out was deleting the job and recreating it.
Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.
Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.
New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.
Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Review follow-up on the #94033 salvage: guard the equality gate against a
future 'helpful' datetime normalization — any rewrite of next_run_at must
invalidate the marker, and normalizing would weaken that.
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
extract it next to _get_script_timeout/_get_media_send_timeout. Using
load_config() (deep-merge) also removes the default-drift hazard flagged
in review: cron.session_db_timeout_seconds now resolves from
DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
10.0 staying in sync with it, and drops the distant-state coupling to
run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
new position the store init happens inside the guarded worker, not before
it).
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.
Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.
Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
- _target_mirror_eligible accepts a precomputed origin_match so the sole
production caller stops re-resolving origin + re-running the origin
match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
precedence (standalone correctness: per-job False must beat raw global
True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
'not inchannel_continuable' and reword it for the widened seed scope.
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).
The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.
Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
gate: origin unchanged; origin_fallback eligible under the same flags
as origin (per-job attach_to_session wins, else global
cron.mirror_delivery); explicit platform:chat targets eligible ONLY
under per-job attach_to_session — the global flag never activates
them, so it cannot start writing transcript entries into arbitrary
explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
session keys are user-isolated, so a seed without a user_id (origin-
less job into a shared channel) would create an orphan session no
reply resolves to — those targets fall back to the plain mirror. DM
targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.
Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.
15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.
Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).
Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
- Make repair_explicit_computer_use_media_paths fail-open internally
(cosmetic repair must never abort delivery); drop the cron-only
try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
gateway.media_repair seam; pre-existing tests keep pinning the
gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
Follow-up to the salvaged fix from PR #94439:
- Extract the repair into gateway/media_repair.py (shared module) and
re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
background tasks (_run_background_task_inner) and cron job delivery
(cron/scheduler.py) — both call agent.run_conversation directly and
never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
content first instead of regex-scanning the raw string, which yielded a
doubled-backslash path artifact and rewrote the response to a path the
model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
the compression-fallback last-user slice (incl. no-user fail-closed).
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
The hosted-provider misfire catch-up (fire_overdue_jobs) fired any runnable
overdue job with no one-shot grace check, so a stored past-due one-shot
bypassed ONESHOT_GRACE_SECONDS and executed arbitrarily late after downtime.
Sibling site of the due-scan gate from #89571; pins both directions with
tests.
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.
- Grace gate: a once-kind job whose next_run_dt is more than
ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
record with a diagnostic file so it stops being scanned and the miss is
operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
process: skip this scan but KEEP the record so its mark_job_run can land
(avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
grace) so an explicitly re-run stale one-shot fires.
Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
Review follow-up on #84203. Both points reproduce; neither was a regression
from the first pass, but both are live bypasses of the same guard.
**Privilege and namespace wrappers were missing.** The allowlist covered the
coreutils-shaped wrappers but not the privilege ones, so each of these ran a
lifecycle script straight past the walk:
pkexec bash ~/restart.sh
runuser -u root -- bash ~/restart.sh
setpriv --reuid=0 -- bash ~/restart.sh
systemd-run --scope bash ~/restart.sh
nsenter --target 1 --mount bash ~/restart.sh
unshare -r bash ~/restart.sh
Added `pkexec`, `su`, `runuser`, `setpriv`, `systemd-run`, `nsenter` and
`unshare`, each with the value-taking options that would otherwise be
mistaken for the command (`nsenter -t 1`, `systemd-run -p X=1`,
`runuser -u root`, …).
**An option can carry a command STRING, not an argv tail.** `env -S` and
`su`/`runuser` `-c` take shell source. The peel treated the operand as an
opaque value and skipped it, so `env -S 'bash ~/restart.sh'` was never
scanned — the string went unread rather than being recursed into.
`_STRING_COMMAND_OPTIONS` now names those options and their values are
re-scanned as shell source, the same treatment `sh -c` payloads already get.
They are read at the ORIGINAL command token, before the transparent-prefix
peel, because peeling past `su`/`env` would discard the very option carrying
the command. `--opt value` and `--opt=value` are both handled.
Scope, stated plainly: this is an enumerated allowlist, not a general
solution to "wrapper that execs its tail". A wrapper outside the set, or a
value-taking option outside these tables, still resolves to no reference —
that fails open, exactly as it did before this PR, and it is a miss rather
than a false block. The reviewer offered "extend the set with tests, or
document that the list is heuristic"; this does the first and states the
second.
Tests: 23 new cases (220 in the file) — every added wrapper against a script
reference including the value-operand option forms, both command-string
option spellings for env/su/runuser, and the same wrappers around ordinary
work (`pkexec systemctl status nginx`, `su -c 'ls -la'`, `env -S 'echo hi'`,
`nsenter -t 1 -m ps aux`) which must stay allowed. 15 fail on the tree
before this commit.
False positives re-checked at scale: the 9,258 command lines from this
repo's own scripts and docs give an identical verdict set before and after —
0 new false positives, 0 lost detections, 0 exceptions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
`_mask_data_sink_arguments` exempts lifecycle text living in the arguments of
executables that cannot run them (`grep`, `rg`, `journalctl`, `sqlite3`, …),
so hunting for a restart string in logs is diagnostics rather than a command.
The exemption is dropped when an argument looks like an escape back into
execution — including anything starting with a dot, because sqlite3 spells
its escapes as dot-commands (`.shell`, `.system`).
But `.`, `./x` and `../x` are ordinary path operands, and
grep -r 'systemctl restart hermes-gateway' .
is the most ordinary recursive search there is. The leading-dot test treated
its `.` operand as a sqlite3 escape, disabled masking for the whole segment,
and blocked the command outright — the exact false-positive class the
exemption exists to prevent, on the shape most likely to hit it. Searching a
relative subdirectory (`./logs`, `../archive`) fails the same way, as does a
relative sqlite3 database path (`sqlite3 ./stats.db "SELECT ..."`).
Require a dot followed by a NAME character (`^\.[A-Za-z]`) so a dot-command
still defeats the exemption while a relative path stays a path. A dotfile
operand (`.env`) still reads as a dot-command — conservative, and unchanged
from today's behavior.
This narrows a security guard in the permissive direction, so the escape
hatches are pinned explicitly: with a relative-path operand present,
`.shell`/`.system`, psql's `\!`, a pipe into `sh`/`bash`/`sudo sh`/`xargs`,
command substitution, and a `;`/`&&` continuation all still block. Only the
segment's own data arguments are masked, and only when nothing in it can
reach execution.
Tests: 18 new cases in tests/hermes_cli/test_gateway_restart_loop.py — the
relative-path shapes that must now be allowed, plus the ten escape-hatch
shapes that must still block. The allow cases fail on the unfixed tree.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
`sudo`, `env`, `nohup`, `timeout` and friends exec their argument tail, so
the command that actually runs sits further right. Three guards read only the
first token of a segment, saw the wrapper, and never inspected what it runs:
bash ~/restart.sh → blocked
sudo bash ~/restart.sh → allowed
launchctl submit -l com.x -- helper → blocked
sudo launchctl submit -l com.x -- helper → allowed
Same foot-gun, one word of prefix. That reaches both enforcement points —
`cron.jobs.create_job` and `tools/terminal_tool.py` under `_HERMES_GATEWAY=1`
— and defeats the label-independent submit block that #62891 added precisely
because a persistent helper is the indirect route to a restart loop.
`_peel_transparent_prefixes()` walks past a bounded chain of these wrappers,
skipping their own options, their value-taking options (`sudo -u deploy`,
`stdbuf -o0`), `VAR=value` assignments, a `--` end-of-options separator, and
`timeout`'s duration operand, then returns the index of the real command. It
is applied to the referenced-script walk, the `sh -c` payload walk, and the
`launchctl submit`/`bootstrap` block.
In the referenced-script walk the peel is ADDITIVE — the segment is read at
the original token and again at the peeled one — because peeling must never
remove a reference the un-peeled read would have found. A local script named
`./timeout` is a script, not the coreutils wrapper, and consuming it as a
prefix would have silently stopped scanning it. (The other two call sites
need no such care: no wrapper name is also a shell name or `launchctl`, so
peeling there can only add.) That split is why the per-index logic now lives
in `_references_at()`.
This is not a new reading of shell syntax for this module — `_PIPE_TO_INTERPRETER`
already treats `sudo ` as transparent for the pipe case (`... | sudo sh`).
This generalises the same reading to the command position.
Deliberately NOT applied to the data-sink masking in
`_mask_data_sink_arguments`: peeling there would widen an exemption, and the
conservative reading is the safe one.
No false positives: peeling only changes which token is treated as the
command, so a wrapper around ordinary work resolves to a non-shell executable
and yields nothing, exactly as before (`sudo apt-get update`,
`timeout 60 curl ...`, `nice -n 10 make -j4`, a bare `env`).
Tests: 40 new cases in tests/hermes_cli/test_gateway_restart_loop.py — every
wrapper form against a script reference, a dot-source, a nested `sh -c`
payload and `launchctl submit`, plus the benign wrapped commands, a wrapped
clean script, and the `./timeout`-style lookalike names that pin the additive
reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
Harden the id-keyed-map flatten with an id-preserving merge:
{**value, "id": value.get("id") or key} — an inline "id" wins,
otherwise the map key is adopted (external tools often key by id and
omit the inline copy; plain list(values) would emit id-less records
that collide or get dropped downstream). Non-dict junk values are
skipped with a warning instead of crashing the load. The self-heal
rewrite persists the id-merged, junk-free records.
Tests: key adopted when no inline id (and inline id wins over a
differing key), non-dict junk skipped with warning + list_jobs
survives + self-heal persists only valid records, all-junk map
flattens to [].
Layer on the load-boundary flatten: when load_jobs() encounters an
ID-keyed jobs map ({"jobs": {"<job_id>": {...}, ...}} — written by
external tools or hand edits, never by save_jobs()), it now not only
flattens to the list contract but persists the canonical
{"jobs": [...]} form back to disk via the existing auto-repair path
(save_jobs), so the store self-heals and subsequent reads are
idempotent.
Note: _peek_jobs_unlocked() intentionally does NOT tolerate the dict
shape — it returns None so the save path never shrink-merges against
an unrepaired baseline. The flatten + repair live only at the
load_jobs() boundary.
Regression tests cover the flatten, the reported list_jobs() traceback
path, idempotent on-disk repair, and the empty-map edge case.
Salvaged from PR #92994.
Co-authored-by: a-yeyang <88581400+a-yeyang@users.noreply.github.com>
get_due_jobs() fires purely off the stored next_run_at <= now, with no
check that the stored instant is still an occurrence of the schedule's
current expression. A direct jobs.json edit that narrows schedule.expr
(e.g. daily "0 7 * * *" -> weekdays "0 7 * * 1-5") keeps the stored
next_run_at computed under the old expression, so the job fires on days
the new expression excludes. The within-grace fire and the catch-up
"run once now" path both inherit the wrong instant.
Add a best-effort stale-schedule guard on the fire path: when the stored
next_run_at is not an occurrence of the current cron expression,
re-anchor it via compute_next_run() from the current expression and skip
the fire. Non-cron kinds, missing expr, croniter unavailability, and
malformed input all report a match so the fire path keeps its existing
semantics. Recomputation uses the current expression, so the re-anchor
converges and cannot defer a valid job forever.
Fixes#93049
A newer main-side sniff fast-path skipped any file with a NUL in its
head, short-circuiting before the magic-number check the salvaged fix
added — reintroducing the #77927 bypass. Key the fast-path on
executable magic only; NUL-bearing text falls through to the tail
logic (magic check, size-before-strip, NUL-strip, scan).
The #76762 binary check treats any NUL byte in the first chunk as "compiled
binary, nothing to scan":
if b"\x00" in data:
return None, False
"Contains a NUL" and "is a compiled binary" are different questions, and the
gap between them is a guard bypass. `bash` executes a *text* script straight
past an embedded NUL, so one pad byte disables the entire scan while the
script still runs:
#!/bin/bash
# pad<NUL>
hermes gateway restart
scan("bash padded.sh") -> False (not blocked)
bash padded.sh -> executes the lifecycle command
This shape was blocked before #76762, so the crash fix traded a loud failure
for a silent one.
Keying the check on a leading `#!` is not sufficient: a shebang-less file with
a NUL on any line but the first also executes normally. (A NUL on line 1 of a
shebang-less file is the one shape bash rejects, exit 126 — but that same file
is still executable via `. file`.)
Fix: identify binaries by MAGIC NUMBER — ELF, Mach-O (incl. byte-swapped and
universal/fat), PE/COFF, static archive, gzip, zip — with a shebang always
winning. A NUL-bearing *text* file is scanned with its NULs stripped;
stripping can only splice tokens together, never apart, so it fails closed.
File extensions are deliberately not consulted, so a suffixless shell script
is still scanned.
The size check now runs BEFORE the strip: stripping shrinks the buffer, so
checking afterwards would let an oversized file slip under the threshold and
skip the fail-closed branch. (Caught by
test_oversized_nul_bearing_text_still_fails_closed, which failed on the first
cut of this patch.)
Return values are unchanged, so this does not conflict with the in-flight
crash-class fixes to the same function.
Tests (tests/hermes_cli/test_gateway_restart_loop.py), 3 of which fail on main:
- test_nul_padded_script_is_still_scanned
- test_nul_padded_script_without_shebang_is_scanned
- test_oversized_nul_bearing_text_still_fails_closed
- test_elf_binary_is_not_scanned_as_script (#76762 stays fixed)
- test_macho_binary_is_not_scanned_as_script (incl. fat binary)
- test_clean_script_without_lifecycle_command_not_blocked
`_iter_referenced_shell_scripts` recognises the `source` builtin so a script
pulled in with `source ./restart.sh` gets scanned for lifecycle commands. The
POSIX dot operator is the same builtin, but it was not caught:
if executable_name in {".", "source"}:
`executable_name` is `Path(executable).name`, and `Path(".").name` is the
**empty string** -- pathlib normalises "." to the current directory, whose name
is "". So the set membership never matched for `.`, the sourced script was
never added to the reference walk, and its contents were never scanned.
Verified against current main:
. /tmp/restart.sh -> not blocked (script never scanned)
source /tmp/restart.sh -> blocked
bash /tmp/restart.sh -> blocked
where /tmp/restart.sh contains a `hermes gateway restart` line. Sourcing runs
the script in the current shell, so the dot spelling is not merely equivalent
to `source` -- it is the more common form in practice.
Fix compares the raw token as well as the basename:
if executable in {".", "source"} or executable_name == "source":
Keeping the `executable_name == "source"` arm preserves the existing behaviour
for a path-qualified spelling, while the raw-token test catches `.` without
relying on pathlib normalisation.
Tests (tests/hermes_cli/test_gateway_restart_loop.py):
- test_dot_operator_sourced_script_is_scanned -- the regression; fails on main
- test_source_builtin_sourced_script_is_scanned -- `source` stays blocked
- test_dot_operator_clean_script_not_blocked -- widening the check must not
false-block an innocent `. ./activate.sh`
Found while auditing the guard after #76762. Scoped deliberately to this one
defect; the NUL-padded-script bypass I found in the same audit is a separate
PR.
Runbook prose inside a quoted-delimiter heredoc feeding a data sink
(cat > file <<'EOF') is documentation, not a command this shell will
execute. Mask provably-inert heredoc bodies (tools/shell_heredoc's
conservative stripper, already used by terminal_tool) before scanning.
Fails open on any ambiguity: executable and unquoted-delimiter heredocs
stay scanned. Salvaged from PR #88336 by @zgqq (the heredoc half; its
Branch D boundary and dir-token halves already landed/were fixed).
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.
A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:
for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
label=${item%%:*}; plist=${item#*:}
launchctl bootout "gui/$uid/$label"
launchctl bootstrap "gui/$uid" "$plist"
done
The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.
cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).
`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.
We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.
Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.
Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The separator-class anchor broke two fail-closed tests (binary bytes
decode to U+FFFD adjacent to the CLI name; remote head-c reads). A
negative lookbehind excluding path/word chars keeps the #77173 fix
while preserving every fail-closed content-scan path.
A file path with embedded spaces (/docs/... with lifecycle words in the
filename) matched Branch A via the path tail and hard-blocked innocent
commands. Anchor the CLI name at command position (start, separator, or
substitution opener). Salvaged from PR #77536 by @eaglezzz0522-cloud,
reapplied onto the current pattern with subshell coverage and tests.
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.
contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.
tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.
Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.
Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Branch B of _GATEWAY_LIFECYCLE_PATTERN enumerated launchd verbs but omitted
`bootout` - the modern replacement for the `unload` it already listed, and
the paired inverse of the `bootstrap` it already listed. `remove` (legacy
sibling of bootout) and `disable` (what makes an unload durable) were
missing for the same reason.
This matters because the two enforcement layers are not interchangeable. In
tools/terminal_tool.py under _HERMES_GATEWAY == "1":
- the cron.lifecycle_guard hard block is documented as applying
unconditionally ("force=True cannot help here")
- detect_dangerous_command below it is explicitly skipped when force=True
detect_dangerous_command already flags all three verbs, so the default path
was covered - but with force=True inside the gateway they reached execution
while stop/unload/kickstart did not. SIGTERM then propagates to the child
before the command completes and the service may never come back, which is
the state described in #74973.
The label anchor (\bhermes[.\-]?gateway) is unchanged, so unrelated services
such as `launchctl bootout gui/501/ai.hermes.update-checker` stay runnable.
Adds TestLifecycleGuardLaunchctlParity, which pins the one-directional
invariant: anything the bypassable approval layer flags, the unbypassable
hard block must also catch. Deliberately not equality - the hard block is
legitimately stricter (it also covers load/restart, which the approval layer
leaves alone). Verified failing on the parent commit for exactly bootout,
remove and disable.
Closes#80260
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Add _split_logical_lines() to split on newlines outside quotes, fixing
false positives from quoted multi-line payloads (e.g. python -c "...")
being torn into fragments and scanned as referenced scripts.
- Make _iter_command_segments() use logical line splitting with fallback
to per-physical-line tokenization for unbalanced quotes.
- Add missing word boundaries to _GATEWAY_LIFECYCLE_PATTERN:
* Branch A: trailing \b after restart|stop
* Branch D: leading \b before p?kill to prevent matching "skill"
(and similar words ending in "kill")
Fixes#92372: gateway lifecycle guard false-blocks on prose inside
a referenced data file.
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.
The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.
Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.
Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.
Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.
Fixes#88655
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.
- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
persistent memory