Commit Graph

592 Commits

Author SHA1 Message Date
Victor Kyriazakos 8e8112687b fix(cron): don't reference nonexistent 'hermes cron trigger' in relay-fronted errors
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos 2d1d65de46 fix(cron): accurate error for relay-fronted delivery with no live gateway (NS-773)
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
2026-08-27 19:52:17 -07:00
kshitijk4poor 939dec1348 fix: harden _is_recoverable_error_job against schedule=None
job.get("schedule", {}).get("kind") crashes with AttributeError when
schedule is present but explicitly None (disk corruption edge case).
Use (job.get("schedule") or {}).get("kind") instead, which safely
returns False for None. This pattern is already used at other sites
in the file (e.g. cron/scheduler.py line 170).
2026-08-28 02:18:25 +05:30
pierrenode ba4c2d5253 fix(cron): make a recurring job stuck in state=error recoverable again
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:

* state=completed: a one-shot that genuinely has no more occurrences,
  ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
  fails to produce a next occurrence (e.g. the croniter package is
  missing at runtime). _mark_job_run_locked's own comment is explicit:
  "Recurring jobs must NEVER be silently disabled" (issue #16265) — the
  job is left enabled=True specifically so it keeps being a live,
  recoverable job once the underlying issue resolves.

Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:

* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
  its own is_terminal_job() check) never runs, because the check itself
  skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
  ... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
  itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
  and claim steps the scheduler's own dispatch loop calls immediately
  after get_due_jobs() for anything that DOES make it into the due list —
  both also refuse the job, so even a manually-recovered next_run_at
  would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
  broken recurring job through the normal path.

The only way out was deleting the job and recreating it.

Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.

Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.

New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.

Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
2026-08-28 02:18:25 +05:30
kshitijk4poor 42ac29eacc docs(cron): comment accuracy — booking is fail-open on probe errors; cross-ref status vocabulary
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
2026-08-27 20:39:30 +05:30
liuhao1024 23f597a8f5 fix(cron): verify a persisted final assistant message before booking complete
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).

Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.

Fixes #93820
2026-08-27 20:39:30 +05:30
kshitijk4poor 046a47a301 docs(cron): pin the manual_run_at comparison as intentionally string-exact
Review follow-up on the #94033 salvage: guard the equality gate against a
future 'helpful' datetime normalization — any rewrite of next_run_at must
invalidate the marker, and normalizing would weaken that.
2026-08-27 20:39:20 +05:30
fangliquanflq 7e64a48303 fix(cron): preserve recurring manual run intent 2026-08-27 20:39:20 +05:30
kshitijk4poor cced6fa360 refactor(cron): extract _get_session_db_timeout alongside sibling timeout resolvers
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
  extract it next to _get_script_timeout/_get_media_send_timeout. Using
  load_config() (deep-merge) also removes the default-drift hazard flagged
  in review: cron.session_db_timeout_seconds now resolves from
  DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
  10.0 staying in sync with it, and drops the distant-state coupling to
  run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
  new position the store init happens inside the guarded worker, not before
  it).
2026-08-27 19:07:15 +05:30
Heath Harris 3113f6056b fix(cron): open the session store only after wake-gate and validation early-returns
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.

Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.

Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
2026-08-27 19:07:15 +05:30
Ayush Nangia 2f0f01192d fix(cron): tree-kill script timeout descendants via agent.deadline.kill_process_tree
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.

The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.

Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
  lost" path orphaned setsid grandchildren the same way (whole-bug-class
  rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
  that exits right at the deadline doesn't log a spurious "no signal"
  warning (mirrors _terminate_cron_script_process); pinned by
  test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
  CI load could eat the whole 1s window before the spawner wrote its
  pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
  old path gave a 1s SIGTERM grace window; intended for a deadline-
  expiry hard stop (both docstrings say "hard stop")

Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.

Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
2026-08-27 14:28:06 +05:30
kshitijk4poor ded9470990 refactor(cron): fold simplify-review findings into mirror eligibility
- _target_mirror_eligible accepts a precomputed origin_match so the sole
  production caller stops re-resolving origin + re-running the origin
  match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
  precedence (standalone correctness: per-job False must beat raw global
  True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
  'not inchannel_continuable' and reword it for the widened seed scope.
2026-08-26 16:06:09 +05:30
kshitijk4poor 8313185449 fix(cron): unify in_channel flatten and seed behind one continuable gate
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
2026-08-26 16:06:09 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
Teknium 1fe0f2f3ac feat(cron): import-error cron failures now name gateway code skew and the one-command fix (#95294 part 3)
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.

Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).

Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
2026-08-26 01:23:15 -07:00
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
kshitijk4poor 105999a0c9 refactor(gateway): unify computer-use repair call sites after review
- Make repair_explicit_computer_use_media_paths fail-open internally
  (cosmetic repair must never abort delivery); drop the cron-only
  try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
  no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
  gateway.media_repair seam; pre-existing tests keep pinning the
  gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
2026-08-25 13:06:43 +05:30
kshitijk4poor bb0d5503c2 fix(gateway): widen computer-use media path repair to sibling surfaces
Follow-up to the salvaged fix from PR #94439:

- Extract the repair into gateway/media_repair.py (shared module) and
  re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
  background tasks (_run_background_task_inner) and cron job delivery
  (cron/scheduler.py) — both call agent.run_conversation directly and
  never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
  content first instead of regex-scanning the raw string, which yielded a
  doubled-backslash path artifact and rewrote the response to a path the
  model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
  gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
  the compression-fallback last-user slice (incl. no-user fail-closed).
2026-08-25 13:06:43 +05:30
kshitijk4poor 0eda2ba0c8 fix: remove dead code, deduplicate error constants, fix skill key check
Follow-up to PR #92189 salvage:
- Remove unused job_no_agent_without_script() function (dead code)
- Replace inline NO_AGENT_WITHOUT_SCRIPT_ERROR string in _validate_job_mode_invariants with the constant
- Replace scheduler inline reason string with EMPTY_PAYLOAD_ERROR constant
- Add 'skill' (singular) to job_payload_is_empty 'in job' presence check
2026-08-24 15:47:42 +05:30
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
Teknium 29c5a12e04 fix(cron): warn loudly when the due-scan removes a consumed one-shot that already ran (#93524)
Extracted from PR #93641. Pre-#93615 stores (or hand edits) can carry a
re-armed record whose budget was never reset; the due-scan guard removes it
without firing — correct under the refusal+explicit-re-arm policy, but the
removal must be operator-visible. WARNING now names the remediation
('hermes cron resume <job> --run-now'); the never-ran dead-tick recovery
case keeps its quiet INFO. Diagnosis credit: @liuhao1024 (#93543),
@aniruddhaadak80 (#93585).
2026-08-24 03:13:30 -07:00
leosiedler a0ca7c1920 feat(cron): add explicit one-shot re-arm 2026-08-24 00:34:26 -07:00
leosiedler c3a63a16f1 fix(cron): refuse to run terminal jobs 2026-08-24 00:34:26 -07:00
Teknium ed8ee9a871 fix(cron): misfire backstop honors the one-shot grace window (#93526)
The hosted-provider misfire catch-up (fire_overdue_jobs) fired any runnable
overdue job with no one-shot grace check, so a stored past-due one-shot
bypassed ONESHOT_GRACE_SECONDS and executed arbitrarily late after downtime.
Sibling site of the due-scan gate from #89571; pins both directions with
tests.
2026-08-23 23:33:29 -07:00
spfcraze b37a5bc0df fix(cron): due-scan must not dispatch a one-shot past its grace window
create_job / update_job / resume_job all reject a one-shot whose run time is
more than ONESHOT_GRACE_SECONDS in the past ("will never fire"), and
_recoverable_oneshot_run_at never recovers such a schedule — but
_get_due_jobs_locked dispatched ANY one-shot whose *persisted* next_run_at was
in the past, even hours later (gateway down past the window, host asleep,
hand-edited jobs.json). A wall-clock one-shot then ran hours late, violating
the "will never fire" contract enforced everywhere else.

- Grace gate: a once-kind job whose next_run_dt is more than
  ONESHOT_GRACE_SECONDS in the past is never appended to the due list.
- If no run_claim/fire_claim exists (nothing was ever dispatched), retire the
  record with a diagnostic file so it stops being scanned and the miss is
  operator-visible.
- If a (possibly stale) claim exists, a run may still be in flight in another
  process: skip this scan but KEEP the record so its mark_job_run can land
  (avoids re-introducing mid-flight record deletion).
- Manual re-trigger still works: trigger_job sets next_run_at=now (inside
  grace) so an explicitly re-run stale one-shot fires.

Tests (tests/cron/test_oneshot_grace_due_scan.py): stale-not-due+retired,
within-grace-still-due, stale+claim-skipped-but-kept, retriggered-is-due, and
recurring-jobs-unaffected.
2026-08-23 23:33:29 -07:00
RickyYii 6b3a7af73d fix(security): cover privilege wrappers and command-string options
Review follow-up on #84203. Both points reproduce; neither was a regression
from the first pass, but both are live bypasses of the same guard.

**Privilege and namespace wrappers were missing.** The allowlist covered the
coreutils-shaped wrappers but not the privilege ones, so each of these ran a
lifecycle script straight past the walk:

    pkexec bash ~/restart.sh
    runuser -u root -- bash ~/restart.sh
    setpriv --reuid=0 -- bash ~/restart.sh
    systemd-run --scope bash ~/restart.sh
    nsenter --target 1 --mount bash ~/restart.sh
    unshare -r bash ~/restart.sh

Added `pkexec`, `su`, `runuser`, `setpriv`, `systemd-run`, `nsenter` and
`unshare`, each with the value-taking options that would otherwise be
mistaken for the command (`nsenter -t 1`, `systemd-run -p X=1`,
`runuser -u root`, …).

**An option can carry a command STRING, not an argv tail.** `env -S` and
`su`/`runuser` `-c` take shell source. The peel treated the operand as an
opaque value and skipped it, so `env -S 'bash ~/restart.sh'` was never
scanned — the string went unread rather than being recursed into.

`_STRING_COMMAND_OPTIONS` now names those options and their values are
re-scanned as shell source, the same treatment `sh -c` payloads already get.
They are read at the ORIGINAL command token, before the transparent-prefix
peel, because peeling past `su`/`env` would discard the very option carrying
the command. `--opt value` and `--opt=value` are both handled.

Scope, stated plainly: this is an enumerated allowlist, not a general
solution to "wrapper that execs its tail". A wrapper outside the set, or a
value-taking option outside these tables, still resolves to no reference —
that fails open, exactly as it did before this PR, and it is a miss rather
than a false block. The reviewer offered "extend the set with tests, or
document that the list is heuristic"; this does the first and states the
second.

Tests: 23 new cases (220 in the file) — every added wrapper against a script
reference including the value-operand option forms, both command-string
option spellings for env/su/runuser, and the same wrappers around ordinary
work (`pkexec systemctl status nginx`, `su -c 'ls -la'`, `env -S 'echo hi'`,
`nsenter -t 1 -m ps aux`) which must stay allowed. 15 fail on the tree
before this commit.

False positives re-checked at scale: the 9,258 command lines from this
repo's own scripts and docs give an identical verdict set before and after —
0 new false positives, 0 lost detections, 0 exceptions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii a19e1bae10 fix(cron): stop a relative path from disabling the data-sink exemption
`_mask_data_sink_arguments` exempts lifecycle text living in the arguments of
executables that cannot run them (`grep`, `rg`, `journalctl`, `sqlite3`, …),
so hunting for a restart string in logs is diagnostics rather than a command.
The exemption is dropped when an argument looks like an escape back into
execution — including anything starting with a dot, because sqlite3 spells
its escapes as dot-commands (`.shell`, `.system`).

But `.`, `./x` and `../x` are ordinary path operands, and

    grep -r 'systemctl restart hermes-gateway' .

is the most ordinary recursive search there is. The leading-dot test treated
its `.` operand as a sqlite3 escape, disabled masking for the whole segment,
and blocked the command outright — the exact false-positive class the
exemption exists to prevent, on the shape most likely to hit it. Searching a
relative subdirectory (`./logs`, `../archive`) fails the same way, as does a
relative sqlite3 database path (`sqlite3 ./stats.db "SELECT ..."`).

Require a dot followed by a NAME character (`^\.[A-Za-z]`) so a dot-command
still defeats the exemption while a relative path stays a path. A dotfile
operand (`.env`) still reads as a dot-command — conservative, and unchanged
from today's behavior.

This narrows a security guard in the permissive direction, so the escape
hatches are pinned explicitly: with a relative-path operand present,
`.shell`/`.system`, psql's `\!`, a pipe into `sh`/`bash`/`sudo sh`/`xargs`,
command substitution, and a `;`/`&&` continuation all still block. Only the
segment's own data arguments are masked, and only when nothing in it can
reach execution.

Tests: 18 new cases in tests/hermes_cli/test_gateway_restart_loop.py — the
relative-path shapes that must now be allowed, plus the ten escape-hatch
shapes that must still block. The allow cases fail on the unfixed tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
RickyYii 5921ba8c06 fix(security): see through wrapper prefixes in the gateway lifecycle guards
`sudo`, `env`, `nohup`, `timeout` and friends exec their argument tail, so
the command that actually runs sits further right. Three guards read only the
first token of a segment, saw the wrapper, and never inspected what it runs:

  bash ~/restart.sh                      → blocked
  sudo bash ~/restart.sh                 → allowed
  launchctl submit -l com.x -- helper    → blocked
  sudo launchctl submit -l com.x -- helper → allowed

Same foot-gun, one word of prefix. That reaches both enforcement points —
`cron.jobs.create_job` and `tools/terminal_tool.py` under `_HERMES_GATEWAY=1`
— and defeats the label-independent submit block that #62891 added precisely
because a persistent helper is the indirect route to a restart loop.

`_peel_transparent_prefixes()` walks past a bounded chain of these wrappers,
skipping their own options, their value-taking options (`sudo -u deploy`,
`stdbuf -o0`), `VAR=value` assignments, a `--` end-of-options separator, and
`timeout`'s duration operand, then returns the index of the real command. It
is applied to the referenced-script walk, the `sh -c` payload walk, and the
`launchctl submit`/`bootstrap` block.

In the referenced-script walk the peel is ADDITIVE — the segment is read at
the original token and again at the peeled one — because peeling must never
remove a reference the un-peeled read would have found. A local script named
`./timeout` is a script, not the coreutils wrapper, and consuming it as a
prefix would have silently stopped scanning it. (The other two call sites
need no such care: no wrapper name is also a shell name or `launchctl`, so
peeling there can only add.) That split is why the per-index logic now lives
in `_references_at()`.

This is not a new reading of shell syntax for this module — `_PIPE_TO_INTERPRETER`
already treats `sudo ` as transparent for the pipe case (`... | sudo sh`).
This generalises the same reading to the command position.

Deliberately NOT applied to the data-sink masking in
`_mask_data_sink_arguments`: peeling there would widen an exemption, and the
conservative reading is the safe one.

No false positives: peeling only changes which token is treated as the
command, so a wrapper around ordinary work resolves to a non-shell executable
and yields nothing, exactly as before (`sudo apt-get update`,
`timeout 60 curl ...`, `nice -n 10 make -j4`, a bare `env`).

Tests: 40 new cases in tests/hermes_cli/test_gateway_restart_loop.py — every
wrapper form against a script reference, a dot-source, a nested `sh -c`
payload and `launchctl submit`, plus the benign wrapped commands, a wrapped
clean script, and the `./timeout`-style lookalike names that pin the additive
reading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cSddnhxiUmdGbgyKnpg8p
2026-08-23 18:42:59 -07:00
Teknium f2639f8872 fix(cron): preserve map keys as ids and skip junk values when flattening id-keyed jobs.json
Harden the id-keyed-map flatten with an id-preserving merge:
{**value, "id": value.get("id") or key} — an inline "id" wins,
otherwise the map key is adopted (external tools often key by id and
omit the inline copy; plain list(values) would emit id-less records
that collide or get dropped downstream). Non-dict junk values are
skipped with a warning instead of crashing the load. The self-heal
rewrite persists the id-merged, junk-free records.

Tests: key adopted when no inline id (and inline id wins over a
differing key), non-dict junk skipped with warning + list_jobs
survives + self-heal persists only valid records, all-junk map
flattens to [].
2026-08-23 18:27:32 -07:00
Teknium ec4b3bc06a fix(cron): self-heal id-keyed jobs.json to canonical list form on load
Layer on the load-boundary flatten: when load_jobs() encounters an
ID-keyed jobs map ({"jobs": {"<job_id>": {...}, ...}} — written by
external tools or hand edits, never by save_jobs()), it now not only
flattens to the list contract but persists the canonical
{"jobs": [...]} form back to disk via the existing auto-repair path
(save_jobs), so the store self-heals and subsequent reads are
idempotent.

Note: _peek_jobs_unlocked() intentionally does NOT tolerate the dict
shape — it returns None so the save path never shrink-merges against
an unrepaired baseline. The flatten + repair live only at the
load_jobs() boundary.

Regression tests cover the flatten, the reported list_jobs() traceback
path, idempotent on-disk repair, and the empty-map edge case.

Salvaged from PR #92994.

Co-authored-by: a-yeyang <88581400+a-yeyang@users.noreply.github.com>
2026-08-23 18:27:32 -07:00
WK Wong 5a24dcf4f2 fix(cron): normalize id-keyed jobs stores on load 2026-08-23 18:27:32 -07:00
liuhao1024 8f9abc9873 fix(cron): re-anchor stale next_run_at after direct jobs.json schedule edits
get_due_jobs() fires purely off the stored next_run_at <= now, with no
check that the stored instant is still an occurrence of the schedule's
current expression. A direct jobs.json edit that narrows schedule.expr
(e.g. daily "0 7 * * *" -> weekdays "0 7 * * 1-5") keeps the stored
next_run_at computed under the old expression, so the job fires on days
the new expression excludes. The within-grace fire and the catch-up
"run once now" path both inherit the wrong instant.

Add a best-effort stale-schedule guard on the fire path: when the stored
next_run_at is not an occurrence of the current cron expression,
re-anchor it via compute_next_run() from the current expression and skip
the fire. Non-cron kinds, missing expr, croniter unavailability, and
malformed input all report a match so the fire path keeps its existing
semantics. Recomputation uses the current expression, so the re-anchor
converges and cannot defer a valid job forever.

Fixes #93049
2026-08-23 18:25:35 -07:00
Teknium a9e46229b2 fix: sniff fast-path keys on binary magic only, not NUL presence
A newer main-side sniff fast-path skipped any file with a NUL in its
head, short-circuiting before the magic-number check the salvaged fix
added — reintroducing the #77927 bypass. Key the fast-path on
executable magic only; NUL-bearing text falls through to the tail
logic (magic check, size-before-strip, NUL-strip, scan).
2026-08-23 18:24:36 -07:00
Meng Chee 92edb861be fix(cron): close NUL-padded script bypass in lifecycle guard
The #76762 binary check treats any NUL byte in the first chunk as "compiled
binary, nothing to scan":

    if b"\x00" in data:
        return None, False

"Contains a NUL" and "is a compiled binary" are different questions, and the
gap between them is a guard bypass. `bash` executes a *text* script straight
past an embedded NUL, so one pad byte disables the entire scan while the
script still runs:

    #!/bin/bash
    # pad<NUL>
    hermes gateway restart

    scan("bash padded.sh")  -> False   (not blocked)
    bash padded.sh          -> executes the lifecycle command

This shape was blocked before #76762, so the crash fix traded a loud failure
for a silent one.

Keying the check on a leading `#!` is not sufficient: a shebang-less file with
a NUL on any line but the first also executes normally. (A NUL on line 1 of a
shebang-less file is the one shape bash rejects, exit 126 — but that same file
is still executable via `. file`.)

Fix: identify binaries by MAGIC NUMBER — ELF, Mach-O (incl. byte-swapped and
universal/fat), PE/COFF, static archive, gzip, zip — with a shebang always
winning. A NUL-bearing *text* file is scanned with its NULs stripped;
stripping can only splice tokens together, never apart, so it fails closed.
File extensions are deliberately not consulted, so a suffixless shell script
is still scanned.

The size check now runs BEFORE the strip: stripping shrinks the buffer, so
checking afterwards would let an oversized file slip under the threshold and
skip the fail-closed branch. (Caught by
test_oversized_nul_bearing_text_still_fails_closed, which failed on the first
cut of this patch.)

Return values are unchanged, so this does not conflict with the in-flight
crash-class fixes to the same function.

Tests (tests/hermes_cli/test_gateway_restart_loop.py), 3 of which fail on main:

- test_nul_padded_script_is_still_scanned
- test_nul_padded_script_without_shebang_is_scanned
- test_oversized_nul_bearing_text_still_fails_closed
- test_elf_binary_is_not_scanned_as_script       (#76762 stays fixed)
- test_macho_binary_is_not_scanned_as_script     (incl. fat binary)
- test_clean_script_without_lifecycle_command_not_blocked
2026-08-23 18:24:36 -07:00
Meng Chee da30db8e8c fix(cron): scan dot-operator sourced scripts in lifecycle guard
`_iter_referenced_shell_scripts` recognises the `source` builtin so a script
pulled in with `source ./restart.sh` gets scanned for lifecycle commands. The
POSIX dot operator is the same builtin, but it was not caught:

    if executable_name in {".", "source"}:

`executable_name` is `Path(executable).name`, and `Path(".").name` is the
**empty string** -- pathlib normalises "." to the current directory, whose name
is "". So the set membership never matched for `.`, the sourced script was
never added to the reference walk, and its contents were never scanned.

Verified against current main:

    . /tmp/restart.sh        -> not blocked   (script never scanned)
    source /tmp/restart.sh   -> blocked
    bash /tmp/restart.sh     -> blocked

where /tmp/restart.sh contains a `hermes gateway restart` line. Sourcing runs
the script in the current shell, so the dot spelling is not merely equivalent
to `source` -- it is the more common form in practice.

Fix compares the raw token as well as the basename:

    if executable in {".", "source"} or executable_name == "source":

Keeping the `executable_name == "source"` arm preserves the existing behaviour
for a path-qualified spelling, while the raw-token test catches `.` without
relying on pathlib normalisation.

Tests (tests/hermes_cli/test_gateway_restart_loop.py):

- test_dot_operator_sourced_script_is_scanned -- the regression; fails on main
- test_source_builtin_sourced_script_is_scanned -- `source` stays blocked
- test_dot_operator_clean_script_not_blocked -- widening the check must not
  false-block an innocent `. ./activate.sh`

Found while auditing the guard after #76762. Scoped deliberately to this one
defect; the NUL-padded-script bypass I found in the same audit is a separate
PR.
2026-08-23 18:24:36 -07:00
zgqq 9aa0721b23 fix: inert heredoc bodies no longer trip the gateway lifecycle guard (#88336)
Runbook prose inside a quoted-delimiter heredoc feeding a data sink
(cat > file <<'EOF') is documentation, not a command this shell will
execute. Mask provably-inert heredoc bodies (tools/shell_heredoc's
conservative stripper, already used by terminal_tool) before scanning.
Fails open on any ambiguity: executable and unquoted-delimiter heredocs
stay scanned. Salvaged from PR #88336 by @zgqq (the heredoc half; its
Branch D boundary and dir-token halves already landed/were fixed).
2026-08-23 18:10:31 -07:00
Artur Hapantsou b34edd6b01 fix: execute_code and argv-list payloads no longer bypass the gateway lifecycle guard (#68289)
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
2026-08-23 18:01:59 -07:00
KeaneYan 1c791cbfe6 fix(gateway): resolve uninstall lifecycle guard conflict 2026-08-23 18:01:59 -07:00
BotUser 679e07a074 fix(gateway): close order-dependency + missing-verb gap in launchctl lifecycle guards
The gateway-lifecycle guards in cron/lifecycle_guard.py (Branch B, the
unconditional hard-block used by cron creation and the terminal tool when
_HERMES_GATEWAY=1) and tools/approval.py's launchctl rule both matched
`launchctl <verb> ... hermes[.-]?gateway` as a single sequential regex,
requiring the hermes-gateway label to appear literally AFTER the verb.

A shell command that builds the label earlier in the string — e.g. a
for-loop reading labels from a list defined before the actual launchctl
call — defeats that ordering entirely:

    for item in 'ai.hermes.gateway-apollo:...' 'ai.hermes.gateway:...'; do
      label=${item%%:*}; plist=${item#*:}
      launchctl bootout "gui/$uid/$label"
      launchctl bootstrap "gui/$uid" "$plist"
    done

The literal text "hermes.gateway" only ever appears in the for-list,
never after "bootout" — so `[^\n]*\bhermes[.\-]?gateway` never matches at
the verb's position, even though the command unambiguously targets the
gateway's own launchd label.

cron/lifecycle_guard.py's verb list also didn't include `bootout` at all
(present in tools/approval.py's list and covered by its own test suite —
`launchctl bootout ai.hermes.gateway` is explicitly asserted as dangerous
there — so the omission in the sibling file looks like list drift between
the two guards rather than an intentional exclusion).

`bootout` is the verb that actually deregisters a launchd job (unlike
kickstart/stop, which just bounce a still-registered one), so a command
using it evades both guards, then removes the service from launchd with
no supervisor left to bring it back — worse than a simple restart-loop.

We hit this for real: a gateway self-restart (triggered from a chat
request to change the default model) used a raw terminal `launchctl
bootout`/`bootstrap` loop across 4 launchd labels instead of the normal
`hermes gateway restart` path. It slipped past both guards, self-bootout
killed the process mid-drain before its own follow-up bootstrap could
run, and all 4 gateway profiles ended up fully deregistered from launchd
with zero user approval (approvals.mode: manual was configured) until
someone manually re-bootstrapped them.

Fix: both guards now check "a launchctl lifecycle verb appears somewhere
AND a hermes-gateway label appears somewhere", independent of order, and
cron/lifecycle_guard.py's verb list gains bootout/kill/disable/remove to
match tools/approval.py's existing set. Internal recovery code
(hermes_cli/gateway.py's own `subprocess.run(["launchctl", "bootout",
...])` calls) is unaffected — these guards only scan shell-command
strings composed by the agent's terminal/cron tools, not the CLI's
trusted internal subprocess argument lists.

Adds regression tests in both test files reproducing the exact incident
command (label built in an earlier for-loop segment, referenced only via
`$label` at the point of the verb).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 17:56:36 -07:00
Teknium 20e308fea7 fix: use lookbehind anchor so binary-decoded and remote-read content still scans
The separator-class anchor broke two fail-closed tests (binary bytes
decode to U+FFFD adjacent to the CLI name; remote head-c reads). A
negative lookbehind excluding path/word chars keeps the #77173 fix
while preserving every fail-closed content-scan path.
2026-08-23 17:50:48 -07:00
eaglezzz0522-cloud 180f981125 fix: lifecycle guard Branch A anchors the CLI name at command position (#77173 path false positive)
A file path with embedded spaces (/docs/... with lifecycle words in the
filename) matched Branch A via the path tail and hard-blocked innocent
commands. Anchor the CLI name at command position (start, separator, or
substitution opener). Salvaged from PR #77536 by @eaglezzz0522-cloud,
reapplied onto the current pattern with subshell coverage and tests.
2026-08-23 17:50:48 -07:00
Axel Vanni 6d501c2958 fix(cron): make gateway lifecycle matching shell-token aware (#80269)
The hard block matched raw command text, but a shell resolves quote
splicing (`kick"start"`) and backslash escaping (`kick\start`) into the
literal verb before execution. So `launchctl kick"start" -k
gui/501/ai.hermes.gateway` ran exactly as the blocked `kickstart` form
while both the non-bypassable block and the approval detector missed it —
leaving an approval-bypassing gateway self-lifecycle operation reachable.

contains_gateway_lifecycle_command now runs a second pass over
shlex-tokenized command segments, where quotes and escapes are already
resolved. It stays anchored on a hermes-gateway identifier, so prose and
non-gateway hermes services are unaffected. Because this function is the
single choke point _contains_unsafe_gateway_action calls at every
recursion level, referenced-script and `sh -c` payload scanning inherit
the fix.

tools/approval.py had the same gap for quote splices: backslash escapes
are stripped by _normalize_command_for_detection, but quote splicing in an
ARGUMENT position is not touched by _deobfuscate_shell_word_for_detection
(scoped to command-position words, deliberately — widening it would let
quoted prose match the destructive patterns). It now delegates to the
fixed guard as a last check, so an ordinary pattern match still wins and
keeps its more specific reason string.

Tests: quoted, single-quoted and backslash-spliced verbs across the
launchctl/systemctl/hermes branches, the spliced gateway identifier
itself, a splice nested in an `sh -c` payload (resolves one level deeper,
asserted at the recursive entry point terminal_tool actually calls), plus
negative cases proving prose and non-gateway labels stay unblocked.

Verified on Windows: no regressions — the 10 remaining failures across
tests/tools/test_approval.py, tests/hermes_cli/test_gateway_restart_loop.py
and tests/cron are identical on the unmodified baseline (POSIX file modes,
symlink privileges, and /bin/bash script paths).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
Axel Vanni 320d884d88 fix(cron): cover bootout/remove/disable in the gateway lifecycle guard
Branch B of _GATEWAY_LIFECYCLE_PATTERN enumerated launchd verbs but omitted
`bootout` - the modern replacement for the `unload` it already listed, and
the paired inverse of the `bootstrap` it already listed. `remove` (legacy
sibling of bootout) and `disable` (what makes an unload durable) were
missing for the same reason.

This matters because the two enforcement layers are not interchangeable. In
tools/terminal_tool.py under _HERMES_GATEWAY == "1":

  - the cron.lifecycle_guard hard block is documented as applying
    unconditionally ("force=True cannot help here")
  - detect_dangerous_command below it is explicitly skipped when force=True

detect_dangerous_command already flags all three verbs, so the default path
was covered - but with force=True inside the gateway they reached execution
while stop/unload/kickstart did not. SIGTERM then propagates to the child
before the command completes and the service may never come back, which is
the state described in #74973.

The label anchor (\bhermes[.\-]?gateway) is unchanged, so unrelated services
such as `launchctl bootout gui/501/ai.hermes.update-checker` stay runnable.

Adds TestLifecycleGuardLaunchctlParity, which pins the one-directional
invariant: anything the bypassable approval layer flags, the unbypassable
hard block must also catch. Deliberately not equality - the hard block is
legitimately stricter (it also covers load/restart, which the approval layer
leaves alone). Verified failing on the parent commit for exactly bootout,
remove and disable.

Closes #80260

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 16:28:56 -07:00
webtecnica c595d3564a fix(cron): block profile-flag gateway restart/stop when self-targeting (#78028) 2026-08-23 16:24:03 -07:00
webtecnica a74eb2dd41 fix(lifecycle_guard): quote-aware command segmentation and word boundaries
- Add _split_logical_lines() to split on newlines outside quotes, fixing
  false positives from quoted multi-line payloads (e.g. python -c "...")
  being torn into fragments and scanned as referenced scripts.

- Make _iter_command_segments() use logical line splitting with fallback
  to per-physical-line tokenization for unbalanced quotes.

- Add missing word boundaries to _GATEWAY_LIFECYCLE_PATTERN:
  * Branch A: trailing \b after restart|stop
  * Branch D: leading \b before p?kill to prevent matching "skill"
    (and similar words ending in "kill")

Fixes #92372: gateway lifecycle guard false-blocks on prose inside
a referenced data file.
2026-08-23 16:18:38 -07:00
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Jack Lau dd03471858 fix(cron): nudge review of escaped-run failures too
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.

The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.

Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.

Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.

Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.

Fixes #88655
2026-08-22 03:19:08 +05:30
Teknium a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Adolanium fc9cbc872d fix(cron): do not load MEMORY.md into scheduled jobs
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.

Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
2026-08-21 13:24:43 +05:30