Commit Graph

427 Commits

Author SHA1 Message Date
Teknium 13fb87af92 refactor(cron): decompose scheduler run_job/_deliver_result and dedupe delivery lanes
cron/scheduler.py (8535 -> 6353):
- _deliver_result split into per-target helpers: _resolve_target_transport (live/relay/standalone
  transport + enablement), _deliver_via_live_adapter (_live_route_metadata for Telegram DM-topic
  vs forum routing, _live_send_text with the cancel()-based timeout disambiguation,
  _live_send_media, _seed_live_delivery_sessions), _standalone_send/_deliver_standalone. The
  three interpreter-shutdown skip branches and the repeated log+append+continue pattern collapse
  into one _note_target_error / one shutdown message; _TargetDelivery carries per-target state.
- Thread/channel session seeding unified into _seed_cron_session (was two near-identical
  functions).
- run_job decomposed into _run_no_agent_job, _apply_monitor_gate, _load_cron_job_config
  (_CronJobConfig), _resolve_job_runtime, _check_model_drift, _open_cron_session_db,
  _run_agent_with_watchdog, _finalize_cron_session, plus one _run_doc_header and one _audit
  closure for the success/failure paths.
- _run_one_job_body: ownership-lost bookkeeping, delivery composition and outcome classification
  extracted; tick: _acquire_tick_lock/_release_tick_lock, _maybe_reap_dead_owners,
  _sweep_stale_inflight_for_tick, _process_due_job, _submit_with_guard.
- _build_job_prompt: context_from injection and skill loading extracted; one
  _prepend_context_block for the four fenced-data blocks.
- One _start_heartbeat_thread for the script- and fire-claim heartbeat threads.
- Dropped unreachable return in SharedRouteAdapters.get; boolean-return and nested-if shapes
  collapsed.
- Comments/docstrings compacted by hand, rationale kept (fd-leak reason for the late SessionDB
  close callback, title-persistence rules, no_agent classification gate, inactivity-vs-provider
  timeout ordering, stale-claim force-release, interruption token keying).
2026-09-02 13:31:43 -07:00
kshitijk4poor 95f62ca3bf fix(cron): validate failure_deliver at preflight and dashboard update lanes
Follow-up to the failure_deliver salvage (#100375):

- _preflight_check_delivery also checks the failure lane, so a typo'd
  failure_deliver platform blocks at config-validation time instead of
  surfacing only when a failure occurs — exactly when the notice must
  not be lost. Duplicate lanes are checked once.
- The dashboard cron-update normalizer treats failure_deliver like
  deliver (text normalization; empty clears the optional override
  instead of coalescing), closing the one update path that could write
  an unnormalized value into jobs.json.

4 guard tests; both fixes mutation-checked (neutralize -> red, restore -> green).
2026-09-02 20:16:14 +05:30
Victor Kyriazakos fd35e1ec5a fix(cron): delivery bookkeeping reads the failure lane it actually routed through
Review findings (Salt, NS-788):

B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.

S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.

S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.

Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
2026-09-02 20:16:14 +05:30
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium a6351a71e5 fix(cron): route a credentialless satellite's cron delivery through the primary adapter for exact profile_routes targets (#101113)
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.

- cron/scheduler.py: factor the preflight's primary-config route loader into
  `_primary_profile_routes_for_current_home()` (one owner for both halves,
  so route semantics cannot drift) and add `SharedRouteAdapters`, a
  read-only view over the primary adapter map that resolves an adapter for
  a (platform, target) ONLY when an enabled primary route with a
  chat_id/thread_id maps that exact target to the current profile —
  using the same `ProfileRoute.matches` predicate as inbound routing.
  `_deliver_result` resolves the transport per target from it; everything
  else (unmatched chat, disabled route, route for another profile, no
  primary adapter, guild-only route) is a miss and never uses the primary
  bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
  gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
  fallback: with no matching route the view is falsy and delivers nothing.

Fixes #101113
2026-09-02 06:27:24 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
muhifni 1cd736ff63 fix(terminal): scope terminal config per turn under profile multiplexing
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.

Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:

- tools/terminal_scope.py: ContextVar holding the routed profile's
  COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
  TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
  resolves ONLY from it - an omitted key yields the defined default,
  never os.environ. Unreadable/malformed policy installs a refusal
  scope; terminal_tool / execute_code refuse instead of running under
  ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
  _profile_runtime_scope, tui_gateway session/build/turn scopes, cron
  per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
  (_get_env_config, _resolve_container_task_id shared key, orphan
  reaper lifetime, degraded mode), gateway/platforms/base.py docker
  media translation (volumes, shared key, persistence), runtime_cwd /
  agent_init / skill_utils / code_execution_tool / file_tools cwd
  anchors, prompt_builder / browser_tool / env_probe backend checks,
  gateway footer, @-refs and slash-command cwd. env_probe resolves the
  backend in the caller's context, since the probe worker thread does
  not inherit the ContextVar.

Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.

Fixes #68559
Fixes #94200
Fixes #101132
Fixes #95470

Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
2026-09-02 05:34:28 -07:00
Teknium 00a7115a02 fix(cron): make cron push-notify configurable (cron.delivery.notify) and surface UNVERIFIED live deliveries in cron list/doctor
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.

An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.

Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
2026-09-02 00:56:52 -07:00
yoma 4b69ba22fc fix(cron): mark live deliveries as final notifications 2026-09-02 00:56:52 -07:00
kaiomp dd72b42ba4 fix(cron): require positive evidence for live-adapter delivery confirmation
A cron job fired, the scheduler logged "delivered to telegram:<chat> via
live adapter", and nothing reached Telegram (#77763). The log line was not
evidence of a send:

* the silence-narration filter returns {"success": True, "delivered": False}
  (a successful *drop*), and the dict-normalization branch read only
  "success", so a filtered message counted as delivered;
* an empty payload (no text, no media) skipped the send entirely and still
  fell into the "delivered" branch;
* the log line named the chat but not the lane, so a wrong-thread delivery
  and a phantom one are indistinguishable after the fact.

_confirm_adapter_delivery now inspects both result shapes: an explicit
`delivered: False` is a rejection even with a truthy `success`, and a
success with no message_id and no raw_response is accepted but logged as
UNVERIFIED. The empty-payload case fails closed into the existing
standalone/warn handling, and the delivered log carries thread= and
message_id=.

Failing closed on the live lane is only half the fix on a native target:
the standalone fallback sent the same empty payload, and the Telegram
adapter returns SendResult(success=True) for empty content without an API
call — a phantom live delivery became a phantom standalone one. Both
_send_to_platform call sites now sit behind one skip guard, so "empty
payload fails closed" holds on every lane (#77763).
2026-09-02 00:56:52 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium be597fc730 fix: extend corrupt-config fail-closed guard to gateway, serve, and cron surfaces
Follow-up to the salvaged #81988 CLI guard (issue #81952):
- gateway/run.py::main() refuses startup (exit 2) on unparseable config.yaml
- hermes serve headless path (cmd_dashboard) gets the same guard
- cron run_job() fails the job with the guard error before AIAgent
  construction (no_agent script jobs exempt — no token spend)
- HERMES_IGNORE_USER_CONFIG=1 / --ignore-user-config escape hatch honored
  on every surface
2026-09-01 07:00:22 -07:00
Teknium c74cf2333c fix: restore _inactivity_watchdog_loop dropped in rebase conflict resolution 2026-08-31 10:42:39 -07:00
HexLab98 85bc25c949 fix(terminal): bound env.execute wait so a wedged poll cannot disable every timer
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
2026-08-31 10:42:39 -07:00
Andrew Bagrin b7c59bda54 fix(cron): isolate per-execution working directories 2026-08-31 09:59:39 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
Teknium 3a351a9665 feat(worktree): pushed open-PR lanes reclaim their disk; cron tick prunes worktrees
Two growth leaks closed:

1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
   trees, ~18GB on the reporting box): managed installs fetch with a
   single-branch refspec, so pushed PR branches never get refs/remotes/*
   entries and read as 'unpushed' forever. When a clean tree's branch head
   EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
   is redundant: reap the TREE, keep the BRANCH ref (shielded from the
   orphaned-branch pass). Anything diverged/unverifiable stays preserved.
   Applied to both the startup pruner and hermes worktree prune/list.

2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
   gateway-driven boxes accumulated trees for days. The scheduler tick now
   dispatches the same conservative pruner on a daemon thread, throttled
   to once per 6h, against the install checkout + job-workdir repos that
   have a .worktrees/ dir.
2026-08-31 07:27:50 -07:00
Teknium 43ca78aa40 fix(cron): keep SessionDB kwargs-free in the context-preserving worker
The late-result close callback (#72782) retrieves the future's SessionDB
and closes it; passing an explicit db_path kwarg broke the hanging-init
test's mock shape. contextvars.copy_context().run(SessionDB) alone is
sufficient — SessionDB resolves its default path from get_hermes_home(),
which reads the profile ContextVar.
2026-08-31 06:02:32 -07:00
liuhao1024 d5d0613778 fix(cron): profile_routes rescue for satellite-profile delivery preflight
Under gateway.multiplex_profiles the primary gateway's in-process ticker
fires satellite-profile jobs and delivers through the primary's live
adapters (#69377) — the satellite home intentionally holds no platform
credentials (its own token would be a duplicate_credential fatal).
_preflight_check_delivery loads the gateway config of the job's OWN home,
so a profile_routes-routed platform reads as unconnected there and the
job is permanently blocked before any LLM call with a misleading
"not connected" error (#97476).

When the own-home config reports a platform unconnected, consult the
primary home's profile_routes: an enabled route matching the platform
that points at the profile currently being served means delivery is the
primary gateway's to make — pass the check. The primary config.yaml is
read directly (both top-level and nested gateway. forms) instead of via
load_gateway_config() so no primary platform config leaks into the
satellite process's environment. Lookup failures and missing configs
fail closed (the block stands).
2026-08-31 06:02:32 -07:00
Kevin dfa5004a04 fix(cron): preserve profile context during session DB init 2026-08-31 06:02:32 -07:00
Jason S. Mow edd6ad53ab fix(cron): resolve multiplex home-channel chat id from the owning profile's secret scope
The #83182 fix moved the secret scope to span execute→deliver so the
delivery path resolves platform TOKENS (e.g. TELEGRAM_BOT_TOKEN) from the
job-owning profile's .env instead of the host process os.environ. But
cron delivery also resolves the DESTINATION via the legacy
<PLATFORM>_HOME_CHANNEL env mirror, and _env_home_target_chat_id /
_get_home_target_thread_id read it through raw os.getenv — NOT
get_secret. So under multiplex, a socrates cron job whose tick was won
by the default gateway process resolved DISCORD_HOME_CHANNEL from the
default profile's os.environ and landed in the default pax-hermes-chat
instead of socrates-hermes-chat.

Make the chat-id and thread-id resolution legs read through get_secret
when a profile secret scope is installed (mirroring the token leg), so
the owning profile's home channel wins. Non-multiplex / no-scope callers
keep reading os.environ as before.

Verified: with os.environ=default chat and the socrates scope active,
_env_home_target_chat_id('discord') returns the socrates chat; with no
scope it returns the os.environ legacy value.
2026-08-31 06:02:32 -07:00
ghosty93 b8c18b8cb7 docs(cron): explain why the secret-scope reset must stay function-level
Review follow-up: the comment records that the reset scopes delivery,
deferred-agent teardown, claim-loss handling and bookkeeping, and that an
earlier inner-finally placement left _deliver_result unscoped — so a
tidy-up does not silently reintroduce the bug.
2026-08-31 06:02:32 -07:00
ghosty93 07f6518b6d fix(cron): retain profile secret scope through delivery 2026-08-31 06:02:32 -07:00
Patrickk 792dbea777 fix(cron): forward request_overrides into scheduled-job agents
Salvaged from #56876 (cron half only; the delegation half is superseded
by #98237). run_job's ephemeral AIAgent constructor passed api_key /
base_url / provider / api_mode from the resolved runtime but dropped
request_overrides, so cron jobs on custom providers silently lost
extra_body / extra_headers request settings.
2026-08-29 19:12:51 -07:00
kshitijk4poor 0dc9367163 fix(cron): widen deleted-profile protection to all cron mkdir sites
Replace #96637's inline active_profile_homes() closure with #96508's
module-level _existing_profile_homes() filter (testable in isolation).
Widen _ensure_cron_dir from 3 to 12 mkdir sites across cron/ so every
directory creation fails closed for deleted named profiles, not just
the 3 originally protected. Add _is_named_profile_path() that checks
'profiles' in path parts (works for subdirs like cron/output/<job> and
scripts/ that the original parent.name heuristic couldn't reach).

Co-authored-by: misterdas <das7514@gmail.com>
2026-08-28 13:39:19 +05:30
Gille 000d22b9db fix(cron): keep deleted profiles from returning 2026-08-28 13:39:19 +05:30
Teknium 31579f781e fix(cron): transient run prompt survives the relay-fronted gateway forward
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
2026-08-27 20:53:02 -07:00
Victor Kyriazakos 8e8112687b fix(cron): don't reference nonexistent 'hermes cron trigger' in relay-fronted errors
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
2026-08-27 19:52:17 -07:00
Victor Kyriazakos 2d1d65de46 fix(cron): accurate error for relay-fronted delivery with no live gateway (NS-773)
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
2026-08-27 19:52:17 -07:00
kshitijk4poor 42ac29eacc docs(cron): comment accuracy — booking is fail-open on probe errors; cross-ref status vocabulary
Review follow-up on the #93829 salvage: the block header said 'fail-closed'
while probe-error behavior deliberately keeps cron_complete (fail-open);
and the pathological-status tuple now cross-references the classifier's
vocabulary in hermes_state so drift is caught at the source.
2026-08-27 20:39:30 +05:30
liuhao1024 23f597a8f5 fix(cron): verify a persisted final assistant message before booking complete
The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).

Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.

Fixes #93820
2026-08-27 20:39:30 +05:30
kshitijk4poor cced6fa360 refactor(cron): extract _get_session_db_timeout alongside sibling timeout resolvers
Review follow-ups on the #96290 salvage:
- The inline env->config->default ladder was the third copy of the pattern;
  extract it next to _get_script_timeout/_get_media_send_timeout. Using
  load_config() (deep-merge) also removes the default-drift hazard flagged
  in review: cron.session_db_timeout_seconds now resolves from
  DEFAULT_CONFIG (config_defaults.py) instead of relying on the hardcoded
  10.0 staying in sync with it, and drops the distant-state coupling to
  run_job's raw _cfg local.
- Trim the relocated comment's stale claim about _submit_with_guard (at the
  new position the store init happens inside the guarded worker, not before
  it).
2026-08-27 19:07:15 +05:30
Heath Harris 3113f6056b fix(cron): open the session store only after wake-gate and validation early-returns
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.

Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.

Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
2026-08-27 19:07:15 +05:30
Ayush Nangia 2f0f01192d fix(cron): tree-kill script timeout descendants via agent.deadline.kill_process_tree
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.

The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.

Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
  lost" path orphaned setsid grandchildren the same way (whole-bug-class
  rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
  that exits right at the deadline doesn't log a spurious "no signal"
  warning (mirrors _terminate_cron_script_process); pinned by
  test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
  CI load could eat the whole 1s window before the spawner wrote its
  pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
  old path gave a 1s SIGTERM grace window; intended for a deadline-
  expiry hard stop (both docstrings say "hard stop")

Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.

Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
2026-08-27 14:28:06 +05:30
kshitijk4poor ded9470990 refactor(cron): fold simplify-review findings into mirror eligibility
- _target_mirror_eligible accepts a precomputed origin_match so the sole
  production caller stops re-resolving origin + re-running the origin
  match it computed one line earlier (tests keep the self-contained path).
- Document why the fallback branch restates _cron_mirror_delivery_enabled
  precedence (standalone correctness: per-job False must beat raw global
  True) instead of collapsing it to the call-site-coupled 'return True'.
- Retarget the stale in_channel warn branch from 'not origin_target' to
  'not inchannel_continuable' and reword it for the widened seed scope.
2026-08-26 16:06:09 +05:30
kshitijk4poor 8313185449 fix(cron): unify in_channel flatten and seed behind one continuable gate
Review finding: the thread-flatten stayed gated on origin_target while the
seed gained fallback/explicit eligibility — a threaded origin_fallback or
opted-in explicit target would deliver into the thread while the seed
created the flat session (the exact split-surface drift the flatten
comment warns about). One shared inchannel_continuable gate now drives
both, with _inchannel_seed_allowed folded in; is_dm_target hoisted above
the flatten and deduplicated.
2026-08-26 16:06:09 +05:30
Victor Kyriazakos 580daa7b96 fix(cron): mirror continuable-cron briefs for origin-fallback and opted-in explicit targets
A managed cron (created by a provisioning script, not from a live gateway
chat) never captures an origin. With cron.mirror_delivery: true and
deliver: origin, its brief was delivered to the home channel — the
user's own DM — but the transcript mirror and the in_channel session
seed were silently skipped: _target_matches_origin returns False for an
empty origin, and the whole continuable machinery keys off that check.
A user replying to the brief landed in a session with no record of it.
Field report 2026-08-17 (enterprise, Slack DM surface).

The June origin-scoping refactor (c06ceb3232) was written to exclude
broadcasts, and the exclusion is kept. What changes is the
classification: a home-channel FALLBACK for deliver=origin is the user's
primary conversation standing in for the origin, not a broadcast.

Changes:
- Delivery targets carry a resolution-provenance tag (_resolved_from:
  origin / origin_fallback / explicit; broadcast expansions untagged).
- _target_mirror_eligible replaces the bare origin check at the mirror
  gate: origin unchanged; origin_fallback eligible under the same flags
  as origin (per-job attach_to_session wins, else global
  cron.mirror_delivery); explicit platform:chat targets eligible ONLY
  under per-job attach_to_session — the global flag never activates
  them, so it cannot start writing transcript entries into arbitrary
  explicitly-addressed chats. 'all'/bare-platform stay never-eligible.
- Dedup OR-merges provenance so 'origin,all' resolving to the same chat
  keeps eligibility regardless of token order.
- _inchannel_seed_allowed guards the flat-session seed: group-channel
  session keys are user-isolated, so a seed without a user_id (origin-
  less job into a shared channel) would create an orphan session no
  reply resolves to — those targets fall back to the plain mirror. DM
  targets (keys don't embed user_id) always seed.
- cronjob tool schema text updated to describe the new attach scope.

Behavioral note: origin-less deliver=origin jobs under global
mirror_delivery now activate the full continuable path — on default
'thread' surface this opens a dedicated thread in the home channel
where the brief previously posted flat. That is the documented
continuable behavior; the silent flat post was the bug.

15 new tests (tests/cron/test_mirror_origin_fallback.py): eligibility
matrix (origin/fallback/explicit/all/bare/other-chat), dedup order
both ways, end-to-end mirror via _deliver_result for all four shapes,
origin regression control, seed user_id guard.
2026-08-26 16:06:09 +05:30
Teknium 1fe0f2f3ac feat(cron): import-error cron failures now name gateway code skew and the one-command fix (#95294 part 3)
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.

Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).

Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
2026-08-26 01:23:15 -07:00
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
kshitijk4poor 105999a0c9 refactor(gateway): unify computer-use repair call sites after review
- Make repair_explicit_computer_use_media_paths fail-open internally
  (cosmetic repair must never abort delivery); drop the cron-only
  try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
  no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
  gateway.media_repair seam; pre-existing tests keep pinning the
  gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
2026-08-25 13:06:43 +05:30
kshitijk4poor bb0d5503c2 fix(gateway): widen computer-use media path repair to sibling surfaces
Follow-up to the salvaged fix from PR #94439:

- Extract the repair into gateway/media_repair.py (shared module) and
  re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
  background tasks (_run_background_task_inner) and cron job delivery
  (cron/scheduler.py) — both call agent.run_conversation directly and
  never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
  content first instead of regex-scanning the raw string, which yielded a
  doubled-backslash path artifact and rewrote the response to a path the
  model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
  gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
  the compression-fallback last-user slice (incl. no-user fail-closed).
2026-08-25 13:06:43 +05:30
kshitijk4poor 0eda2ba0c8 fix: remove dead code, deduplicate error constants, fix skill key check
Follow-up to PR #92189 salvage:
- Remove unused job_no_agent_without_script() function (dead code)
- Replace inline NO_AGENT_WITHOUT_SCRIPT_ERROR string in _validate_job_mode_invariants with the constant
- Replace scheduler inline reason string with EMPTY_PAYLOAD_ERROR constant
- Add 'skill' (singular) to job_payload_is_empty 'in job' presence check
2026-08-24 15:47:42 +05:30
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Jack Lau dd03471858 fix(cron): nudge review of escaped-run failures too
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.

The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.

Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.

Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.

Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.

Fixes #88655
2026-08-22 03:19:08 +05:30
Teknium a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Adolanium fc9cbc872d fix(cron): do not load MEMORY.md into scheduled jobs
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.

Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
2026-08-21 13:24:43 +05:30
Victor Kyriazakos 4e1dd1a74b feat(cron): per-job reasoning_effort override in job definitions
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.

- cron/jobs.py: new optional job field, validated at the storage choke
  point against the canonical grammar via the shared
  hermes_constants.parse_reasoning_effort (spelling-only; capability
  clamping stays owned by the provider transports at send time, same as
  config-set effort). Empty string clears on update; invalid values
  raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
  > agent.reasoning_overrides > agent.reasoning_effort at fire time,
  after the auth-fallback model swap (the pin is model-independent by
  design). A stored value that no longer parses warns and falls back to
  config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
  (create and update), conditional key in _format_job, schema documents
  grammar/precedence/transport clamping/clear semantics. Agent-settable,
  unlike model/provider pins: it cannot redirect spend to a different
  model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.

Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
2026-08-20 19:56:14 -07:00