Commit Graph

82 Commits

Author SHA1 Message Date
Søren L. Hansen 0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
Teknium 37dcc0a6e8 fix(gateway): a secondary API_SERVER_KEY no longer skips the profile; start/install/status honour the live multiplexer
Under gateway.multiplex_profiles the default gateway serves every profile, yet
four startup/status paths still reasoned from the wrong source:

* A secondary profile's API_SERVER_KEY (which the docs REQUIRE for /p/<profile>/
  auth) auto-enabled api_server in that profile's config, so
  _load_secondary_profile_config raised SecondaryPortBindingConfigError and the
  whole profile was skipped. gateway/config_env.py::_enable_from_env now leaves
  `enabled` alone for port-binding platforms while a multiplexer loads a
  NON-default profile (home override + multiplex flag, the same signal
  gateway.config uses for scoped reads); the credential still lands in extra so
  the shared listener can authenticate the prefix. Default profile unchanged.

* "Is this profile served?" was re-derived from the default config.yaml plus
  GATEWAY_MULTIPLEX_PROFILES as seen by the CLI process. `hermes -p coder ...`
  loads coder's .env, so an env-only opt-in on the default profile was invisible
  (guard never fired, status said stopped) and an allowlist edit flipped the
  answer before the restart. named_profile_served_by_running_multiplexer now
  reads the pid-verified default gateway_state.json served_profiles (written by
  _record_served_profiles) first and falls back to config derivation only when
  the key is absent. The record helpers live in hermes_cli/gateway_multiplex_served.py.

* The served-profile guard ran only inside `gateway run`. `hermes -p X gateway
  start|install|restart` reached the service manager, whose unit then exited 78
  forever (systemd parks it while the CLI prints "started"; launchd KeepAlive
  respawns every 30 s). The service verbs now run the same guard up front
  (exit 78, same message) and accept --force; the Desktop /api/gateway/start
  route returns 409 for a served profile instead of spawning a doomed child.

* Status surfaces disagreed: `hermes -p coder status` said stopped, `hermes -p
  coder cron status` said "cron jobs will NOT fire" while `cron list` said fine,
  and the default `hermes status` never listed served profiles. Both now route
  through the probe / the recorded served set. The -p/--profile matcher in
  _scan_gateway_pids and gateway.status._command_line_belongs_to_profile compares
  the flag token for equality (`-p ops` no longer claims -- or lets `gateway
  stop` SIGTERM -- an `-p ops-2` gateway).

Docs: multi-profile-gateways.md now describes the start/install refusal, the
--force flags, the API_SERVER_KEY behaviour and the single default-home
gateway_state.json (the per-profile runtime_status.json claim was wrong).

Fixes #100397
Addresses #89726 #97360 #71344

(cherry picked from commit d002c1864a7b6a22c53758b16b7b0cc79aea2edf)
2026-09-11 15:27:23 -07:00
Teknium 5280fe9987 fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.

Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
2026-09-07 16:48:29 -07:00
Teknium 39ed610f8c feat(cron): create paused jobs without a scheduling race
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.

The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.

Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
2026-09-07 08:27:50 -07:00
Teknium eeb7671e69 simplify(compat): hermes_cli small facades — drop 7 re-exports/aliases (+relay_runtime alias module), repoint 12 callers/tests 2026-09-03 13:05:57 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium c25960666d refactor(hermes_cli): group D layout — join lone closing-bracket lines (AST-neutral) 2026-09-03 00:04:40 -07:00
Teknium 07617a8f92 refactor(hermes_cli): cron/debug/dashboard_procs — inline single-use locals, collapse try/pass to suppress, compact rationale comments (every WHY kept) 2026-09-02 23:58:03 -07:00
Teknium b7c5e3a205 refactor(hermes_cli): cron doctor summary line width 2026-09-02 23:03:28 -07:00
Teknium 02586dbed0 refactor(hermes_cli): cron incidents rows table, debug share flow tightening, copilot env-var loop invert 2026-09-02 22:58:35 -07:00
Teknium 7ff29f4837 refactor(hermes_cli): cron.py — split cron_list into _job_rows/_job_warnings/_last_run_display 2026-09-02 22:42:26 -07:00
Teknium 7ea4002495 refactor(hermes_cli): cron.py — lift manual-run verdict into _run_outcome, flatten cron_resume mode check 2026-09-02 22:40:25 -07:00
Teknium f1526d59d5 refactor(hermes_cli): group D — contextlib.suppress for swallow-only try/except blocks 2026-09-02 22:26:44 -07:00
Teknium 2cac493fec refactor(hermes_cli): group D — drop blank line after local imports (AST-neutral) 2026-09-02 22:10:21 -07:00
Teknium e4f8ab0e8e refactor(hermes_cli): group D — flatten notepad/doctor/lateness/lockfile branches, single-raise HTTP error mapping 2026-09-02 21:48:27 -07:00
Teknium 87d30c4701 refactor(hermes_cli): cron/debug — merge consecutive prints into single literals (byte-identical output) 2026-09-02 21:30:58 -07:00
Teknium ffe8b15a3e refactor(hermes_cli): group D — collapse doctor/edit/notepad/sweep branches, debug subcommand dispatch, wrap new long lines 2026-09-02 21:21:22 -07:00
Teknium 8f047aced9 refactor(hermes_cli): debug.py — lift tail-reader out of _capture_log_snapshot; cron status lock-probe collapse 2026-09-02 21:13:57 -07:00
Teknium c599b9ef9b refactor(hermes_cli): cron.py — table-driven job listing rows, state badges and detail lines 2026-09-02 20:59:24 -07:00
Teknium 7a643acfc9 refactor(hermes_cli): group D — merge fragmented string literals, hug call layouts (AST-neutral) 2026-09-02 20:51:45 -07:00
Teknium be5654efb2 refactor(hermes_cli): group D — AST-neutral bracket hugging / arg packing 2026-09-02 20:33:47 -07:00
Teknium e6f5f109f9 refactor(hermes_cli): cron.py — fold subcommand wrappers into dispatch table, dedupe unverified-target rendering, compact narrative comments 2026-09-02 20:12:37 -07:00
Teknium 2c7f5d12b5 refactor(hclib): cron/update/install lifecycle — cron status, update_* receipts and recovery, install repair, service manager 2026-09-02 14:45:25 -07:00
Victor Kyriazakos c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium 00a7115a02 fix(cron): make cron push-notify configurable (cron.delivery.notify) and surface UNVERIFIED live deliveries in cron list/doctor
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.

An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.

Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
2026-09-02 00:56:52 -07:00
Justin Wilson 8fd76fd1d6 fix(cron): surface delivery_failed instead of last_status ok
A successful agent run whose delivery failed used to persist
last_status=ok and bury the failure in last_delivery_error. CLI list
painted that as green and the run looked identical to a quiet success.

Record last_status=delivery_failed instead, keep last_delivery_error,
do not increment failure_streak, and teach cron list/doctor not to
treat it as ok.

Fixes #83993
2026-09-02 00:52:58 -07:00
Teknium 71c4bcf7af fix(cron): surface missed-fire catch-up lateness in hermes cron list / status
The catch-up machinery already re-ran jobs missed during gateway downtime,
but the late execution rendered as an ordinary on-time success — no
scheduled-vs-actual time, no lateness, no disposition (issue #99879, the
visibility half).

- Due-scan now persists a `last_dispatch` stamp on every recurring dispatch:
  scheduled_at, dispatched_at, lateness_seconds, and kind
  (on_time / late / catch_up, classified against the ticker tolerance and
  the schedule's catch-up grace window). Manual triggers and one-shots are
  not stamped (no scheduled instant to be late against / retired beyond
  grace).
- `hermes cron list` renders a per-job Dispatch line; late/catch-up runs
  show "⚠ catch-up after missed fire: scheduled ..., ran ... (31m late)".
- `hermes cron status` calls out jobs whose last dispatch was late or a
  catch-up, in both the built-in ticker and external provider paths.

CLI surface only — no new tools, no policy engine.

Addresses the visibility half of #99879.
2026-09-01 08:31:37 -07:00
Teknium f2f7a3bf15 feat(cron): doctor flags overdue next_run_at as silent non-firing
Widens the salvaged cron doctor with the highest-value fleet check:
an active job whose next_run_at is parked >15min in the past is not
firing (dead ticker, downed gateway, wedged fire-claim). Also registers
doctor in the docs (cron guide + CLI reference) and resolves the salvage
onto current main alongside runs/incidents/notepad.
2026-08-31 10:00:29 -07:00
joe102084 b028fe632e feat: add cron doctor health check 2026-08-31 10:00:29 -07:00
Jay. (neocode24) 9a7732b45f fix(cron): stale ticker yields its tick to a fresh gateway
A long-lived process whose checkout was updated underneath it (hot git
pull, interrupted hermes update) serves mixed sys.modules. When such a
stale process races a fresh gateway for the cron tick lock and wins the
minute, every agent job it dispatches can die on ImportErrors whose real
cause is staleness — and the fresh gateway's ticker skips the same minute
as lock-loser, so the user's scheduled job fires broken or not at all.

tick() now checks, BEFORE acquiring the tick lock:

  skew detected (boot fingerprint != disk revision)
    AND this process does not own the gateway runtime lock
    AND that lock is held (a fresh gateway is alive)
      -> raise CronTickYielded, skipping the tick entirely

Each arm alone keeps the old behavior:
- skew + self-owned lock -> proceed (delivery-path stale-code hint stays
  the surface for gateway-owned dispatches)
- skew + no lock holder -> proceed (desktop-standalone users must not
  lose their only ticker to a silent yield)
- skew None (non-git install, no boot fingerprint, probe failure) ->
  proceed; yielding is a certainty claim, never a guess

The yield RAISES instead of returning 0 so the provider loops record it
via record_ticker_error and mark the heartbeat success=False — a yielded
tick must not look like a healthy one (hermes cron status shows why),
mirroring the EMFILE propagation contract (#87644). Yield logging is
throttled to once per skew episode. Self-healing: when the fresh gateway
dies, its lock releases and the stale ticker's next tick proceeds.

Multiplex loop: a yield for one profile no longer cancels sibling
profiles' ticks in the same cycle; only the yielding profile records an
unsuccessful beat.

gateway/status.py gains owns_gateway_runtime_lock() —
is_gateway_runtime_lock_active() is True for the lock's own owner too, so
a caller deciding whether to yield to a FRESH gateway needs the
in-process handle as the discriminator.
2026-08-31 09:59:07 -07:00
686f6c61 e703717513 fix(cron): treat a live multiplexer as gateway-alive for satellite profiles
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
2026-08-31 07:28:39 -07:00
kokhlo 82d7a13003 fix(cron): profile isolation — systemd filter + heartbeat guard
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
  PID: 4165
  Ticker heartbeat: 39s ago

  4 active job(s)
  Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.

Root causes:
1.  executed ,
   which enumerated the entire systemd fleet regardless of ,
   violating the docstring "only PIDs belonging to the current profile".
2.  checked  when
   (no heartbeat file) should trigger a warning — instead, it fell through
   to the "✓ Gateway is running" green branch.

Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
  when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
  explicit yellow warning: "ticker has not reported a heartbeat".

Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
2026-08-31 06:02:32 -07:00
kshitijk4poor 1e5fb70fb5 fix(cron): widen the lock-first liveness check to 'hermes cron status'
Sibling site of the salvaged #95947 fix (same file): cron_status
declared 'Gateway is not running — cron jobs will NOT fire' from a bare
find_gateway_pids() miss even while the runtime lock proved the gateway
alive. Now the not-running verdict requires both the scan AND the lock
to read dead; when only the lock answers, the pid line falls back to
the recorded gateway pid (or is omitted).

Two regression tests pin the false-alarm suppression and the genuine
not-running warning.
2026-08-27 16:10:57 +05:30
kshitijk4poor f2d043eb08 test(cron): pin lock-first liveness + harden lock-probe failure
Follow-ups to the salvaged #95947 cron commit:

- Wrap the lock probe in its own try/except: a crashing probe is
  'unknown', not 'dead' — the pid scan still decides instead of the
  whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
  held + empty pid scan -> alive (the reported false alarm); lock
  inactive -> pid-scan fallback both ways; crashing lock probe still
  falls back.
- patch_liveness now pins the lock probe inactive by default so the
  pre-existing pid-scan tests stay deterministic on machines where a
  real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-27 16:10:57 +05:30
Magnus Lundstedt b71c3cc13c fix(cron): trust the gateway runtime lock for builtin-ticker liveness
_builtin_gateway_liveness() decides whether the builtin cron ticker can fire by
PID-scanning via find_gateway_pids(). That scan can transiently return empty even
while the gateway is up (e.g. just after a restart), so the in-gateway cronjob tool
emits a false 'Gateway is not running — jobs won't fire' while jobs are firing on time.

Prefer the gateway runtime lock: it is held for exactly the gateway's lifetime (a
reliable liveness signal), and inside the gateway process it short-circuits to True,
so the in-gateway check can never false-alarm. Fall back to the PID scan only when the
lock reads inactive (the external-CLI path).
2026-08-27 16:10:57 +05:30
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
kshitijk4poor ddbd928ee4 refactor(cron): parameterize liveness warning plurality, drop fragile string replace
Simplify-pass follow-ups on the #87033 fix:
- _gateway_liveness_notice(plural=) authors both wording variants at one
  site; removes the exact-substring .replace() that would silently no-op
  if the create-path text is ever edited.
- Collapse the operator-precedence-trap conditional in list to a plain
  'if jobs' — an empty list has nothing inert and now skips the probe.
- Fix docstring/code mismatch (builder returns gateway_running: True on
  the happy path) and drop the dead try/except in
  _warn_if_gateway_not_running (the helper never raises).
2026-08-24 16:38:47 +05:30
kshitijk4poor 5843b2f595 fix(cron): share liveness helper with CLI and extend it to cronjob list
Follow-ups for the salvaged #93098:
- Move the tri-state liveness heuristic into hermes_cli.cron
  (_builtin_gateway_liveness) so the CLI warning and the cronjob tool
  share one implementation instead of two drifting copies.
- Surface gateway_running/warning on the list action too — an agent
  inspecting jobs in a gateway-less environment has the same silent-
  inert-job failure mode (#87033) as create. Empty lists stay quiet.
2026-08-24 16:38:47 +05:30
leosiedler a0ca7c1920 feat(cron): add explicit one-shot re-arm 2026-08-24 00:34:26 -07:00
Victor Kyriazakos 4e1dd1a74b feat(cron): per-job reasoning_effort override in job definitions
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.

- cron/jobs.py: new optional job field, validated at the storage choke
  point against the canonical grammar via the shared
  hermes_constants.parse_reasoning_effort (spelling-only; capability
  clamping stays owned by the provider transports at send time, same as
  config-set effort). Empty string clears on update; invalid values
  raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
  > agent.reasoning_overrides > agent.reasoning_effort at fire time,
  after the auth-fallback model swap (the pin is model-independent by
  design). A stored value that no longer parses warns and falls back to
  config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
  (create and update), conditional key in _format_job, schema documents
  grammar/precedence/transport clamping/clear semantics. Agent-settable,
  unlike model/provider pins: it cannot redirect spend to a different
  model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.

Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
2026-08-20 19:56:14 -07:00
Teknium cb1b1da219 fix: surface missed cron fires as last_fire_error on the job record
On hosted deployments a scheduled fire that cannot be forwarded to the
gateway api_server (dead 8642 listener, gateway down) was invisible
outside gui.log: no execution row is created because the claim never
happens, so `cronjob list` showed a healthy job that silently missed
days of scheduled runs (4 consecutive nightly misses in the field,
diagnosed only by log grep).

Changes:
- cron/jobs.py: note_fire_forward_failure() durably stamps
  last_fire_error ({at, detail}) on the job record; mark_job_run clears
  it on the next successful run so it always describes current
  auto-fire health (mirrors preflight_alerted/drift_alerted).
- hermes_cli/web_routers/cron.py: the dashboard fire webhook stamps the
  job on the gateway-unreachable path, best-effort (never disturbs the
  503/Retry-After retry contract or the OOF-266 intentional-stop drop).
- tools/cronjob_tools.py: _format_job carries last_fire_error so the
  agent-facing cronjob list surfaces it.
- hermes_cli/cron.py: `hermes cron list` prints a red
  "Missed scheduled fire" line.
- web/: dashboard CronPage renders the miss; api.ts type updated.
- gateway/run.py: one-time startup warning when an external cron
  provider is active but the api_server adapter is not running (the
  fire path is dead-on-arrival; most common cause is API_SERVER_KEY
  missing from an unsupervised gateway relaunch).
- website/docs: cron doc section on missed fires.
2026-08-17 11:29:10 -07:00
kshitij 0a8e703701 refactor(cron): dedup EMFILE tick-failure handling; share the fd-exhaustion text matcher
/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
  loops (_start and _start_multiplex) along with duplicated function-local
  imports — extracted _note_tick_failure() next to _backoff_wait_seconds
  so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
  _is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
  _is_fd_exhaustion_text() out and use it from both.

11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
2026-08-17 16:55:25 +05:30
kshitij 80dc1836c4 fix: single-owner fd reclamation, shared backoff helper, clean CLI tick failure
Follow-ups on the #87796 salvage:

- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
  lock-failure raise site — the ticker loop's except handler already runs
  reclamation once per failed tick, so the raise-site call doubled the
  gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
  duplicated verbatim in start() and _start_multiplex() into a module-level
  _backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
  cleanly (exit 1) instead of dumping a traceback — tick() raising on real
  lock-acquisition failures is new behavior from this fix.
2026-08-17 16:55:25 +05:30
webtecnica 815934ae52 fix(cron): scheduler self-heals after EMFILE instead of stalling silently (#87644)
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.

- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
2026-08-17 16:55:25 +05:30
Teknium ea29702749 feat(cron): --continuity / --no-continuity flags on hermes cron create/edit
CLI parity for the continuity toggle:

- subcommands/cron.py: --continuity on create; --continuity / --no-continuity
  tri-state pair on edit (same store_const pattern as --no-agent/--agent)
- cron.py: forwarded to the cronjob tool; created/edited job summaries print
  a "Continuity: on" line
- cronjob_tools._format_job: reports continuity as an explicit boolean and
  strips the reserved 'self' entry from the reported context_from list
- cron-job.ts: form reader accepts both shapes (raw store record with 'self'
  inside context_from, or formatted record with the explicit flag)
- docs: CLI flag examples in the continuity section

E2E (real argparse -> cron_create/cron_edit -> jobs.json in temp HERMES_HOME):
create --continuity stores ['self']; edit --no-continuity clears; edit
--continuity restores; default-off unchanged. 91 cron/tool tests + 16 CLI
cron tests + vitest 10/10 pass.
2026-08-16 22:09:28 -07:00
Teknium 07a5179158 Inspired by Poke: nudge review of repeatedly-failing recurring cron jobs
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.

- cron/jobs.py: persist a failure_streak counter in mark_job_run —
  incremented on agent failure, reset on success; delivery failures don't
  count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
  delivered failure summary once a recurring job's streak reaches
  cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
  nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
  failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.

Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
2026-08-16 22:08:59 -07:00
Teknium 0fc2a10d82 fix(cron): stop one-shot CLI cron run from orphaning the job; reap dead-owner claims on tick
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.

Two-part fix:

1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
   channel stateless (scoped ContextVar set/reset around the call) before
   invoking the cron API, so `async_delivery_supported()` gates off
   `_try_dispatch_background_run` and the run executes synchronously to
   completion in the CLI process — the same behavior `hermes -z` already
   gets via declare_stateless_channel().

2. cron/scheduler.py: tick() now periodically invokes
   recover_interrupted_executions() (previously only run at scheduler
   startup), so execution rows whose exact owner process is provably dead
   (pid + process start time check in _owner_is_live) are reaped to
   'unknown' by the long-lived gateway ticker without a restart.
   Throttled to once per 300s so idle 60s ticks don't pay a ledger
   connection every cycle.

Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.

Fixes #86721
2026-08-15 02:44:14 -07:00
webtecnica 2d81236f7f fix(cli): report background-dispatch cron runs without false 'failed' (#83340) 2026-08-14 21:55:14 -07:00
rjvandeve c7a5de7d6e fix(cron): make pause authoritative against half-paused records
pause_job already sets enabled=false atomically with state/paused_at, but
get_due_jobs only checked enabled — so a contradictory record
(enabled=true + paused_at/state=paused) still fired. That was the 07-30
outage failure mode: list looked frozen, fleet kept merging.

- is_job_runnable / effective_job_state: pause markers gate fire; display
  derives from the scheduler-honoured enabled flag so half-paused never
  renders as [paused]
- get_due_jobs self-heals enabled=false + logs error on contradiction
- claim_job_for_fire uses is_job_runnable (paused_at counts too)
- list/format paths use effective_job_state
- behavioural tests: pause blocks due fire; half-pause self-disables
2026-08-08 13:48:00 +05:30
Teknium 04e8a661f2 feat(cron): per-job durable notepad — KV scratchpad surviving scheduled runs
- cron/notepad.py: SQLite-backed cron_notepad(job_id, key, value,
  updated_at) store in its own profile-local db (cron/notepad.db),
  following the executions.py connection/transaction pattern. APIs:
  set_note/get_note/delete_note/list_notes/clear_notepad +
  render_notepad_section. Documented size caps: 16KB per value,
  128-char keys, 64KB per job total; oversized writes raise ValueError.
- cron/scheduler.py: inject non-empty notepads into the job prompt at
  the context_from data-injection seam as a clearly-labeled
  "Job notepad (persistent across runs)" section that also documents
  the CLI write path. Empty notepad renders "" — byte-stable prompts
  for jobs that never use the feature.
- hermes_cli/cron.py + hermes_cli/subcommands/cron.py:
  `hermes cron notepad <job_id> [get|set|delete|list]` under the
  existing cron subcommand tree (no new top-level command, no new
  model tool — the agent writes via terminal + CLI).
- tests/cron/test_notepad.py: CRUD, durability, cap enforcement,
  prompt injection, byte-stable empty case, read-failure resilience,
  CLI handler + dispatch (TDD; watched fail first).

Inspired by: Amp (Sourcegraph) cron notepad (idea-level, proprietary —
zero code).
2026-08-07 08:57:48 -07:00