Commit Graph

389 Commits

Author SHA1 Message Date
Teknium 1fe0f2f3ac feat(cron): import-error cron failures now name gateway code skew and the one-command fix (#95294 part 3)
When an agent cron job dies with an import-class error (cannot import
name / ModuleNotFoundError / ImportError), the failure summarizer — which
runs inside the gateway process — now consults gateway.code_skew: if the
process booted on a different revision than disk HEAD, the delivered
message appends 'gateway is running stale code (booted on X, disk is at
Y) — run hermes gateway restart'. Turns the reported two-day mystery
(15 missed jobs, identical ImportError, no explanation) into a one-line
fix instruction on the first failure.

Fail-safe by construction: skew detection returns None on non-git
installs and processes without a boot fingerprint, the probe seam
swallows every exception, and no_agent script jobs (fresh subprocess,
consistent imports) fall through to the generic cleaner — their
ImportErrors are the script's own problem, and blaming gateway skew
there would send the reader to the wrong place (same mode-gating as the
provider branches).

Reuses gateway/code_skew.py (the /model-switch skew detector) rather
than adding a second fingerprint reader.
2026-08-26 01:23:15 -07:00
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
kshitijk4poor 105999a0c9 refactor(gateway): unify computer-use repair call sites after review
- Make repair_explicit_computer_use_media_paths fail-open internally
  (cosmetic repair must never abort delivery); drop the cron-only
  try/except so all three call sites are identical one-liners.
- Drop cron's redundant 'MEDIA:' pre-check (helper early-returns).
- Document the intentional lazy BasePlatformAdapter import (verified:
  no cycle either way; keeps module import cheap for cron processes).
- Point the two new regression tests at the canonical
  gateway.media_repair seam; pre-existing tests keep pinning the
  gateway.run re-export shim.
- Docstring: matching is case-insensitive, say so.
2026-08-25 13:06:43 +05:30
kshitijk4poor bb0d5503c2 fix(gateway): widen computer-use media path repair to sibling surfaces
Follow-up to the salvaged fix from PR #94439:

- Extract the repair into gateway/media_repair.py (shared module) and
  re-export under the historical private name in gateway/run.py.
- Wire the repair into the two bypassed delivery surfaces: gateway
  background tasks (_run_background_task_inner) and cron job delivery
  (cron/scheduler.py) — both call agent.run_conversation directly and
  never pass the main turn chokepoint.
- Fail closed on malformed/truncated JSON tool results: parse JSON-looking
  content first instead of regex-scanning the raw string, which yielded a
  doubled-backslash path artifact and rewrote the response to a path the
  model never wrote.
- Deduplicate the tool_name_by_call_id builder (three verbatim copies in
  gateway/run.py) into the shared module; hoist the abs-path prefix regex.
- Add regression tests: malformed-JSON fail-closed (mutation-checked) and
  the compression-fallback last-user slice (incl. no-user fail-closed).
2026-08-25 13:06:43 +05:30
kshitijk4poor 0eda2ba0c8 fix: remove dead code, deduplicate error constants, fix skill key check
Follow-up to PR #92189 salvage:
- Remove unused job_no_agent_without_script() function (dead code)
- Replace inline NO_AGENT_WITHOUT_SCRIPT_ERROR string in _validate_job_mode_invariants with the constant
- Replace scheduler inline reason string with EMPTY_PAYLOAD_ERROR constant
- Add 'skill' (singular) to job_payload_is_empty 'in job' presence check
2026-08-24 15:47:42 +05:30
cycorld 350fb975b9 fix(cron): prevent empty payload loop and protect against blank name overwrite
- Reject cron jobs with empty runnable payload (blank prompt, no script, no skills) on create and update
- Auto-pause legacy unrunnable jobs at schedule time to prevent infinite fire loops
- Prevent blank name string in cron update tool from unintentionally wiping job names
- Add comprehensive test coverage (34 tests)
2026-08-24 15:47:42 +05:30
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Jack Lau dd03471858 fix(cron): nudge review of escaped-run failures too
A recurring job that fails at the scheduler layer - an exception escaping
run_one_job's body before the agent is ever constructed - has delivered a
failure alert since 4668750fa. It has never carried the repeated-failure
review nudge the normal agent-failure delivery carries: the nudge (#80752,
2026-08-06) predates that second delivery site by eight days and only ever
composed the first one.

The streak itself is layer-agnostic. mark_job_run increments failure_streak
for an escaped failure exactly as it does for an agent failure, and the
escape handler calls it. So the counter climbs correctly and shows up in
`hermes cron list`, but the chat message that spends it is unreachable for a
job whose failures ALL escape - a half-applied update leaving a bad import,
a provider client that cannot construct. Those are precisely the failures
that repeat identically on every tick, so the operator gets the same one-line
error every 10 minutes indefinitely and is never told the automation itself
is worth reviewing or pausing.

Compose the nudge at the escape handler's delivery exactly as the normal
path does. It stays config-gated and threshold-gated by the same helper, so
a first-time escaped failure reads exactly as it did before.

Docs said the streak counts "runs where the agent failed", which is what the
reporter read and reasonably concluded their failures were out of scope. The
counter never worked that way; correct the sentence to match the code.

Tests: two cases on the escaped-failure delivery path - streak at threshold
appends the nudge (fails on the unfixed handler with the bare summary), and
streak below threshold delivers the unchanged one-liner, so the guard also
proves the nudge is not unconditional. The existing nudge tests only ever
exercised the helper in isolation, which is why the second delivery site
could be added without it.

Fixes #88655
2026-08-22 03:19:08 +05:30
Teknium a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Adolanium fc9cbc872d fix(cron): do not load MEMORY.md into scheduled jobs
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.

Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
2026-08-21 13:24:43 +05:30
Victor Kyriazakos 4e1dd1a74b feat(cron): per-job reasoning_effort override in job definitions
A cron job can now pin its own reasoning (thinking) effort, independent
of the global agent.reasoning_effort and per-model reasoning_overrides.
Heavy scheduled analyses can run at high while cheap recurring jobs run
at minimal, without touching the fleet-wide default.

- cron/jobs.py: new optional job field, validated at the storage choke
  point against the canonical grammar via the shared
  hermes_constants.parse_reasoning_effort (spelling-only; capability
  clamping stays owned by the provider transports at send time, same as
  config-set effort). Empty string clears on update; invalid values
  raise ValueError before anything persists. Not a drift-guard axis.
- cron/scheduler.py: _resolve_job_reasoning_config resolves per-job pin
  > agent.reasoning_overrides > agent.reasoning_effort at fire time,
  after the auth-fallback model swap (the pin is model-independent by
  design). A stored value that no longer parses warns and falls back to
  config resolution instead of killing the tick.
- tools/cronjob_tools.py: reasoning_effort on BOTH mutation verbs
  (create and update), conditional key in _format_job, schema documents
  grammar/precedence/transport clamping/clear semantics. Agent-settable,
  unlike model/provider pins: it cannot redirect spend to a different
  model.
- hermes cron create/edit --reasoning-effort (empty string clears).
- Docs: cron feature page tip + CLI reference rows.

Tests: tests/cron/test_cron_reasoning_effort.py (32) — store contract,
scheduler precedence incl. byte-identical absent-field behavior and
garbage fallback, tool create/update/clear/error paths, schema surface.
2026-08-20 19:56:14 -07:00
Chris Fontes 5046282867 feat(config): resolve_turn_limit — first-class 'none'/'unlimited' for agent.max_turns
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.

This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:

  - int/float → int(raw) (floats truncated)
  - numeric string ('120') → int(raw)
  - 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
    whitespace-tolerant) → sys.maxsize sentinel
  - YAML None/null → default (90)
  - bool/list/dict/garbage → default (with debug log)

All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.

The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.

Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
2026-08-20 04:50:39 -07:00
Ben Barclay 4308c453ac fix(cron): stamp persisted origin scope_id onto origin-matching delivery metadata
The seed-key fix made the SESSION scoped, but the delivery leg still
dropped the scope: cron route_metadata carried only job_id (+thread),
DeliveryRouter stamps scope_id only for the configured HOME channel, and
the RelayAdapter's per-chat scope cache is cold after a gateway restart
(learned from inbound only). A scoped Slack origin that is not the home
chat therefore egressed with NO tenant discriminator, and the connector's
fail-closed guard could reject the brief before delivery — the
delivery-leg sibling of the seed-key scope gap.

Copy origin.scope_id into the live text and media routing metadata for
ORIGIN-MATCHING targets only (setdefault — never overrides router/home
stamping). Fan-out/broadcast targets are excluded by the origin gate: a
fan-out target's tenant is not the origin's, and a wrong scope is worse
than none.

Tests: restart-shaped positive (scoped non-home origin -> scope_id on
routed metadata, RED before this fix) and legacy negative (scope-less
origin stamps nothing).
2026-08-20 20:42:20 +10:00
Ben Barclay 6cd1ed2e78 test(relay): rename misnamed precedence test; document the flat-key fallback nuance
test_flat_key_wins_over_subblock asserted the OPPOSITE of its name (the
sub-block wins, matching _relay_slack_extra). Rename to what it proves.
Also note in _resolve_cron_surface_mode why its fallback differs from
_relay_slack_extra's all-or-nothing sub-dict: the flat key is the legacy
staging shape, and a flat knob applies to every fronted platform, gated
only by the per-platform D6 capability check.
2026-08-20 20:13:29 +10:00
Ben Barclay a78af23cca docs(cron): document the in_channel carve-out on the mirror opt-in
_cron_mirror_delivery_enabled still promised 'cron deliveries live only
in the cron job's own session' as the unconditional default, but the
in_channel continuable surface now seeds the target session regardless
of attach_to_session/cron.mirror_delivery (the seed IS the continuation
feature, and in_channel is itself opt-in). State the carve-out where the
guarantee is documented.
2026-08-20 20:13:10 +10:00
Ben Barclay 162b23c3e2 fix(relay): D6 in_channel capability gate resolves the destination platform's descriptor
RelayAdapter.supports_inchannel_continuable is a scalar adopted from the
PRIMARY identity's handshake descriptor, but one RelayAdapter fronts N
platforms and the connector advertises the bit per platform. Reading the
scalar for every logical platform both leaked a Slack-primary True onto
other fronted platforms (activating the flat surface their descriptor
never advertised) and suppressed a non-primary platform's advertised
True (forcing thread mode on capable Slack behind a Discord primary).

Add supports_inchannel_continuable_for_platform(platform): resolves the
platform's own negotiated descriptor via descriptor_for_platform (the
same Phase 1.5 seam max_message_length uses), scalar fallback only when
the per-platform descriptor is unavailable. The scheduler's D6 gate
prefers the query when the adapter provides it; native adapters keep
the class-attribute path byte-identically.

Tests: two-platform descriptor matrix (primary-True no-leak,
non-primary-True honored, unknown-platform scalar fallback).
2026-08-20 20:12:52 +10:00
Ben Barclay 20c56f82b8 fix(cron): in_channel thread-flatten uses the seed's gate (origin_target)
The seed was decoupled from the mirror opt-in (in_channel is the
continuation surface regardless of attach_to_session), but the
thread-id-clearing gate above it still read mirror_this_target. With the
advertised default config (attach_to_session=false, cron.mirror_delivery
unset) and an origin carrying a real thread_id, the brief kept delivering
INTO the origin thread while the flat (thread_id=None) session got
seeded — brief and continuation surface in different places, so a plain
reply never saw it.

Flatten on the same gate as the seed: origin_target (with the existing
live_adapter_ready guard). Fan-out/broadcast targets are unaffected.

Test drives _deliver_result with a thread-carrying origin and default
knobs, asserting on the routed DeliveryTarget.thread_id — RED on the old
gate, GREEN now.
2026-08-20 20:07:44 +10:00
Ben Barclay 8b6cf434cb fix(cron): carry Slack workspace scope_id into continuable seed keys
build_session_key embeds the workspace segment (scope_id) in every Slack
dm/group/thread key, but both cron seed helpers built their SessionSource
without it: the seeded row keyed agent:main:slack:dm:<chat>:<thread> while
a real scoped reply keys agent:main:slack:dm:<team>:<chat>:<thread> — a
row no reply ever resolves to. DMs were rescued only incidentally by the
legacy-key claim-once migration; scoped channels/threads got continuation
amnesia, and identical channel ids in two workspaces could collide.

Capture HERMES_SESSION_SCOPE_ID into the cron origin (_origin_from_env —
the session-context var async_delegation already snapshots), add scope_id
to _seed_cron_thread_session/_seed_cron_channel_session, and pass the
origin's scope at all three seed call sites.

Tests: scoped dm-thread / channel-thread / flat-channel seed-vs-reply key
equality through the real build_session_key, plus a two-workspace
non-collision guard.
2026-08-20 20:04:49 +10:00
Victor Kyriazakos fbf5eb8dd9 fix(cron): DM cron thread seed keys through the DM arm — thread-typed seed row never matched the DM reply's key
Live incident (Alice canary 2026-08-20, job 8e21a957b77b): the continuable
thread seed created its session with chat_type='thread', but a Slack DM
in-thread reply arrives chat_type='dm' and build_session_key routes DM
threads through the DM arm (...:dm:<chat>:<thread>). Seed row and reply row
never matched — the reply had no brief in context (continuation amnesia).

is_dm on _seed_cron_thread_session selects the seeded chat_type at both call
sites (opened-thread and the companion in_channel thread seed); channel
threads are unchanged. Sibling lane of the flat seed's is_dm fix. Tests pin
the key-equality contract: seeded key == the key the reply builds.
2026-08-20 01:40:47 +00:00
Victor Kyriazakos 1d86dccad1 fix(cron): seed continuable delivery independently of mirror opt-in
Continuability is explicit via attach_to_session or cron.mirror_delivery;
those knobs control transcript mirroring for the ordinary thread/default
surface. However, the in_channel surface must still receive its delivery
text to seed the continuation session. Previously mirror_text was populated
only when the optional mirror knob was enabled, so in_channel jobs created
with the default false settings passed an empty string to the seed helper,
which returned False. The live symptom was a delivered cron message with no
continuation context; Alice reproduced it three times (latest ef7bd2869d15).

Keep cleaned delivery text available for continuable surface seeding while
retaining mirror_enabled for the separate _maybe_mirror_cron_delivery path.
Targeted cron/in_channel regression suite: 1008 passed, 1 skipped.
2026-08-20 01:40:23 +00:00
Victor Kyriazakos 4673836f79 fix(cron): deterministic in_channel seed + companion thread-surface seed
Two live failures from the Alice canary (2026-08-19, jobs 28a24afebd81 /
83b93f8be379), both leaving a continuable in_channel cron with amnesia:

1. Seed mirrored via origin heuristics and silently dropped the brief.
   _seed_cron_channel_session created the flat session row, then
   mirror_to_session RE-DISCOVERED the target via find_session_by_origin —
   whose multi-candidate bail-out returns None on a populated chat (flat
   session + N per-message thread sessions sharing one chat_id, mixed
   user_ids). Receipt: 'in_channel seed did NOT land on slack:D0BJTDCSR7C'.
   mirror_to_session now accepts an explicit session_id and both cron seeds
   pass the exact row they just created; origin-scan remains the fallback
   for callers that genuinely don't know the target.

2. The brief's OWN THREAD was never seeded. in_channel delivers flat, but
   a flat Slack message still invites a thread reply (the natural mobile
   affordance — exactly what the user did). That reply keys to
   (chat, thread=<brief ts>), which no seed touched. The delivery's
   message_id now anchors a companion _seed_cron_thread_session so BOTH
   reply surfaces (plain channel message AND in-thread reply) continue the
   job.

Also: thread-seed failures upgraded debug→WARNING (silent seed failure IS
the user-facing bug), and the thread seed reports landed/not-landed.

Regression tests drive both against the live failure shapes: exact-session
mirror asserted via session_id kwarg; thread companion asserted via the
SendResult message_id anchor. Clean-fixture blind spot noted: the E2E
harness used a fresh store with one row, which is why heuristic rediscovery
looked fine pre-production.
2026-08-20 01:40:23 +00:00
Victor Kyriazakos d46b3533d4 chore(cron): loud diagnostics on the in_channel seed path
Seed failure was logger.debug — invisible in production while being the
exact 'agent has no idea about its own brief' symptom. WARNING on: seed
exception (with reason), seed returning False, and in_channel delivery to
a non-origin target.
2026-08-20 01:40:23 +00:00
Victor Kyriazakos 3c52d3589f fix(cron): in_channel seed must not require the attach_to_session mirror opt-in
Live regression (Alice, 2026-08-19): a continuable cron with
cron_continuable_surface=in_channel delivered its brief flat, but the
flat-session seed was gated on mirror_this_target = mirror_enabled AND
origin-match. Without attach_to_session (and with cron.mirror_delivery
defaulting False) the seed never ran; the next plain reply resolved to a
blank (slack, chat, None) session and the agent had no idea about its own
delivery message — not continuable in channel OR thread (in_channel mode
correctly skips thread creation, so there was no thread session either).

in_channel IS the continuation surface, not a mirror nicety: gate the seed
on origin-match alone, and resolve origin_user_id for any origin-matching
target so the seeded key still carries the scheduling user on
per-user-isolated chats. attach_to_session remains the opt-in for the
separate default-surface mirror behavior.

Regression test drives the delivery path with attach_to_session=False and
asserts the seed fires with the right user_id (fails on pre-fix code).
2026-08-20 01:40:23 +00:00
Victor Kyriazakos 85b89451f6 feat(relay): flat in_channel continuable cron surface on the relay lane
Field report (enterprise side-by-side, 2026-08-18, finding 1 — the relay-only
blocker): on relay-fronted Slack, cron briefs always deliver into a dedicated
thread; the flat continuable surface (cron_continuable_surface: in_channel)
that native Slack supports is inert, so plain DM replies never continue the
job and the main conversation never sees the brief.

Three gaps closed:
- CapabilityDescriptor gains supports_inchannel_continuable (default False,
  additive within contract_version 1; from_json ignores it from old
  connectors, old gateways filter it as unknown). The connector advertises
  it per platform at handshake.
- RelayAdapter maps the bit onto the adapter capability surface in both the
  constructor and _apply_descriptor (renegotiation), so the scheduler's D6
  fail-safe gate sees it exactly like native Slack's class attribute.
- _resolve_cron_surface_mode replaces the scheduler's inline flat-key read:
  native keeps the shipped flat shape; the relay lane reads the same
  per-logical-platform sub-block as the documented relay Slack knobs
  (platforms.relay.extra.slack.cron_continuable_surface), sub-block wins,
  scoped so a slack block cannot leak onto other fronted platforms.

The seed path needs no changes: RelayAdapter inherits set_session_store
(wired by the generic adapter boot loop) and _seed_cron_channel_session
keys the flat session off the logical platform_name.

12 new tests: descriptor default/from_json/legacy-absence, adapter mapping
constructor + renegotiation, and the surface-knob matrix (native flat key,
relay sub-block, per-platform scoping, precedence, defaults).
2026-08-19 13:44:27 +00:00
Teknium bc76f62c20 feat(cron): configurable media-send timeout + non-empty failure reasons
Follow-up on the salvaged commits from PRs #87965 and #87967
(@AiwendilInTheWoods):

- Promote the media-send timeout to the standard resolution pattern:
  HERMES_CRON_MEDIA_SEND_TIMEOUT env var, then
  cron.media_send_timeout_seconds in config.yaml, then 300s default
  (mirrors script_timeout_seconds; .env stays secrets-only).
- Register the config key in DEFAULT_CONFIG and document both surfaces
  (environment-variables reference + cron user guide).
- Fold the empty-str() exception fallback into the error string recorded
  in delivery_errors (post-#88631 the reason reaches the run status, not
  just the log line).
- Tests: timeout resolution precedence + TimeoutError reason fallback.
2026-08-17 17:51:08 -07:00
AiwendilInTheWoods d9ec9dd3fd feat(cron): make the media-send timeout configurable
The media delivery path used a hardcoded future.result(timeout=30).
Large attachments legitimately exceed it with no way to raise the limit.
Read HERMES_CRON_MEDIA_SEND_TIMEOUT, matching the existing
HERMES_CRON_SCRIPT_TIMEOUT / HERMES_CRON_TIMEOUT /
HERMES_CRON_SESSION_DB_TIMEOUT convention in the same module.
2026-08-17 17:51:08 -07:00
AiwendilInTheWoods 64b7c96e45 fix(cron): media-send failure logs an empty reason on timeout
TimeoutError carries no message and str(TimeoutError()) is the empty
string, so the media-send warning rendered with nothing after the colon.
Fall back to the exception class name when str(e) is empty.
2026-08-17 17:51:08 -07:00
Victor Kyriazakos 22f0f22298 fix(cron): manual runs no longer silently drop media attachments
Field report (enterprise, v0.20.0): cron jobs delivering text + PDF/image
attachments to Slack DMs deliver both on scheduled ticks but text-only on
manual `hermes cron run <job-id>`. Same box, same token, same scopes —
the divergence is process context and error visibility, not credentials.

Three defects, one bug class (attachment failures invisible + policy
divergence between the gateway process and standalone processes):

1. Standalone lane swallowed warnings: platform standalone senders
   (Slack files_upload_v2, Discord, ...) report per-file upload failures
   in result['warnings'] while returning success=True for the delivered
   text leg. _deliver_result only read result['error'], so the run was
   marked ok and the attachment vanished without a trace. Warnings now
   surface into delivery_errors (and the job's last_error).

2. Live-adapter lane swallowed media failures: _send_media_via_adapter
   logged failures at WARNING and returned None. It now returns per-file
   error strings and _deliver_result records them — text-delivered-but-
   attachment-failed is a visible partial failure on BOTH lanes.

3. Media-policy env bridge was gateway-only: gateway.strict /
   media_delivery_allow_dirs / trust_recent_files were translated from
   config.yaml to the env vars validate_media_delivery_path reads ONLY in
   gateway startup. A CLI-process manual run filtered attachment paths
   under a different policy — in strict/allowlisted deployments the exact
   reported symptom (scheduled delivers, manual drops, silently). The
   translation now lives in gateway/media_policy.apply_media_policy_env
   (idempotent, env-wins, never raises); gateway startup delegates to it
   and _deliver_result applies it before filtering. Attachments dropped
   by the policy filter are also reported in the run status instead of
   only a stderr WARNING.

On v0.20.0 specifically the failure was double-blind: the pre-9cf2cbd382
isinstance(resp, dict) gates meant upload failures were undetectable in
the sender AND unsurfaced by the scheduler. 9cf2cbd382 (in 2026.8.13)
fixed detection; this fixes visibility and policy parity.

8 new tests (tests/cron/test_media_delivery_parity.py): warnings→errors,
clean-delivery control, media-reaches-sender control, live-adapter
failure/dropped-path reporting, bridge helper semantics, strict+allowlist
end-to-end in a non-gateway process, and the .env-strict/config-allowlist
split that reproduces the field symptom. Mutation check: disabling the
warnings loop and the bridge fails exactly the 2 guarding tests.
2026-08-17 17:19:50 -07:00
kshitij 3baf6c14a5 fix: guard against string schedule in _clear_run_claim_best_effort
The PR's guard used `(job.get('schedule') or {}).get('kind')` which
crashes with AttributeError when schedule is a raw string (e.g.
'every 5m'), as happens in test_parallel_pool.py fixtures and any
job created via create_job(schedule='every 1h'). Use the
isinstance guard pattern already used at lines 5181 and 5325.
2026-08-17 17:31:26 +05:30
kshitij baa1cfb89d perf(cron): skip run_claim clear for recurring jobs on dispatch failure
/simplify-code finding: only one-shots carry a run_claim, yet the three
dispatch-failure paths called clear_run_claim unconditionally — each call
acquires _jobs_lock (blocking cross-process flock) and does a full
load_jobs read just to return False for any non-'once' job. The trigger
is exactly a failure storm (interpreter shutdown, EMFILE with N due
jobs): N serialized flock+file reads at the moment the process can least
afford I/O, all guaranteed no-ops for the majority job kind.

Gate at the call site on schedule.kind == 'once'; new mutation-checked
test proves recurring dispatch failures skip the claim I/O entirely.

9/9 tests green; ruff clean.
2026-08-17 17:31:26 +05:30
kshitij 70fc5a5ee2 fix: guard claim cleanup best-effort + add #86522 regression tests
Follow-ups on the #87591 salvage:

- cron/scheduler.py: wrap the three clear_run_claim call sites in a
  best-effort helper — clear_run_claim does load_jobs/save_jobs file I/O,
  and on the interpreter-shutdown path (or with a corrupt store) it could
  itself raise, defeating the skip-cleanly purpose of these early exits.
  A claim that can't be cleared simply expires at the TTL, as before.
- tests/cron/test_oneshot_dispatch_failure_run_claim.py (new): 8 tests —
  clear_run_claim unit contract (one-shot cleared / already-clear noop /
  recurring never touched / unknown id), all three dispatch-failure paths
  through a real tick() clear the claim, and a raising clear_run_claim
  does not crash the tick. Mutation-verified: reverting the fix makes the
  suite fail.
2026-08-17 17:31:26 +05:30
RelaxJonh 6d85f214ac fix(cron): clear run_claim for one-shot jobs on dispatch failure (#86522)
get_due_jobs() stamps a run_claim on one-shot jobs before returning
them as due, and mark_job_run() clears it on successful completion.
When dispatch itself fails (interpreter shutdown, executor submit
error, execution-creation error) the job never reaches mark_job_run
and the stale claim blocks re-dispatch until the TTL expires
(default 30 min).

Add clear_run_claim() to jobs.py and call it on every early-exit
path in _submit_with_guard so the job stays due and fires on the
next healthy tick — matching the existing scheduler comment's
promise.

Fixes #86522
2026-08-17 17:31:26 +05:30
kshitij 55e2dffe16 refactor(cron): two-phase ledger query + reuse shared terminal-state/timestamp helpers
/simplify-code findings on the salvage stack:
- efficiency HIGH: latest_executions() ran a SQLite connect + DDL + query
  every tick for the whole duration of ANY running job, even when every
  claim had a live future and the result was never consulted. Two-phase
  now: snapshot (job_id, future) under _running_lock, query the ledger
  only for claims whose future is missing/pending/done — the healthy
  steady state pays zero DB work per tick.
- quality: inline ("completed", "failed", "unknown") tuple duplicated
  cron/executions._TERMINAL_STATES (drift risk) — import the constant.
- reuse: hand-rolled naive-timestamp normalization in _row_belongs_to_claim
  duplicated cron.jobs._ensure_aware's legacy-naive policy — reuse it.
- quality: dropped the tautological 'if fut is None or pending or done'
  re-check (control only reaches it after the live-future continue) and
  collapsed the two copy-pasted release blocks into one with a computed
  reason.

24/24 tests green; mutation check re-verified on the final stack
(defeating the ownership guard fails exactly the 2 race-guard tests).
2026-08-17 16:56:54 +05:30
kshitij 6370134bfe fix: ledger-terminal release only for rows belonging to THIS claim
Follow-ups on the #87259 salvage:

- cron/scheduler.py: the ledger-terminal reconciliation now requires the
  terminal execution row's claimed_at to be >= the in-memory claim's
  registration time (_running_since). Without this, the latest terminal
  row for a recurring job is usually the PREVIOUS run's outcome — a fresh
  claim in the try_register_running_job -> create_execution window (or a
  finished run whose worker finally block hasn't released yet) would be
  force-released and the job double-dispatched. Unparseable/missing
  claimed_at fails closed to the age-based bound.
- cron/scheduler.py: take the _running_job_ids snapshot for the ledger
  query under _running_lock — list() over a set concurrently mutated by
  try_register/release_running_job can raise RuntimeError.
- tests: existing reconciliation tests updated to the claimed_at contract;
  two new race-guard tests (previous-run terminal row never releases a
  fresh claim; missing claimed_at fails closed). Mutation-verified:
  removing the ownership guard fails both.
2026-08-17 16:56:54 +05:30
devops 9b9cfbc1fd fix(cron): reconcile stale in-flight claim against executions ledger (t_8b5480b3)
The age-only stale-claim sweep (t_3778a491, already on main) force-releases
an in-memory _running_job_ids claim only once it is older than
max(2*interval, 30m). A leaked claim that is YOUNG (inside its allowance)
while the durable executions ledger already proves the last run ended stays
wedged: the job is returned as due every tick, _submit_with_guard short-
circuits on 'already running', and next_run_at keeps fast-forwarding with no
execution — the exact 2026-08-14 recurring-router incident (t_20e23f84),
which survived a gateway restart because the in-memory age bound alone could
not see a run the ledger had already finished.

sweep_stale_inflight now reconciles each in-flight claim against the durable
executions ledger (cron/executions.db): if the job's MOST RECENT execution
row is terminal (completed/failed/unknown), the run provably ended, so the
claim is stale by construction regardless of its in-memory age and is force-
released. This is a persisted-state recovery path: the ledger is written by
the worker that ran the job and read by ANY ticker process (including one
that started AFTER the leak), so a leaked claim is recoverable without
force-run/resume and without depending on which process holds it in memory.
A ledger-terminal release is authoritative — it does not write a synthetic
mark_job_run failure (the ledger already records the outcome).

Added TestLedgerTerminalReconciliation (4 tests): young+terminal -> released
(RED on main, GREEN here), no-ledger-row -> not released, running-row -> not
released, old+terminal -> released once without synthetic failure.
2026-08-17 16:56:54 +05:30
kshitij 0a8e703701 refactor(cron): dedup EMFILE tick-failure handling; share the fd-exhaustion text matcher
/simplify-code findings on the salvage stack:
- the classify+reclaim+counter block was pasted verbatim into both ticker
  loops (_start and _start_multiplex) along with duplicated function-local
  imports — extracted _note_tick_failure() next to _backoff_wait_seconds
  so both loops share one implementation.
- hermes_cli/cron.py's EMFILE hint reimplemented the text half of
  _is_fd_exhaustion with a case-SENSITIVE variation (drift risk) — split
  _is_fd_exhaustion_text() out and use it from both.

11 EMFILE tests + 54 provider/ticker tests green; ruff clean.
2026-08-17 16:55:25 +05:30
kshitij 80dc1836c4 fix: single-owner fd reclamation, shared backoff helper, clean CLI tick failure
Follow-ups on the #87796 salvage:

- cron/scheduler.py: drop the _reclaim_fds_best_effort call at tick()'s
  lock-failure raise site — the ticker loop's except handler already runs
  reclamation once per failed tick, so the raise-site call doubled the
  gc.collect() pause on every EMFILE failure.
- cron/scheduler_provider.py: extract the exponential-backoff math
  duplicated verbatim in start() and _start_multiplex() into a module-level
  _backoff_wait_seconds() helper.
- hermes_cli/cron.py: `hermes cron tick` now reports a propagated OSError
  cleanly (exit 1) instead of dumping a traceback — tick() raising on real
  lock-acquisition failures is new behavior from this fix.
2026-08-17 16:55:25 +05:30
webtecnica 815934ae52 fix(cron): scheduler self-heals after EMFILE instead of stalling silently (#87644)
tick() swallowed a real OSError at tick-lock acquisition as 'another
instance holds the lock', so fd exhaustion (EMFILE/ENFILE) made the
scheduler return 0 — recorded as a successful tick — while no job ever
ran again. Heartbeat and success markers stayed fresh, masking the stall.

- propagate lock-acquisition OSError to the ticker loop (records + backs off)
- detect fd exhaustion, attempt gc.collect() + raise soft nofile limit
- exponential backoff so an exhausted process stops hammering the store
- preserve genuine lock contention (EWOULDBLOCK) silent-skip behavior
- 11 regression tests
2026-08-17 16:55:25 +05:30
konsisumer 17fa4e2944 fix(cron): direct drift remediation to user-owned pins 2026-08-17 02:27:37 -07:00
Teknium 47d7661aa8 Inspired by Amp: cron self-context — context_from='self' gives recurring jobs run-to-run continuity
Amp's 'Right on Schedule' (Jul 21 2026) lets scheduled agents wake up with
their saved context and continue where they left off. Hermes cron jobs run
in isolated sessions with per-run amnesia; the existing context_from chain
mechanism only referenced OTHER jobs. This adds the special value 'self'
(and treats a job's own literal id the same way): the job's most recent
output is injected with continuity framing so recurring scouts/monitors
dedupe against what they already reported and continue where they left off.

- cron/scheduler.py: resolve 'self'/own-id in _build_job_prompt with
  continuity framing instead of upstream-job framing
- tools/cronjob_tools.py: allow 'self' through create/update validation
  (can't be validated against the store — the job doesn't exist yet at
  create time); schema description documents the value
- tests: 6 new tests incl. sabotage-verified failures without the fix
- docs: self-context section in cron.md
2026-08-16 22:09:28 -07:00
Teknium 07a5179158 Inspired by Poke: nudge review of repeatedly-failing recurring cron jobs
Poke (poke.com) 'encourages users to review recurring automations that
haven't been acted upon'. Hermes' equivalent pain point is a recurring
cron job that fails run after run: each failure delivers the same one-line
error with no signal that the automation itself needs attention.

- cron/jobs.py: persist a failure_streak counter in mark_job_run —
  incremented on agent failure, reset on success; delivery failures don't
  count. Back-compat: missing field reads as 0.
- cron/scheduler.py: _failure_streak_nudge() appends a review nudge to the
  delivered failure summary once a recurring job's streak reaches
  cron.failure_nudge_threshold (default 3, 0 disables). One-shots never
  nudge.
- hermes_cli/cron.py: 'hermes cron list' shows '(N failures in a row)' on
  failing jobs with streak >= 2.
- docs: new 'Repeated-failure review nudge' section in cron.md.

Tests: 17 passed (TestMarkJobRun + TestFailureStreakNudge); E2E verified
with real cron store in temp HERMES_HOME.
2026-08-16 22:08:59 -07:00
SHT 795e035f60 fix(cron): coerce script_path to str in the NUL guard so it can never crash (#86829, review #86832)
"\x00" in script_path raises TypeError when a caller passes a non-str
(e.g. a pathlib.Path, which is not iterable) — the guard itself would
crash the scheduler. All current call sites pass plain str, but the
guard must be crash-proof: str() first, then check. Adds a regression
test running a real script through _run_job_script with a Path argument,
which fails with TypeError on the pre-fix guard.
2026-08-16 06:29:03 -07:00
SHT 40586082e5 fix(cron): reject NUL-bearing script paths before any Path call (#76762 class)
_run_job_script wrapped only expanduser() in its ingestion try/except.
On Linux an unexpandable NUL-bearing value raises inside that call, so it
landed on the clean fail-with-report path; on Windows expanduser() never
expands '~user' (and so never raises), and the NUL surfaces later as an
uncaught ValueError from resolve()/exists() — crashing the scheduler.

Align with cron.lifecycle_guard._expand_candidate_path, which already
documents this as the whole-class fix (the per-syscall catching produced
#76762, #77703, #77780, #78256): reject '\\x00' eagerly at the ingestion
boundary so both platforms fail identically.
2026-08-16 06:29:03 -07:00
Merge_Conflict - Pasi 309cf2c5e2 fix: honor JSON-array string forms for skills.disabled and agent.disabled_toolsets
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.

Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.

Fixes #86661
2026-08-16 06:24:56 -07:00
SHT 58d6bf2d61 fix(cron): log and document the .pth bootstrap fallback (#86816 review)
- WARN when the venv site-packages layout is unresolvable and the script
  falls back to plain PYTHONPATH execution, so 'editable installs
  invisible' failures are diagnosable.
- Docstring: note that runpy does not set __package__/__spec__ the way a
  direct python script.py invocation does.
2026-08-16 06:24:06 -07:00
SHT 029a0d8c76 fix(cron): process .pth files for Windows uv-venv script jobs (#86567)
_windows_cron_python_invocation bypasses the uv venv launcher (to avoid
flashing a console window) and re-attaches the venv via PYTHONPATH — but
PYTHONPATH entries are plain sys.path additions and never get .pth
processing, so editable installs (pip install -e) were invisible to cron
script jobs (ModuleNotFoundError).

Bootstrap the script with site.addsitedir() on the venv site-packages,
then exec it as __main__ via runpy.run_path, preserving the script
directory on sys.path (python script.py semantics). Falls back to a
plain invocation when the venv layout is unresolvable.
2026-08-16 06:24:06 -07:00
Teknium 0fc2a10d82 fix(cron): stop one-shot CLI cron run from orphaning the job; reap dead-owner claims on tick
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.

Two-part fix:

1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
   channel stateless (scoped ContextVar set/reset around the call) before
   invoking the cron API, so `async_delivery_supported()` gates off
   `_try_dispatch_background_run` and the run executes synchronously to
   completion in the CLI process — the same behavior `hermes -z` already
   gets via declare_stateless_channel().

2. cron/scheduler.py: tick() now periodically invokes
   recover_interrupted_executions() (previously only run at scheduler
   startup), so execution rows whose exact owner process is provably dead
   (pid + process start time check in _owner_is_live) are reaped to
   'unknown' by the long-lived gateway ticker without a restart.
   Throttled to once per 300s so idle 60s ticks don't pay a ledger
   connection every cycle.

Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.

Fixes #86721
2026-08-15 02:44:14 -07:00
yuric 569b7a34b1 fix(cron): bound post-run cleanup 2026-08-14 21:55:14 -07:00
Teknium a5bb1bcde3 fix(cron): make transient-DNS classification platform-safe
The errno literal set {8, 7, 11, 51, 60, 61, 65} mixed macOS getaddrinfo
constants with errno values: on Linux socket.EAI_NONAME is -2 and
EAI_AGAIN is -3, so genuine DNS failures were missed while unrelated
OSErrors carrying errno 8/11 (ENOEXEC/EAGAIN semantics differ) could be
misclassified as transient. Classify socket.gaierror against the EAI_*
constants and plain OSError against named errno constants instead.

Addresses the platform-portability review on #83977.
2026-08-14 21:55:14 -07:00