Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:
- suspect flag: a command timeout marks the session key suspect;
the NEXT _get_session_info for that key health-checks and recycles
the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
socket still answers, recycle the session only; if it is
unresponsive (or no socket exists to probe — conservative), tree-
kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle
Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).
Co-authored-by: luyifan <al3060388206@gmail.com>
Sibling site of the salvaged #95947 fix (same file): cron_status
declared 'Gateway is not running — cron jobs will NOT fire' from a bare
find_gateway_pids() miss even while the runtime lock proved the gateway
alive. Now the not-running verdict requires both the scan AND the lock
to read dead; when only the lock answers, the pid line falls back to
the recorded gateway pid (or is omitted).
Two regression tests pin the false-alarm suppression and the genuine
not-running warning.
Follow-ups to the salvaged #95947 cron commit:
- Wrap the lock probe in its own try/except: a crashing probe is
'unknown', not 'dead' — the pid scan still decides instead of the
whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
held + empty pid scan -> alive (the reported false alarm); lock
inactive -> pid-scan fallback both ways; crashing lock probe still
falls back.
- patch_liveness now pins the lock probe inactive by default so the
pre-existing pid-scan tests stay deterministic on machines where a
real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
_builtin_gateway_liveness() decides whether the builtin cron ticker can fire by
PID-scanning via find_gateway_pids(). That scan can transiently return empty even
while the gateway is up (e.g. just after a restart), so the in-gateway cronjob tool
emits a false 'Gateway is not running — jobs won't fire' while jobs are firing on time.
Prefer the gateway runtime lock: it is held for exactly the gateway's lifetime (a
reliable liveness signal), and inside the gateway process it short-circuits to True,
so the in-gateway check can never false-alarm. Fall back to the PID scan only when the
lock reads inactive (the external-CLI path).
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
Follow-up to the #94755 salvage: every runGroupChatRounds exit recorded
'settled', so a room that died at the round/message/continuation cap looked
identical to genuine consensus. Track the exit kind and record 'capped'
(with label + glyph) when a cap ended the drive, and pin the exit-path
wiring with source-contract tests.
- GROUP_CHAT_MAX_CONTINUATIONS=2 caps continuation rounds independently of
the message cap, so pathological @mention chains can't consume the room's
whole budget on handoffs.
- unaddressedGroupMentions now orders by log INDEX instead of entry id:
ids are UUIDs (groupChatEntryId), not monotonic — string comparison could
both re-drive answered members and miss stranded ones.
- New unaddressed-mentions.test.mjs exercises the REAL function via the
vm-slice pattern (replacing concept-only helpers) and pins the ordering
fix with a UUID-vs-log-order case.
In Bot Mode group chats, a member reply that @mentions a teammate never
drove the cited bot when the current round went quiet: the
'spokeThisRound === 0' early exit treated a zero-reply round as 'everyone
passed' and settled the room, even though the reply's @mention was
pending. The same silent settle happened whenever GROUP_CHAT_MAX_ROUNDS
or GROUP_CHAT_MAX_MESSAGES landed between the mention and the next
round (#94478).
Fix:
- unaddressedGroupMentions() detects member-to-member citations in the
thread whose cited member has not posted anything after the citing
entry (self-mentions excluded; user sends re-drive everyone anyway).
- The quiet-round exit now checks for such pending handoffs and runs one
bounded continuation round driving exactly those cited members — same
holds/stranded/epoch/cap machinery as ordinary rounds, so nothing new
is trusted.
- If the continuation also produces nothing (pass, failure, or cap), the
room settles as before; behavior only changes where a bot was actually
called and never answered.
Regression tests in plugins tests family (2 new); full hermes-bots
plugin suite stays green (558/558).
Fixes#94478
1. Quoting: the spawn payload wrapped expandRemotePath() output -- already
a shell-quoted fragment like "$HOME"'/...' -- in shq() again, so the
reservation/lock/owner_file variables hold the quote characters
literally and every mkdir "$reservation" fails forever (~5 min per
attempt spinning in the reservation loop while holding the box-global
update mutex; queued spawns starve behind it). The same double quoting
sits in the stale-reaper identity guards, making every reap REFUSE.
The lockfile-reuse path masks the bug for existing backends, so it
only bites on fresh spawns.
2. Bashism: lockfile publication used ${var//__PID__/$child} -- bash-only
substitution in a payload run under plain sh (dash on Ubuntu), which
aborts the script AFTER the serve was spawned. The client then saw an
unknown failure, ran its error cleanup (deleting the token file), and
the just-booted serve died on the missing token -- orphaning one serve
per attempt. Replaced with a POSIX sed substitution.
Adds two regression tests: payload variables must keep $HOME expandable
(no re-quoting), and the pid substitution must be POSIX sh. Both fail
against the previous code; all 89 remote-lifecycle tests pass with the
fix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The OpenRouter model picker builds its list from a curated set of model
IDs, then filters against OpenRouter's live catalog. minimax/minimax-m3:free
exists on OpenRouter (free tier, 1M context, tool-calling) but was missing
from both the in-repo fallback list and the remote catalog manifest.
Add it to OPENROUTER_MODELS and website/static/api/model-catalog.json so
the free variant surfaces in the picker alongside the paid one.
Independent review finding on the merged #95980:
_apply_live_compression_config only acted on PRESENT keys, so removing
tail_mode / model.context_length / target_ratio / model_thresholds /
proactive_prune_* / protect_last_n / min_tail_user_messages / threshold /
idle_compact_after_seconds from config.yaml left stale values active in
live sessions forever (probe-verified: all six stale after applying
empty mappings).
Absence now restores the normalized default — or the model-derived
value — through the SAME derivation the construction path uses:
- ContextCompressor ctor defaults read off its real __init__ signature
(no hardcoded copies to drift)
- compression.threshold removal re-derives via agent_init's
_resolve_compression_threshold (Codex gpt-5.4/5.5 + spark autoraise
included)
- model.context_length removal drops the config override and forces
re-inference through the deferred get_model_context_length resolution,
which also re-applies the small-context threshold floor
- model_thresholds removal clears stale per-model overrides from the
live threshold; tail_mode falls back to the ctor's 'lean' (the old
present-key path normalized invalid values to 'legacy', diverging
from the compressor's own fallback)
Also fixes proactive_prune_min_reclaim_tokens's present-but-null default
(was 0; the real default is 4096).
Refs #94724
A hermes serve killed mid-update lost every un-flushed in-memory session
(#94724 item 2, reported by @ruangraung): the next RPC failed with
'session-scoped RPC rejected: not in memory (detached/reaped runtime)'
and no store held the transcript. #95576 made serves survive future
updates; this closes the kill path itself:
- install chaining SIGTERM/SIGINT handlers (hermes serve / dashboard
startup, before uvicorn's capture_signals) that first persist
in-memory session transcripts to state.db — bounded by
HERMES_TUI_EXIT_FLUSH_BUDGET_S (default 5s, daemon worker + join) so
a hung SQLite write can never block exit
- _shutdown_sessions (atexit) runs the same bounded flush FIRST, before
the slow per-session teardown a supervisor may SIGKILL mid-way
- the idle-reaper scan piggybacks a periodic incremental flush
(marker-deduped agent._persist_session, running sessions skipped) so
even a SIGKILL loses at most one flush interval — no new timer
subsystem
Refs #94724
The fail-closed owner ladder (#95407) is correct for new sessions, but
legacy unowned rows on registry-topology installs dead-ended in
SessionOwnerResolutionError (reporter's Error B) with their transcripts
fully intact in state.db.
- resolveLegacyOwnerBackfillScope: pick the single-match store for the
server-side owner backfill at enumeration time (serving registered
connection / primary pool); fail closed on multi-candidate topologies.
- maybeBackfillLegacySessionOwners: one-shot per scope per renderer,
fire-and-forget from the #95407 stamp path, logs the stamped count.
- Read-only stored-transcript resume: when session.resume fails closed,
fetch the transcript over id-only REST (ambient first, then registered
backends, read-only probes only) and open the session as a read-only
transcript instead of dead-ending; sends are refused with a notice and
a later successful live resume clears the latch. Wired into the main
pane resume recovery and the session-tile delegate (which now runs the
same fail-closed owner gate as the RPC dispatcher).
Refs #94724
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.
Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
(`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
pass it explicitly to `call_llm`, so the deadline cannot be lost if the
internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
failure branch from DEBUG to WARNING with elapsed time + exception
class, so a wedged guardian call is visible in the logs instead of
silent.
- Failure still returns "escalate" (fail open to the human/pattern
gate) — behavior unchanged, observability only.
Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.
Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>
Apply-ready delta distilled by @andrexibiza: deterministic watcher-consumer
tests (watcher times out while a child is alive, resolves when all are dead),
psutil-unavailable fail-open pin, and probe-failure fail-open handling in
_stdio_children_dead (unknown is never proof that every child exited).
Local: 8 passed on tests/tools/test_mcp_stdio_children_dead.py
_stdio_children_dead returned True ('all children dead') on the first LIVE
pid — the intended False was dead code right below it. Every spawn path
that captures child PIDs (observed in hermes -z oneshots) then failed the
#81995 pre-call fast-fail with 'TimeoutError: MCP stdio subprocess ... has
exited' on every tools/call while the subprocess was demonstrably alive.
Long-lived gateway/dashboard sessions were unaffected only when
_stdio_child_pids was empty (the not-pids short-circuit).
Return False on the first live pid and drop the unreachable line.
TestGetHermesHome.test_default_path asserted ~/.hermes unconditionally,
but the native Windows default is %LOCALAPPDATA%\hermes (see
hermes_constants._get_platform_default_hermes_home). Branch the
assertion by platform so the test passes everywhere.
Salvaged from PR #96003 by @Aoshi-Dev (the parse-guard half of that PR
was superseded by #96169); authorship preserved.
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):
- config_command: catch RuntimeError for set/unset and print a clean
one-line error + exit(1) instead of a raw traceback on the primary
'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
ConsoleCommandError so 'hermes console' and the dashboard console
report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
branch — the old fallthrough claimed 'falling back to default config'
even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.
Fail closed when config.yaml is unparseable or non-mapping before set/unset writes, reuse the readable-config guard to return the parsed mapping, and cover refuse/empty-mapping paths with regression tests.
- Reuse the existing _commit_status variable for the terminal-edge gate
instead of the parallel _compaction_succeeded boolean (derived state).
- Give the commit_fence_cancelled abort the same force_terminal=True
terminal edge as the lock-contended abort, and reword the closure
comment that overstated the lock contender as 'the one exception'.
- Inline the codex app-server path's lifecycle closure: after gating on
success it reduced to a single success-site emit, so the scaffolding
(done-flag + closure + two no-op failure-path calls) was dead.
Eighth review round (the first against the atomic-renewal fix) verdict:
the production code holds - CAS exclusivity across real processes,
lease-extension schedules, clock skew both directions, renew-per-attempt
under 5xx backoff, defer accounting, and the author's mutants all
verified - but one shipped regression test could not fail against the
property it is named for.
test_renewal_extends_the_lease_across_the_post asserted
next_attempt_at >= lease_before under a frozen clock. A renewal that
matches the row but never extends the lease (M4: SET next_attempt_at =
next_attempt_at) satisfies >= trivially, and that mutant double-POSTs:
the un-extended lease expires mid-POST and a second process reclaims.
The reviewer demonstrated M4 surviving the whole suite while producing
a real duplicate send in a two-process schedule.
The test now renews 100s into the lease from an advanced clock and
requires the deadline to move strictly forward to exactly
renewal-clock + 300s. Verified: M4 now fails this test (61 others
unaffected); clean HEAD passes all 62.
No production code change. 277 tests; ruff + footguns clean.
The send_media lane (08b95c3) had no committed regression test: cover
explicit-bool stamping on media frames, the descriptor-platform
fallback when _platform_by_chat is empty (post-restart proactive
sends), and the omitted-key absence case.
When either unfurl key is set, media captions post as a separate
message before the file (the upload API cannot carry unfurl controls)
and native draft streaming falls back to edit-based delivery. Surface
both side effects in the config reference table.
The send and send_media lanes resolved the platform only from
_platform_by_chat, which is empty until an inbound frame arrives (e.g.
after a gateway restart). A proactive send to a Slack chat then missed
the unfurl stamp. Mirror the streaming gate and delivery resolver:
fall back to the negotiated descriptor's platform.
Relay-plane parity: hermes config set / Railway persist YAML booleans
as strings, and _slack_unfurl_kwargs silently dropped them — so
'unfurl_links: "false"' was a no-op on native while working on relay.
Coerce recognized string booleans exactly as _slack_unfurl_hints does;
unrecognized values still drop so junk config keeps Slack's default
instead of accidentally suppressing previews.
Replaces test_send_ignores_non_boolean_unfurl_options (which froze the
dropped-string behavior) with coercion + junk-drop tests.
Live staging (Coatue Slack):
- hermes config set / Railway knobs persist "true" as a string; bots that
omit unfurl_links do NOT inherit the human default, so dropping the
string looked like suppression.
- chat.startStream cannot carry unfurl_*. Native SlackAdapter already
falls back to chat.postMessage; the relay now matches.
Relay-fronted Slack reads platforms.relay.extra.slack.unfurl_links/unfurl_media
and stamps explicit booleans onto the frame metadata; the connector forwards
them to chat.postMessage with no config of its own (mirrors reply_in_thread).
Covers send, send_for_platform (cron/scheduled), and send_media lanes.