The scheduler booked every finished run as end_reason=cron_complete based
on the run lifecycle alone. A job whose agent turn died after a tool
call, mid-API-wait, or without any assistant text still surfaced as a
healthy run — one audited day held 10 such silently-failed sessions
whose run history showed green (#93820).
Before end_session, the session's LAST message row is now classified
through the existing cost-bounded session_lifecycle_statuses helper:
only a real assistant reply (a plain answer or the [SILENT] sentinel —
both assistant-text rows) keeps cron_complete; the positively
recognized pathological statuses (interrupted / error / empty) book the
run as cron_incomplete_no_output with a warning. Unknown values and
probe failures keep the historical reason — classification is
best-effort metadata and must not mislabel a healthy run. The new
end_reason is a free-form forensics string like cron_complete (in no
recovery/reset whitelist), so session recovery semantics are unchanged.
Fixes#93820
Carried from #94034 (closed as duplicate of #94033): explicit regression
that a jobs.json-edited stale next_run_at with NO manual_run_at marker
still re-anchors without firing — the direct #93049 protection case.
Carried from #94770 (closed as duplicate of #94775): black-box tests that
build a real temp HERMES_HOME config.yaml and assert the exact rendered
TimeoutStopSec strings in the generated unit, including the
HERMES_CRON_DRAIN_TIMEOUT env-override case — complementing #94775's
helper-level tests.
cli-config.yaml.example said macOS is "always held at FULL regardless",
which reads as "your setting is ignored on this platform" and would talk
an operator out of choosing EXTRA. _apply_synchronous_pragma only refuses
values BELOW FULL on Darwin; EXTRA is applied normally.
The existing doc guard only asserted the key's presence, so it could not
have caught this. Pin the distinction instead.
Reported by @Enough1122 in review.
`apply_database_pragmas()` reads five sizing pragmas from `database:` and no
durability one, and `_enforce_macos_synchronous_full()` returns early when
`sys.platform != "darwin"`. Between them, nothing in the process ever executes
`PRAGMA synchronous` against state.db on Linux or Windows.
The effective level there is therefore `SQLITE_DEFAULT_WAL_SYNCHRONOUS`, a
compile-time constant of whichever SQLite the interpreter links. Debian and
Ubuntu builds commonly ship it as NORMAL; the bundled build and a plain
source build use FULL. So the durability of state.db is decided by which
python3 the installer found, is invisible from config, and cannot be pinned.
#90837 is three weeks of corruption forensics conducted on Ubuntu under the
stated premise `synchronous=FULL`, with every other cause eliminated live.
That premise is not something the reporter could have verified from config,
because there was no config key to set and no log line to read back.
- `resolve_synchronous_level()` maps the spellings operators actually write
(OFF/NORMAL/FULL/EXTRA, any case, or 0-3) to the PRAGMA integer, and
returns None for anything else. Kept out of the sizing loop on purpose: an
unrecognised `cache_size` harmlessly falls back to a default, an
unrecognised durability level must not.
- `_apply_synchronous_pragma()` applies it, and on Darwin refuses to lower
below FULL. `_enforce_macos_synchronous_full()` runs during
`apply_wal_with_fallback()`, which is earlier than `apply_database_pragmas()`,
so without an explicit floor a configured NORMAL would silently undo #64355
by the accident of running last. Raising to EXTRA on macOS is allowed.
- Unset changes nothing, so no existing install moves.
Tests: 34, covering the parser, application, the unset path, the typo path,
a guardrail that #77630's five keys still apply, and the Darwin floor in both
directions. Removing the wiring fails 6; removing the floor alone fails 1.
Related to #90837
apply_wal_with_fallback treats an on-disk WAL database as authoritative and
says so twice: it never live-downgrades one. The mirror case had no
protection at all. When the on-disk mode is DELETE and the configured mode
is wal, the function flips the database and logs nothing.
journal_mode is a property of the FILE, so that rewrites the header and
persists after the process exits. Setting the mode directly on the file is
something operators do; it was the documented mitigation for the SQLite
3.50.4 WAL-reset bug. A config key that makes the choice durable already
exists (database.journal_mode, #68545), but nothing named it at the moment
the PRAGMA was being undone.
#89293 reports the cost: after upgrading past the vulnerable SQLite,
is_sqlite_wal_reset_vulnerable() stopped short-circuiting into
_apply_delete_for_wal_reset_bug, the flip path went live, and 4 of 5
databases silently returned to WAL with no log line anywhere.
Add a deduped WARNING at both points where the switch succeeds, decided
before the pragma runs since both inputs are only readable while the file is
still in its original state. Log-only: the flip still happens, the return
value is unchanged, and the never-live-downgrade rule is untouched.
WARNING rather than ERROR is deliberate. The reverse direction is ERROR
because dropping to DELETE costs concurrency; this direction is normally the
desirable one (managed_uv treats a database stuck on DELETE as a bug worth
repairing on update). The problem was never the change, it was that the
change was invisible.
The page_count guard is the load-bearing half. A brand-new database also
reports journal_mode=delete and is also about to be switched to WAL, and
every opener applies WAL before creating schema, so without it the warning
would fire on the first run of every install.
Refs #89293
When database.journal_mode=delete is configured but the on-disk DB is already
WAL, apply_wal_with_fallback honors the never-live-downgrade rule and keeps WAL.
That is correct (a live downgrade under open connections causes mixed-mode
corruption), but the operator's configured mode silently has no effect, and on
a WAL-incompatible filesystem (virtiofs/NFS/SMB) the DB then corrupts on the
next crash/sleep exactly what they configured to prevent.
Two code paths return WAL in this situation; both now emit a once-per-process-
per-db_label ERROR telling the operator the config did not apply and they must
convert the DB header offline (stop connections, PRAGMA journal_mode=DELETE):
1. The WAL-reset-vulnerable path (_apply_delete_for_wal_reset_bug): previously
warned only about the vulnerability with an "upgrade SQLite" remedy, which
does not help when the real cause is the filesystem. Emitted after that
warning so the actionable message is last.
2. The read-only probe path (non-vulnerable runtime): previously returned WAL
with no signal at all.
The never-live-downgrade behavior is unchanged (existing test now also asserts
the warning). New tests cover both paths, the per-db_label dedup, and the
require_wal=True edge case.
Real-world impact: a Hermes deployment with state.db on a Podman virtiofs
bind-mount (or any NFS/SMB home) that upgrades across a version where WAL was
the default, then sets journal_mode=delete, sees no corruption protection
until the DB header is converted. This makes the gap visible. See #68545.
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
Review found _create_entry_from_recovered_row builds a minimal entry:
replacing the live object would silently drop model_override, token/cost
counters, resume_pending/queued-work markers, and metadata. Keep the
original entry (routing is unchanged, so no sessions.json rewrite either)
and log at INFO — this is a success path, not a corrective action.
Regression test now asserts state preservation, reopen_session call, and
no save.
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.
`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:
1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
`_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
used by `systemd_restart()` and the updater — had no launchd call site.
2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
systemd branch uses `_get_restart_exit_wait_budget()`
(drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
documents that callers falling back to a hard kill must cover both phases
or they reintroduce #77184.
The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.
Observed on macOS 27.0 / Hermes 0.20.4:
→ Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
⚠ Gateway PID 49787 still running after 0.0s — restart may fail
⚠ Gateway drain timed out after 0s — forcing launchd restart
Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.
The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.
Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.
Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.
Refs: #79096
Fixes two related native-streaming bubble defects surfaced in production:
1. Duplicate bubble on long turns — when keep-alive already refreshed the
6-min reply window, the Layer-2 clock fallback still declined the finalize
frame and forced a proactive send(), duplicating the message. Skip the
clock fallback while keep-alive is active; intermediate-frame failures are
now fully fire-and-forget (only a failed FINAL frame falls back to send()).
2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
notice) called _reset_segment_state(), clearing the cumulative
_accumulated so the next delta + finalize frame carried only the few
chars accumulated after the reset. Native streaming now skips that
reset (commentary still posts as its own message via send()).
b) The adapter-side _BlockChunker.update() 'only grow' guard silently
dropped any cumulative snapshot shorter than its high-water mark, so
after a baseline reset the leading characters were stranded before
_emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
layer entirely; intermediate frames are pure identity-dedup, matching
the fire-and-forget model.
Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.
Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.
Includes the machinery intrinsic to native streaming on WeCom:
- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
frames update it, a finalize frame closes it; native-streaming adapters are
let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
long-connection mode has no documented edit-rate limit); an adapter-level
frame cap is retained. (Early builds gated frames behind a char throttle;
removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
(control vs normal) plus a per-chat token bucket to stay under WeCom's rate
limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
delivered the moment it is emitted; failures logged, not re-sent; delivery
marked once per turn), per-req_id reply queue with ack tracking, and the
timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
the prompt is the last thing on screen and never traps a lingering bubble;
eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
messages; image+text double-callback merged into one turn.
Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
When send_message is invoked from the agent's worker thread (a different
event loop than the gateway's), awaiting the WeCom adapter directly can hang
because the adapter enqueues onto the gateway loop. Dispatch via
run_coroutine_threadsafe onto the gateway loop when the caller loop differs,
with caller-cancellation shielded so an already-enqueued send is not cancelled
mid-flight (which would otherwise cause a false-failure retry -> duplicate).
Recognizes WeCom native chat IDs as explicit send targets and whitelists WeCom
for media delivery. Part of the async queue design this branch introduces.
Add WeCom to the per-platform streaming defaults (DEFAULT_CONFIG display
plumbing) so native streaming is enabled by default for the WeCom adapter,
alongside the existing per-platform flags. Non-secret config lives in
config.yaml (no HERMES_* env vars).
Per the 'when in doubt, optional' rule — site publishing is an
on-request capability, not a weekly daily-driver for most users.
Joins cloudflare-temporary-deploy/page-agent under
optional-skills/web-development (existing category, existing
DESCRIPTION.md kept; the new bundled category dir is dropped).
Install via: hermes skills install official/web-development/publish-site
Precision pass on the comments added by the previous commit: prompt.submit
reaches replace_messages(archive_dropped=True) with an empty prefix on a
confirmed ordinal-0 rewind, which is the production shape that lands a
populated session on message_count = 0. archive_and_compact normally
publishes at least a summary row, so it is pinned as defense in depth
rather than claimed as an equally reachable trigger.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:
* `replace_messages(..., archive_dropped=True)` — the rewind / edit /
regenerate mode added in #82756 so a taken-back turn stays recoverable.
* `archive_and_compact` — in-place compaction, which archives the
pre-compaction transcript under the same session id (#38763).
A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.
Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.
Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.
Fixes#95868
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
Replaces the stale MCP-first AgentMail skill with CLI-first guidance:
self-signup + OTP verification, inbox/message/thread/label/attachment
flows, webhook and WebSocket delivery references, and MCP as an
alternative path. Declares AGENTMAIL_API_KEY (optional) so a stored key
reaches the sandboxed terminal while self-signup stays viable without
one.
Salvaged from PR #60811 — kept in optional-skills/ per the March 2026
decision that third-party-API-key skills are not bundled.
run_job opened state.db (SessionDB) at the top of the function, before the
wake-gate (wakeAgent: false), prompt-injection block, and drift-skip early
returns. Every gated run therefore opened a full SessionDB — read pool,
token-writer machinery, .db/-wal/-shm handles — and returned without
reaching the finally that closes it, relying on GC/__del__ to release the
descriptors. On a gateway whose monitor-gated jobs tick every few minutes,
that is constant wasted open/migrate work and GC-dependent fd lifetime.
Move the init inside the main try, immediately before AIAgent construction,
after every early-return path. The timeout resolution now reuses the _cfg
already loaded for model routing instead of a second load_config() call.
Behavior on the normal (non-gated) path is unchanged: same env/config/default
timeout resolution, same abandoned-worker done-callback close (#72782), and
the existing finally still closes the store after the agent turn.
Salvaged from PR #96290 (cron slice) with a mutation-checked regression test
(fails on main: gated run opens SessionDB; passes with the reorder).
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.
Salvaged from PR #96290 (kanban slice) with regression test.
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.
Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
httpx timeout exceptions (ReadTimeout, WriteTimeout) stringify to "",
which defeats _is_timeout_error's first-line guard (if not error: return
False). The base-layer plain-text fallback then re-sends an already-
delivered message — the user receives it twice.
Replace error=str(exc) with error=str(exc) or type(exc).__name__ at every
httpx-based adapter boundary so the existing matcher ("readtimeout",
"writetimeout") still fires. ConnectTimeout intentionally stays
unmatched: if the connection never opened the message was not delivered,
so retry/fallback remains correct.
Applies to BlueBubbles (send + _create_chat_for_handle), WhatsApp Cloud
(text + interactive + media), QQ Bot (send chunk + keyboard + media), and
Yuanbao media handler — the same latent bug exists in every adapter that
stores error=str(exc) from an httpx call.
- _write_machine_sentinel_line: wrap the print() fallback so a closed
redirected stream (ValueError, not OSError) can't propagate out of the
ready path and kill a healthy serve; document that pythonw port
discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
redirect active all server logging lands on stderr, and an unread
stderr pipe can fill and block the child before the sentinel, flaking
the test at the 120s timeout
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).
Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.
Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
Builds on e11187208f (salvage of #72206 by @luyifan, authorship
preserved) which added the post-timeout session reset. This commit
completes the Phase 3b contract:
- suspect flag: a command timeout marks the session key suspect;
the NEXT _get_session_info for that key health-checks and recycles
the session instead of handing back the poisoned handle (#72205)
- flag cleared on fresh-session store so a healthy new session is
never spuriously recycled; cross-key leakage fixed
- wedged-vs-alive rule (#68139): after a timeout, if the daemon's
socket still answers, recycle the session only; if it is
unresponsive (or no socket exists to probe — conservative), tree-
kill via agent.deadline.kill_process_tree and evict
- negative probe: successful commands never mark or recycle
Tests: tests/tools/test_browser_suspect_recycle.py (20 tests:
mark-once, recycle-then-succeed, success-never-recycles, tree-kill
invoked on the wedged path with pid assertion, flag lifecycle).
Co-authored-by: luyifan <al3060388206@gmail.com>
Sibling site of the salvaged #95947 fix (same file): cron_status
declared 'Gateway is not running — cron jobs will NOT fire' from a bare
find_gateway_pids() miss even while the runtime lock proved the gateway
alive. Now the not-running verdict requires both the scan AND the lock
to read dead; when only the lock answers, the pid line falls back to
the recorded gateway pid (or is omitted).
Two regression tests pin the false-alarm suppression and the genuine
not-running warning.
Follow-ups to the salvaged #95947 cron commit:
- Wrap the lock probe in its own try/except: a crashing probe is
'unknown', not 'dead' — the pid scan still decides instead of the
whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
held + empty pid scan -> alive (the reported false alarm); lock
inactive -> pid-scan fallback both ways; crashing lock probe still
falls back.
- patch_liveness now pins the lock probe inactive by default so the
pre-existing pid-scan tests stay deterministic on machines where a
real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
Independent review finding on the merged #95980:
_apply_live_compression_config only acted on PRESENT keys, so removing
tail_mode / model.context_length / target_ratio / model_thresholds /
proactive_prune_* / protect_last_n / min_tail_user_messages / threshold /
idle_compact_after_seconds from config.yaml left stale values active in
live sessions forever (probe-verified: all six stale after applying
empty mappings).
Absence now restores the normalized default — or the model-derived
value — through the SAME derivation the construction path uses:
- ContextCompressor ctor defaults read off its real __init__ signature
(no hardcoded copies to drift)
- compression.threshold removal re-derives via agent_init's
_resolve_compression_threshold (Codex gpt-5.4/5.5 + spark autoraise
included)
- model.context_length removal drops the config override and forces
re-inference through the deferred get_model_context_length resolution,
which also re-applies the small-context threshold floor
- model_thresholds removal clears stale per-model overrides from the
live threshold; tail_mode falls back to the ctor's 'lean' (the old
present-key path normalized invalid values to 'legacy', diverging
from the compressor's own fallback)
Also fixes proactive_prune_min_reclaim_tokens's present-but-null default
(was 0; the real default is 4096).
Refs #94724
A hermes serve killed mid-update lost every un-flushed in-memory session
(#94724 item 2, reported by @ruangraung): the next RPC failed with
'session-scoped RPC rejected: not in memory (detached/reaped runtime)'
and no store held the transcript. #95576 made serves survive future
updates; this closes the kill path itself:
- install chaining SIGTERM/SIGINT handlers (hermes serve / dashboard
startup, before uvicorn's capture_signals) that first persist
in-memory session transcripts to state.db — bounded by
HERMES_TUI_EXIT_FLUSH_BUDGET_S (default 5s, daemon worker + join) so
a hung SQLite write can never block exit
- _shutdown_sessions (atexit) runs the same bounded flush FIRST, before
the slow per-session teardown a supervisor may SIGKILL mid-way
- the idle-reaper scan piggybacks a periodic incremental flush
(marker-deduped agent._persist_session, running sessions skipped) so
even a SIGKILL loses at most one flush interval — no new timer
subsystem
Refs #94724
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
The smart-approval guardian (`_smart_approve`) gates every flagged
terminal command with a synchronous auxiliary LLM call, but it never
passes `timeout=` and logs nothing on the normal path. In production a
stalled provider response silently froze the agent turn for 62 minutes
with zero log output; the gateway kill-switch eventually fired, and only
an unrelated error surfaced afterwards (#82846; watchdog-style fix in
#72500). The call was invisible by design — nothing logs at the hang
point.
Changes in tools/approval.py:
- Resolve the same configured timeout the client would use internally
(`auxiliary.approval.timeout` via `_get_task_timeout("approval")`) and
pass it explicitly to `call_llm`, so the deadline cannot be lost if the
internal default resolution changes or is misconfigured.
- Log the assessment call and its duration (DEBUG), and promote the
failure branch from DEBUG to WARNING with elapsed time + exception
class, so a wedged guardian call is visible in the logs instead of
silent.
- Failure still returns "escalate" (fail open to the human/pattern
gate) — behavior unchanged, observability only.
Complements #72500 (watchdog hard ceiling) rather than duplicating it:
explicit timeout is the root-cause hardening, logging closes the
silence gap; the watchdog remains the safety net if the SDK-level
timeout itself is defeated.
Tests: explicit timeout forwarded to call_llm (revert-fails), failure
logs WARNING + escalates. 49 approval-adjacent tests pass; one unrelated
test_approval.py failure is pre-existing (fails on clean main too).
The script-timeout path used a site-local process-group kill, which
cannot reach a grandchild that created its OWN session (start_new_session
background jobs, watchdogs). Such descendants kept running after the job
reported failure (#71148, #59549). Migrate the timeout handler to the
unified deadline layer's kill_process_tree (#85147, d6a5cb9725): psutil
snapshots the descendant set before signalling, so own-session
grandchildren are reached too. Fallback to the site-local group kill if
the import ever fails, so the path cannot re-wedge.
The explicit script-timeout message stays the classification anchor
(#85536's contract), keeping cron timeouts distinct from provider
timeouts.
Salvage additions on review (#85125 Phase 4a):
- migrate the sibling kill site too — the cancel_event/"ownership was
lost" path orphaned setsid grandchildren the same way (whole-bug-class
rule); pinned by test_cancel_path_also_tree_kills
- proc.poll() early-return in _terminate_cron_script_tree so a script
that exits right at the deadline doesn't log a spurious "no signal"
warning (mirrors _terminate_cron_script_process); pinned by
test_already_exited_proc_is_left_alone
- acceptance test's script timeout 1s -> 2s: interpreter startup under
CI load could eat the whole 1s window before the spawner wrote its
pid file
- note: kill_process_tree hard-kills (SIGKILL) immediately, whereas the
old path gave a 1s SIGTERM grace window; intended for a deadline-
expiry hard stop (both docstrings say "hard stop")
Based on #86791 by @ayushnangia; cherry-picked to preserve authorship.
Co-authored-by: dante32683 <dante32683@users.noreply.github.com>
Co-authored-by: supotato-ipj <supotato-ipj@users.noreply.github.com>