A session row persists the provider identity a chat actually used. When that
provider is later renamed or removed (e.g. a custom_providers:/providers:
entry deleted, or a provider renamed oldone->newone), Desktop/TUI resume
restores the stale name into agent init and dies with:
agent init failed: Unknown provider '<name>'
while the CLI resumes the same session fine with the configured default.
- runtime_provider: add is_routable_provider() (full resolution chain:
built-in -> providers: -> custom_providers: -> models.dev)
- _stored_session_runtime_overrides: heal a non-routable provider via
canonical_custom_identity (base_url -> model -> configured provider),
drop to the configured default when unrecoverable, and clear the stale
base_url after healing so a dead endpoint cannot override the registry URL
- _start_agent_build: gate deferred-resume overrides on provider routability;
when the stored provider is gone, prefer the model the user picked for THIS
session, else the configured default
- tests: is_routable_provider cases, heal/fallback round-trips, gate checks
Refs #75128
Test-pollution class: runtime_provider is usually imported lazily (inside
switch_model's resolution path), so its first import in a pytest worker can
happen while a test has hermes_cli.config.load_config patched. The
module-level from-import then bound the MagicMock permanently — after the
patch exited, every later caller in the process silently read the dead
test's config. Live victim: MoA aggregator context-length resolution
(resolve_runtime_provider -> AuthError 'Unknown provider custom:example'),
making TestMoAContextLength::test_moa_custom_context_configures_compressor_threshold
fail whenever it shared a process with
TestLocalOllamaModelDiscovery::test_switch_model_on_current_ollama_custom_endpoint_keeps_base_url.
Fix: load_config / get_compatible_custom_providers / normalize_extra_headers
become late-bound delegates resolving hermes_cli.config attributes at call
time. Both patch targets (config.load_config and
runtime_provider.load_config) keep working. Regression tests pin the
late-binding property and fail if the delegates revert to from-imports
(sabotage-verified).
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.
Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
/simplify-code reuse finding: _write_marker reimplemented the
mkstemp->write->os.replace pattern that utils.atomic_write_text already
provides as the repo's shared atomic-text-write helper (and the shared
version adds fsync + cross-device/busy-file fallbacks).
Follow-up to the #95605 salvage, closing the review findings:
- _copy_alias no longer swallows OSError silently: it warns (a leftover
alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
(update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
it now asserts the whole layout (anchor + aliases) is complete, so a
partially-materialized alias set can never read 'active' in doctor —
the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.
5 new regression tests.
Review fixes from kokhlo's live-hardware review:
- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
PYTHONHOME=<venv> boots a staged copy that would otherwise die with
"No module named 'encodings'", papering over the exact prefix
failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
images) still skip; EACCES after our own chmod now refuses the
install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
so the managed-runtime layout (cpython-3.11-macos-* symlinked to
cpython-3.11.15-macos-*) no longer reports stale on a fresh install.
Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).
Re-land:
- Keep the signed real-file copy of bin/python (identifier-pinned
via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
symlinks. Copies boot on every build we could reproduce and keep
the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
plus the venv prefix, abort and leave the live venv untouched
on failure.
Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.
Closes#95596.
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Unlimited sessions used a no-op lease, so a sibling profile backend could
not see that the same durable session was still owned. Track liveness in
the profile registry without imposing a cap, and fail closed when the
registry cannot be inspected.
Co-authored-by: metamindedu <metamind@kakao.com>
/simplify-code follow-ups on the 90502 salvage:
- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
byte-identical — saw_node was effectively write-only. Collapsed to one
arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
from dead PIDs at arm time (POSIX-only liveness probe; Windows never
creates AF_UNIX nodes) so state/ does not accumulate nodes across
os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
was DISPROVED for this call site — asyncio's create_unix_server
os.remove()s an existing node before binding — but the contract is now
pinned by test_producer_rebinds_over_stale_socket_node (a live
producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
longer leaks a directory per test run.
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.
WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.
New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
the loop is frozen for longer than one tick timeout but shorter than
the wedge window -> probe is ALIVE and launchd_restart drains, never
escalates; loop frozen for longer than the window -> WEDGED.
The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.
The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).
The classifier is now two-witness:
- socket answers -> ALIVE (file age irrelevant: a stalled write
or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
agree the loop stopped scheduling)
- legacy payload (no flag) -> unchanged single-witness contract: the
legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity -> UNKNOWN, never escalate
Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
PyYAML loads unquoted names like provider: 2070 as int. GET /api/model/options
then called .lower() on the providers dict key and 500ed, and activate/delete
looked up "2070" and missed. Stringify identity fields so Desktop can assign
and remove those endpoints.
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):
- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
ignored mode'); with it, every ZIP update was refused as 'could not
check the working tree'. The mocked tests could not see this — a new
real-git test creates an actual repo + .gitignore and asserts the guard
runs clean, blocks on an ignored user file, and exempts ignored
preserved entries. --ignored=matching also reports an ignored dir as
one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
rename/copy status codes. Porcelain v1 does not quote plain filenames
with spaces, so an ignored file literally named 'venv -> node_modules'
parsed as two preserved tops and slipped past the guard into the
destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
instead of a comment-synced duplicate set (change-detector test added).
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.
Credit: @JoaoMarcos44, whose #87392 included this hardening.
The generated unit only counted restart_drain_timeout, so a default
cron drain (30s + 10s cleanup) could still be inside budget when
systemd SIGKILLed the cgroup. Size TimeoutStopSec from max(drain,
cron floor + reserve) plus headroom so an in-budget stop is not killed.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.
`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:
1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
`_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
used by `systemd_restart()` and the updater — had no launchd call site.
2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
systemd branch uses `_get_restart_exit_wait_budget()`
(drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
documents that callers falling back to a hard kill must cover both phases
or they reintroduce #77184.
The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.
Observed on macOS 27.0 / Hermes 0.20.4:
→ Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
⚠ Gateway PID 49787 still running after 0.0s — restart may fail
⚠ Gateway drain timed out after 0s — forcing launchd restart
Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.
The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.
Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.
Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.
Refs: #79096
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
Add WeCom to the per-platform streaming defaults (DEFAULT_CONFIG display
plumbing) so native streaming is enabled by default for the WeCom adapter,
alongside the existing per-platform flags. Non-secret config lives in
config.yaml (no HERMES_* env vars).
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:
* `replace_messages(..., archive_dropped=True)` — the rewind / edit /
regenerate mode added in #82756 so a taken-back turn stays recoverable.
* `archive_and_compact` — in-place compaction, which archives the
pre-compaction transcript under the same session id (#38763).
A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.
Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.
Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.
Fixes#95868
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
_sqlite_connect opened a connection via connect_tracked and then ran the
busy_timeout PRAGMA; if that raised, the half-open connection was abandoned
— leaking its fd AND leaving a stale entry in the sqlite_safe_read
live-connection registry (which only clears on close), permanently blocking
byte-level probes of the kanban database. Close before re-raising.
Salvaged from PR #96290 (kanban slice) with regression test.
- _write_machine_sentinel_line: wrap the print() fallback so a closed
redirected stream (ValueError, not OSError) can't propagate out of the
ready path and kill a healthy serve; document that pythonw port
discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
redirect active all server logging lands on stderr, and an unread
stderr pipe can fill and block the child before the sentinel, flaking
the test at the 120s timeout
The same stdout redirect that rerouted the READY sentinel (#96282) also
reroutes the machine-parsed BACKEND_PORT_IN_USE sentinel printed by
_report_port_in_use() — both preflight and probe-to-bind-race callers run
after tui_gateway.server's sys.stdout=sys.stderr swap. Extract the fd-1
write into _write_machine_sentinel_line() and use it at both sentinel
sites; human-facing hint lines stay on print().
Since 6d4e851d8 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue #96282).
Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.
Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.
Delist credit: @orouge97 flagged this in PR #80036.
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
Allowlist hot-path hooks for abandon-on-timeout, keep subagent_stop on the caller thread, suppress re-fires of hung callbacks, and block tools when pre_tool_call times out.
Sibling site of the salvaged #95947 fix (same file): cron_status
declared 'Gateway is not running — cron jobs will NOT fire' from a bare
find_gateway_pids() miss even while the runtime lock proved the gateway
alive. Now the not-running verdict requires both the scan AND the lock
to read dead; when only the lock answers, the pid line falls back to
the recorded gateway pid (or is omitted).
Two regression tests pin the false-alarm suppression and the genuine
not-running warning.
Follow-ups to the salvaged #95947 cron commit:
- Wrap the lock probe in its own try/except: a crashing probe is
'unknown', not 'dead' — the pid scan still decides instead of the
whole tri-state collapsing to None.
- Regression tests (shape adapted from #94155 by @liuhao1024): lock
held + empty pid scan -> alive (the reported false alarm); lock
inactive -> pid-scan fallback both ways; crashing lock probe still
falls back.
- patch_liveness now pins the lock probe inactive by default so the
pre-existing pid-scan tests stay deterministic on machines where a
real gateway holds the real lock.
- contributors mapping for magnus.lundstedt@infidyne.com.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
_builtin_gateway_liveness() decides whether the builtin cron ticker can fire by
PID-scanning via find_gateway_pids(). That scan can transiently return empty even
while the gateway is up (e.g. just after a restart), so the in-gateway cronjob tool
emits a false 'Gateway is not running — jobs won't fire' while jobs are firing on time.
Prefer the gateway runtime lock: it is held for exactly the gateway's lifetime (a
reliable liveness signal), and inside the gateway process it short-circuits to True,
so the in-gateway check can never false-alarm. Fall back to the PID scan only when the
lock reads inactive (the external-CLI path).
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
The OpenRouter model picker builds its list from a curated set of model
IDs, then filters against OpenRouter's live catalog. minimax/minimax-m3:free
exists on OpenRouter (free tier, 1M context, tool-calling) but was missing
from both the in-repo fallback list and the remote catalog manifest.
Add it to OPENROUTER_MODELS and website/static/api/model-catalog.json so
the free variant surfaces in the picker alongside the paid one.
A hermes serve killed mid-update lost every un-flushed in-memory session
(#94724 item 2, reported by @ruangraung): the next RPC failed with
'session-scoped RPC rejected: not in memory (detached/reaped runtime)'
and no store held the transcript. #95576 made serves survive future
updates; this closes the kill path itself:
- install chaining SIGTERM/SIGINT handlers (hermes serve / dashboard
startup, before uvicorn's capture_signals) that first persist
in-memory session transcripts to state.db — bounded by
HERMES_TUI_EXIT_FLUSH_BUDGET_S (default 5s, daemon worker + join) so
a hung SQLite write can never block exit
- _shutdown_sessions (atexit) runs the same bounded flush FIRST, before
the slow per-session teardown a supervisor may SIGKILL mid-way
- the idle-reaper scan piggybacks a periodic incremental flush
(marker-deduped agent._persist_session, running sessions skipped) so
even a SIGKILL loses at most one flush interval — no new timer
subsystem
Refs #94724
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
Follow-ups to the salvaged #71385 guard (which raises RuntimeError from
require_readable_config_before_write on unparseable / non-mapping YAML):
- config_command: catch RuntimeError for set/unset and print a clean
one-line error + exit(1) instead of a raw traceback on the primary
'hermes config set/unset' CLI path.
- console_engine._capture_output: convert escaping RuntimeError into a
ConsoleCommandError so 'hermes console' and the dashboard console
report the refusal instead of crashing the REPL/websocket session.
- _warn_config_parse_failure: add a dedicated 'refuse-write' wording
branch — the old fallthrough claimed 'falling back to default config'
even though the write was refused and the file preserved.
- approval_mode: update the stale SystemExit-only comment.
- Regression tests for the console path and both config_command paths.