- test_shutdown_watchdog: the non-POSIX arm now asserts the witness ARMS
over TCP (port published, no AF_UNIX call, no warning, no socket node)
instead of pinning the old witness-absent fail-safe.
- test_update_wedged_gateway: TestLoopTickTcpWitness exercises the
consumer probe against a real loopback listener — stale file + answering
witness stays ALIVE (#90502 shape), stale + silent is WEDGED, fresh +
silent is UNKNOWN, garbage port never counts as armed. Runs on the
Linux lane so the TCP path is not Windows-CI-only.
`asyncio.start_unix_server` does not exist on Windows (no AF_UNIX event-loop
support in asyncio), so arming the loop-tick witness in
`loop_heartbeat_forever` raised AttributeError on every native-Windows
gateway start. The broad except swallowed it and recorded
`loop_tick_socket=False`, so every stale-heartbeat probe classified the
gateway as UNKNOWN — never WEDGED, never ALIVE-with-stalled-write. The
two-witness interlock from a1c83ef9 (issue #90502 follow-up) has been
effectively disabled on Windows since it landed: a wedged native-Windows
gateway could never be detected, and an alive one could never be
distinguished from a stalled heartbeat write.
On non-POSIX platforms the witness now arms over a TCP loopback server on
127.0.0.1 (OS-assigned dynamic port) instead:
- same protocol — connect, read one byte "1"
- same semantics — pure in-memory, zero disk I/O, answered only while the
loop is dispatching, armed by the loop task itself (an awaited
`asyncio.start_server` is structurally loop-owned exactly like the Unix
variant, so a wedged loop cannot keep answering pings)
- the assigned port is published in the heartbeat payload as
`loop_tick_tcp_port`, and `probe_gateway_loop_liveness` prefers the TCP
witness when the producer published a port, falling back to the AF_UNIX
socket for POSIX/legacy producers
POSIX behavior is unchanged: the AF_UNIX arm (including the stale-node
sweep) stays gated behind `os.name == "posix"` so the missing attribute can
never raise on Windows again. Legacy heartbeats without `loop_tick_tcp_port`
keep the existing socket-node contract untouched.
Tested end-to-end on native Windows: witness arms, port is published,
`_probe_loop_tick_tcp` answers from an external thread while the loop
dispatches, and the existing loop-liveness suite passes unchanged (the
AF_UNIX structural test still passes — the Unix arm text is preserved
inside the POSIX branch).
Adds two tests pinning the new behavior: an E2E test that arms the TCP
witness and probes it (skipped on POSIX, where the Unix arm is the real
witness), and a structural test that the TCP arm stays awaited on the loop
task and the AF_UNIX arm stays POSIX-gated.
- ambientRequestFor(gateway): one adapter for the six copy-pasted
`<T>(method, params) => gateway.request(method, params ?? {})` lambdas
the pollers built to call requestForOwnedSession.
- resetBackgroundPollingGuard() with no argument (primary reconnect,
use-gateway-boot.ts) now also clears healsByStoredId. The latch and the
heal budget share a lifetime: a respawned backend re-mints every runtime
id, so a stored session that had exhausted its 3 heals must be healable
again. Previously only the latch was cleared there.
- Leaf module comment states the actual import convention.
The approvals loop and broadcast_session_info fan out UNSCOPED session.info
frames for every live session. The event router attributes an unscoped
frame to the active session, and approvalReplaySessionId then re-pulls
approval.pending for it. When the active runtime is already gone (4001),
that made every fan-out tick a fresh dead-id request — the dominant source
of the 1,369 post-restart approval.pending rejections in #100639.
approvalReplaySessionId now takes the frame's explicitness and a gone
predicate and returns null for an unscoped replay onto a latched runtime.
An explicitly scoped frame is the runtime speaking for itself and is never
skipped.
Refs #100639
The salvaged #95647 commits predate runtime-gone.ts and shipped their own
gone-latch (session-rpc-guard.ts) beside the one main already had. Fold
them: one latch, one classifier, one clear seam.
- session-gone-latch.ts is a dependency-free leaf holding the latch, the
4001 classifier (now also rejecting "mentions session not found" tool
strings and unwrapping IPC bridge prefixes), and the rebind seam. It
exists because session-request-router — imported by every store — must
clear the latch after a successful session.resume/activate without
pulling the session/tile stores into its import graph.
- runtime-gone.ts re-exports the leaf and keeps the heal logic.
- A successful rebind now also refunds the stored session's heal budget.
markRuntimeGone caps consecutive heals at 3 per stored id and only a
successful process.list refunded it, so a backend that reaps a detached
runtime a few times left the view stuck on a phantom id with every
poller latched and the socket-reconnect global clear (removed by the
salvaged commit) re-arming the storm. #100639: 1,230 approval.pending
4001s on one runtime id in 42 minutes, zero recovery.
Refs #100639
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.
Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.
Addresses #100313
The lazy gateway outbox was missing from the recovery inventory, so a
verified salvage could drop owed replies even when the rows were still
readable. Initialize the destination schema and copy the table.
Drives a real unopenable lock path — a directory where the code expects a
regular file, so open() raises a genuine kernel OSError — rather than
monkeypatching the helpers, standing in for the ENOSPC/EMFILE the field reports
hit without needing to fill a disk.
Both authorities are covered at the primitive and the behavior level:
fts_rebuild_admission refuses admission and rebuild_fts() reports no progress
(asserted against a preceding successful rebuild, so the 0 is the deferral and
not an unrelated no-op); _cross_process_repair_lock refuses the authority and
repair_state_db_schema runs no writable_schema surgery, takes no forensic
backup, and leaves the damaged image byte-identical for the next authorised
pass. A guardrail test pins that a pathless in-memory store is still admitted,
so the fix cannot turn that legitimate no-op into a permanent deferral.
All four deferral assertions fail on the pre-fix code, where the repair test
shows the surgery really did proceed without cross-process authority.
state.db has two cross-process admission authorities gating destructive work
on a file several Hermes processes share: fts_rebuild_admission for full
structural FTS rebuilds, and _cross_process_repair_lock for writable_schema
surgery / VACUUM. Both document themselves as fail-closed, and both honoured
that only for a timed-out acquire. When the lock file could not be open()ed at
all they yielded True and proceeded "with in-process serialisation only" —
which is no cross-process authority whatsoever.
That inversion is reachable exactly when it does the most damage. Creating the
lock file needs a directory entry and an inode, so on a full disk open() raises
ENOSPC — while a sibling that opened ITS handle before the disk filled is still
mid-rebuild or mid-surgery. Every process then ran concurrent destructive work
on the same live DB: precisely the interleaving PR #93200 added these locks to
prevent, and the shape reported in #100368 (disk-full trigger, then a fresh
corruption on every boot with other writers alive, and no re-corruption on a
boot with zero other writers).
Both helpers now yield False on OSError. This routes the error into the
outcome the locks already define and every caller already handles: rebuild_fts
returns 0, _recover_stale_fts leaves canonical writes plus LIKE search
available behind the retryable stale breadcrumb, the startup path detaches FTS
triggers, and repair_state_db_schema re-probes and reports. Nothing reachable
is lost — on a read-only directory the rebuild's and the repair's own writes
could not have committed either. The repair report's error string now names
both ways the authority can be missing, since operators read it directly.
Review point from keeltrace, and it is a real gap: secret redaction and
prompt omission are different contracts, and only the first one is
pattern-shaped.
redact_sensitive_text removes credentials. A provider 4xx that quotes the
request back carries the user's own prose - a paragraph about a person, a
file pulled in by an @ reference - which matches no credential pattern and
so passed through untouched into cause=. The record's stated contract is
that prompt content is not logged, and the previous commit only enforced
the half of it that a regex can see. The existing prompt test could not
catch this: its provider error does not echo the prompt, so it proves the
prompt is not logged directly, not that it cannot arrive by being quoted.
_strip_prompt_echo closes the quoted path directly. Anything the message
shares with the submitted prompt for 24 characters or more becomes
<prompt>. Shingle-set matching rather than a diff, so cost is linear in
both strings on a path that runs for every failed turn and can face an
@-expanded prompt of arbitrary size; the JSON-escaped form of the prompt is
shingled too, because a provider handing back its own request body often
hands it back escaped. The prompt is captured after @-expansion on purpose:
an injected file's contents are exactly the material an echo would carry,
and they are not in the submitted text.
Ordering is load-bearing. The strip runs after the whitespace collapse, so
a re-wrapped quote still matches, and before the length cap, so a quote
cannot survive by being cut mid-run.
What this does not claim: verbatim echo is what it stops. A paraphrase, a
summary, or a re-encoding would survive it. The alternative keeltrace
raised - log only structured provider metadata and drop the message body -
is airtight but costs the diagnosis this PR exists to enable, since the
reporter needed to tell a 402 from a crashed finalizer. Happy to switch if
maintainers prefer the stricter contract.
Tests: the non-secret sentinel keeltrace asked for (a benign phrase present
only in the prompt, echoed by the provider error, asserted absent from the
record), plus guards that a message sharing nothing with the prompt is
untouched, that an overlap below the window is not treated as an echo, that
a prompt shorter than the window cannot blank the message, that a
JSON-escaped echo is stripped, that the strip precedes the length cap, and
that whitespace shape does not hide an echo. The three that cover the new
path fail with the strip removed; the guards pass either way.
Fixes#89117
The whole of #89117 is two log lines:
tui_turn finished: ui_session=0dfcee58 status=error error_retained=True duration=0.9s
A provider 4xx, a budget wall, a billing block and a crashed finalizer all
produce exactly those characters, so an intermittent failure cannot be
triaged from the one record that is guaranteed to exist.
The bookend came from #86865, which added it to trace compression
rotations across #86647 -- identities and a coarse status were the job, and
content was deliberately excluded. What that leaves is a returned-error
path (provider 4xx, budget, billing) which writes no other log line at all.
The exception path at least prints `[gateway-turn] <Type>: <msg>` to
stderr, so the failures that go unlogged are exactly the sub-second ones
this issue is about.
Both failure paths now stash a one-line cause, and the bookend appends it.
The record keeps its shape when nothing failed: a successful turn gains no
new fields.
The cause is redacted with `redact_sensitive_text(force=True)` and capped at
240 characters with a visible ellipsis, because a 4xx body routinely quotes
the request that produced it -- adding the cause without redacting it would
write an Authorization header the user never chose to log. Redaction fails
closed: if the redactor cannot run, the fragment reads `<unredactable>`
rather than the raw message. Whitespace is collapsed so a multi-line
provider body cannot split the record, which is the only property that
makes it greppable for a bug like this one.
12 regression tests. Four mutations proven: disabling the helper fails 9,
dropping redaction fails 2, dropping truncation fails 1, wiring only the
exception path fails 4.
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.
Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):
- CompressionCommitFence gains mark_commit_watermark_fenced() /
commit_watermark_fenced; compress_context marks the fence right after
capturing get_active_message_watermark() under the durable compression
lock (#75316/#87484) — the property that makes a LATE commit safe: rows
appended after compression start survive both commit paths verbatim as
cloned concurrent tail (archive_and_compact watermark= and
publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
the detached worker (already kept alive via
_defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
user's turn proceeds on the uncompressed transcript at the same 10s
budget, and the summary is adopted at the worker's own watermark-fenced
commit boundary. Unfenced workers are cancelled exactly as before —
never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
block preflight adoption via the same-session cooldown); re-attempt
spacing is covered by the durable compression lock
(_session_has_compression_in_flight). If the worker ends WITHOUT
committing, a done-callback restores the flat non-escalating 60s
retry-after; a successful adoption resets the hygiene failure streak.
The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
to describe deferred adoption and the thinking-model case;
config_defaults.py comment updated. Knob stays config.yaml-only.
Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
start watermark; the fence still gates admission and unfenced/late
results are discarded.
New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
turn still released at the budget, no cooldown while running,
streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.
Fixes#97963
PR #99779 gave the streamed chat.completions consumer the host's absolute
compression deadline. The two wires that consume their streams internally
still ran on their own, always-larger budgets after the host gave up:
- Codex Responses: clamp the re-armable watchdog's hard ceiling to the
published host deadline, so a live (re-arming) stream is severed the
instant the host stops waiting instead of at max(600s, 4x timeout).
- Anthropic Messages: the per-event hook now raises at the host deadline
and on an explicit hard cancel; create_anthropic_message lets that
TimeoutError abandon the stream (the with-block closes it) instead of
swallowing it as a callback failure.
Sabotage-verified: without the Codex clamp the new deadline test hangs past
its 25s harness cutoff; without the Anthropic hook the three Anthropic tests
fail.
CompressionCommitFence.set_total_ceiling_seconds documents its deadline as
"shared by the host and worker", but only the host ever read it. The worker's
streamed summary bounds itself with _aux_stream_total_ceiling() instead —
max(600, 4 * aux_timeout) — which is >= the host's total ceiling for every
configured timeout AND starts counting later (after pool admission,
_serialize_for_summary, prompt build and TTFT). A stream that outlives its
abandoned host is therefore not an edge case; it is the guaranteed outcome of
every total-ceiling timeout.
8207862212 closed the first half: a cancelled fence now releases the
compression owner, freeing its pool slot and session lease. Its own comment
leaves the second half open — the isolated provider daemon that holds the
socket keeps streaming "until the auxiliary stream's longer absolute ceiling
expires". With the #99692 reporter's auxiliary.compression.timeout: 600 that
is 2400s of an orphaned ~500K-token summary the fence is already guaranteed to
refuse, and because the session never shrank, every following turn stacks a
fresh orphan on top of the last.
Publish the fence's deadline as an absolute monotonic instant
(CompressionCommitFence.deadline_monotonic) and give the auxiliary layer the
return leg it was missing: aux_stream_deadline() installs it thread-locally,
_ChatStreamAccumulator.feed() stops the stream once it passes, and
_run_protected_sync_provider_call propagates it onto the provider daemon
(thread-locals do not cross that boundary, so an owner-thread-only install
would be inert on exactly the path large-session compression takes).
Absolute, not relative: the deadline is unaffected by however long dispatch and
TTFT took before the accumulator was constructed. Checked as well as — not
instead of — the existing ceiling, so every caller without a host deadline is
byte-for-byte unchanged, and the "timed out" phrasing keeps _is_timeout_error
classification identical to a request timeout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.
Fixes#73175
Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.
The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.
Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.
The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.
Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
exhausted, parked, tools deregistered -- no respawn loop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/proc/<pid>/task/<tid>/children is per-thread. stdio_client() spawns the
MCP subprocess from the background loop thread, so the main-thread-only
read returned an empty set on Linux and _stdio_child_pids/_stdio_pids
never tracked the child: the #81995 dead-child fast-fail, the #96452
respawn signal and the killpg shutdown sweep were all no-ops. Union the
children of every task instead.
The existing test used model='glm-5.1:cloud' which now returns False from
_is_ollama_glm_backend() — the :cloud guard introduced in the fix.
- Rename the existing test to use 'glm-4-9b' (local GLM, no :cloud suffix):
the 3-call continuation path is still exercised for local backends.
- Add test_ollama_glm_cloud_stop_after_tools_does_not_request_continuation:
model='glm-5.1:cloud' at :11434 — asserts stop is honoured at face value
(2 API calls, no synthetic continuation nudge).
Resolves the conflicting regression noted in #98415 review.
Ollama cloud models (model name contains ':cloud') run generation on
Ollama's hosted server; the local 11434 endpoint is only a transparent
proxy that forwards finish_reason faithfully. _is_ollama_glm_backend()
was matching these models because the proxy listens on the same port
as local Ollama, causing _should_treat_stop_as_truncated() to rewrite
a correct finish_reason='stop' into 'length'.
The 4-attempt continuation loop then injects a synthetic user nudge
that reasoning-capable GLM models spend their output budget deliberating
over, producing unpunctuated tails that re-trigger _has_natural_response_ending()
rejection -- a self-reinforcing loop that always exhausts retries.
Fix: add an early return in _is_ollama_glm_backend when the model name
contains ':cloud'. Local GLM inference (no ':cloud') is still caught
by the existing port/URL/provider checks.
Fixes#98406
Two compounding bugs that cause WebUI to discard or misrender agent
responses when using GLM models on Ollama Cloud:
1. _is_ollama_glm_backend() matched "ollama" in base URL, which
included Ollama Cloud (ollama.com). The hosted service correctly
reports finish_reason and is not affected by the local Ollama
stop-reason bug. Exclude "ollama.com" before the substring check.
2. _handle_session_chat_stream() hardcoded "partial": False in the
assistant.completed SSE event instead of reading result.get("partial").
The WebUI could not detect truncation and rendered partial responses
incorrectly (showing only the continuation instead of the full text).
Read the partial flag from the agent result, matching the pattern
used by other SSE paths in the same file.
Fixes#72316
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').
- gateway/run.py: run both hygiene executor hops (detached-agent path and
codex app-server path) inside copy_context().run, keeping the default
executor so a fence-cancelled hung summary never occupies a gateway
agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
class failure — abort and preserve the session instead of dropping the
middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
run_in_executor); UnscopedSecretError classified as access failure.
Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
Some bundled CPython runtime builds strip stdlib ThreadPoolExecutor's
copy_context() propagation, so work submitted to the daemon pool runs in a
bare context. Under the multiplexed gateway this dropped the profile
secret scope in pool workers: the context-compression timeout fence
resolved auxiliary provider keys (SURPLUS_API_KEY) with
UnscopedSecretError, silently degrading LLM compression to lossy
deterministic summaries and driving re-read loops in affected sessions.
Restore stdlib semantics in submit() by snapshotting the caller's context
and running the callable inside it (a no-op re-application on runtimes
that already propagate). Mirrors the gateway's
_run_in_executor_with_context pattern.
Tests: daemon pool worker sees caller contextvars; scoped get_secret works
in a daemon-pool worker under multiplex while scoped misses still fail
closed (no env leak).
- generate() now passes kwargs.get("model") into _resolve_model(), so the
user's hermes tools pick (forwarded by the dispatcher as top-level
image_gen.model) is honored instead of silently dropped (#55893 class;
matches xai/krea/openrouter).
- Setup schema badge "internal" -> "paid" to match every other paid
image backend in the hermes tools picker.
- Tests: caller-model precedence, unknown caller model falls through,
model kwarg reaches the API payload, badge contract.
Adds a bundled image-generation backend for the Meta Model API
(https://api.meta.ai/v1), which is OpenAI-compatible. Exposes the
muse-image-1.0 model via the standard image_generate tool. This is the
image-gen companion to the already-bundled meta-ai chat provider
(plugins/model-providers/meta-ai, PR #88565).
- plugins/image_gen/meta-ai/ — provider registered as `meta-ai`, matching
the chat provider's id. Reuses the openai SDK pointed at Meta's base URL.
- Auth mirrors the chat provider: MODEL_API_KEY (Meta's documented var),
with META_API_KEY / META_MODEL_API_KEY aliases and a META_BASE_URL
override.
- Text-to-image only for now (capabilities gated); base64 (WebP) and URL
responses both handled and saved under $HERMES_HOME/cache/images/.
- Auto-loads as `kind: backend` and appears in `hermes tools` with no
central list edits, matching the other bundled providers.
- tests/plugins/image_gen/test_meta_ai_provider.py — 27 tests (metadata,
auth-alias resolution, base-url override, model resolution, generate
paths incl. b64 save, aspect mapping, URL caching, error handling).
- docs: image-generation feature page + provider-plugin built-in list.
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.
Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.
Fix: publish a per-request retirement token so the worker can tell it has been
retired.
- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
`agent._active_codex_stream_request_token` before handing off to the worker
(codex_responses only) and clears it at all four kill sites plus the worker's
own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
which can raise — every other call site wraps it in try/except, and a leaked
token would let a later worker mistake itself for the owning attempt. The
request-local `_codex_request_retired` mirror also swallows the transport
error our own force-close causes, so the worker's local error cannot replace
the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
`TimeoutError` from `interrupt_check` when it no longer owns the request —
raising rather than breaking, because a break returns the partial `final`.
The four stream callbacks also drop post-retirement frames so an abandoned
attempt cannot stream tokens into the live turn's bubble (the gateway caches
AIAgent instances per session).
`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.
Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.
model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).
Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
The fast-fail gate probed the stdio child watcher by CALLING it —
inspect.isawaitable(_watch_children()) — creating a fresh coroutine on
every stdio MCP tool call that was never awaited (RuntimeWarning spam +
gc churn). Inspect the function instead of invoking it.
Salvaged (unique hunk only) from PR #96044; the bundled
_stdio_children_dead polarity fix was already on main via #94339.
CI's deep pytest tmp_path pushes state/gateway.loop-tick.<pid>.sock past
the sockaddr_un limit; the producer swallows the OSError into
loop_tick_socket=False and the POSIX arm test fails falsely. Use a
mkdtemp under the temp root, same as test_update_wedged_gateway.py.
asyncio.start_unix_server does not exist on Windows, so the ungated
call raised AttributeError on every gateway start — a warning +
traceback in the logs each time, with the witness permanently absent
(the documented deliberate fail-safe). Gate the server creation inside
the existing os.name == "posix" block, mirroring the stale-socket
sweep above it, and log the non-POSIX skip at debug. POSIX behavior
and the two-witness liveness contract are unchanged.
Closes#96956
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>