The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.
The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.
Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.
Backward compatible: existing hook subscribers ignore unknown kwargs.
The preflight handler surfaces the boundary exception's per-request text
(token count, 'provider call was not sent') rather than the in-loop
_COMPRESSION_TIMEOUT_FINAL_RESPONSE constant, which describes a
different state (compression ran and could not reduce). Document the
divergence so it is not 'fixed' into a single message later.
When the turn-start fail-closed boundary (#98424) raises
PreflightCompressionTimedOut, the exception escaped run_conversation to
the surfaces' generic exception handlers. The gateway deliberately never
exposes raw exception text, so users saw 'Sorry, I encountered an
unexpected error... Try again or use /reset' instead of the boundary's
actionable guidance, and the compression_exhausted clean-session
recovery contract (#9893/#35809) never engaged.
Catch it at the build_turn_context callsite and convert it into the
same typed recovery dict the in-loop timeout consumers return
(salvaged #98741 / PR #99710): failed=True, partial=True,
compression_exhausted=True, turn_exit_reason=context_compression_timeout,
with the actionable message in final_response and error.
Regression test proves the exception no longer escapes and the typed
contract fields survive to the caller (mutation-checked: test fails on
main without the handler).
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):
- Per-session heal counters (heal events + messages healed) tracked by the
repair path in agent_runtime_helpers.py, session totals preserved across
10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
configurable via agent.sanitizer_heal_escalation_threshold in
config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
(events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
the conversation loop through _emit_warning (status callback -> gateway
status message / CLI print). Never injected into conversation context or
the wire copy: prompt caching, role alternation, and durable history are
untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
the /debug share // hermes debug report, and the config key surfaced in
hermes dump overrides. errors.log carries the ERROR line for `hermes logs
errors`.
Fill empty non-final user/assistant turns on the wire copy during send-time projection so the sanitizer does not re-heal the same poisoned row every call. When a caller still hits the owner, log WARNING then one ERROR per session window instead of flooding errors.log.
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:
- conversation_compression.py: threading.local + per-agent lock timeout
state (mark/reset/read helpers) upgrading #98424's simple attribute
where overlapping automatic/manual compression entrypoints matter;
the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
overflow (413/400 context_length_exceeded) recovery path end the turn
with the typed compression_exhausted recovery contract instead of
re-sending the unchanged oversized request and re-entering compression
in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
attempt; the #98424 turn-start check reads it through the typed helper.
Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995
- Add rolling status bar metrics:
- cache hit ratio (◈) delta since model/compression reset
(hit = cache_read / prompt, verified against live logs)
- avg latency (◷) and throughput (↑ t/s) over last 10 API calls
(deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
idle, focus, yolo, stash, battery, title
Missing/null -> all enabled (backward compat). Unknown keys ignored.
Title gated via right-align; stash/battery also gated.
- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
overflow trim preserved. Battery also respects display.battery.
No private data; mock data in tests.
Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).
New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.
Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.
Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.
Salvaged from #88960. Fixes#47339.
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.
Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes#76884.
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.
Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.
Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.
Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.
Two recovery layers, deliberately separate:
- Semantic split: new FailoverReason.image_corrupt with
_IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
_classify_by_message. Corrupt bytes route to strip-and-retry, never
to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
still contain image parts gets one strip-and-retry via the existing
_strip_images_from_messages helper, guarded by a new
stripped_images_this_turn one-shot flag on TurnRetryState. This
un-bricks the session for any current or future provider wording
without adding another pattern list to maintain.
Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.
Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.
Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:
- the self-improvement loop, which distils memories and skills by replaying
`messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
Hermes tool iterations, of which there are none.
So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.
The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
Adversarial-review fixes for the #93057 snapshot-compaction PR:
- Fail-closed detachment: only re-enable compression after
bind_session_state successfully severs the engine's parent binding.
A failed rebind keeps the historical compression_enabled=False
behavior and warns, instead of running compaction against a
compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
preflight + pre-API pressure check) until the fork's first provider
response, so the first request replays the full snapshot as the
intended cached read and compaction applies from the second request
on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
pre-fix code) and the existing threshold-crossing test reworked to a
two-request review asserting the warm first request + compacted
second request. 116 tests green across all touched suites; ruff
clean.
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.
Closes#93057
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
Follow-up to @BrunoBza's #93062:
1. Set failed=True only for the new repeated_outer_errors exit reason.
Previously the error exit left failed=False, so finalize_turn reported
completed=True for a turn that actually failed — incorrect.
2. Don't append_message the assistant response at the break. A thinking-
prefill or interim assistant may already be the tail, and appending
would create assistant→assistant role-alternation violation.
finalize_turn (lines 341-353) handles this safely by checking
_tail_role != 'assistant' before appending.
3. Update test to assert failed=True and completed=False for the
repeated_outer_errors exit.
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.
Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.
The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.
Fixes#92450
PR #93269 (kshitijk4poor) landed the outer-handler break for the same
symptom while this branch was in flight. Keep his guard (it covers
shutdown errors from local post-processing and does the resume-hint +
best-effort persist) and keep this branch's inner-retry-handler return
(it fires BEFORE the ⚠️ retry trace, credential rotation, and fallback
attempts that the outer handler never sees). Point his
_is_interpreter_shutdown_error at tools/interpreter_shutdown.py so the
class has exactly one text-matching site, preserving his RuntimeError
type gate and all 7 of his tests.
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
unconditionally computed request_pressure_tokens (zero marginal cost,
covers turn-start AND mid-turn growth) — drops the duplicate every-turn
estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
under the window, so warn -> /compress -> regrow warns again (the dedup
was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
sites (context_compressor always exists; its context_length property
hard-floors positive; the fallback would have been a synchronous
network probe mid-turn that also bypassed config overrides) and the
undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
fake reimplementation left them uncovered) and add dedup, re-arm,
no-rearm-while-over, and multimodal-gate coverage.
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
provider fallback (a deterministic request-shape failure that failover
cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
clamp+compress recovery as plain output-cap 400s.
Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):
- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
tracks the consecutive streak of (tool, canonical args, result-hash); on
the 3rd identical call a compact one-line notice is appended to that tool
RESULT at construction time (cache-safe — tool results are append-only).
Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
changed result, or new turn. Observed on the raw result before the
tool-loop warning suffix so its changing count can't defeat matching.
- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
short reply ENDING on an announced next action ('Let me now…', 'I will
now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
intent-ack continuation path (same interim-assistant + user-nudge
mechanism, same codex_ack_continuations cap of 2), preserving message
alternation — no parallel recovery machinery.
Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.