Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995
- Add rolling status bar metrics:
- cache hit ratio (◈) delta since model/compression reset
(hit = cache_read / prompt, verified against live logs)
- avg latency (◷) and throughput (↑ t/s) over last 10 API calls
(deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
idle, focus, yolo, stash, battery, title
Missing/null -> all enabled (backward compat). Unknown keys ignored.
Title gated via right-align; stash/battery also gated.
- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
overflow trim preserved. Battery also respects display.battery.
No private data; mock data in tests.
Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).
New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.
Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.
Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.
Salvaged from #88960. Fixes#47339.
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.
Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes#76884.
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.
Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.
Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.
Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.
Two recovery layers, deliberately separate:
- Semantic split: new FailoverReason.image_corrupt with
_IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
_classify_by_message. Corrupt bytes route to strip-and-retry, never
to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
still contain image parts gets one strip-and-retry via the existing
_strip_images_from_messages helper, guarded by a new
stripped_images_this_turn one-shot flag on TurnRetryState. This
un-bricks the session for any current or future provider wording
without adding another pattern list to maintain.
Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.
Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.
Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:
- the self-improvement loop, which distils memories and skills by replaying
`messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
Hermes tool iterations, of which there are none.
So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.
The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
Adversarial-review fixes for the #93057 snapshot-compaction PR:
- Fail-closed detachment: only re-enable compression after
bind_session_state successfully severs the engine's parent binding.
A failed rebind keeps the historical compression_enabled=False
behavior and warns, instead of running compaction against a
compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
preflight + pre-API pressure check) until the fork's first provider
response, so the first request replays the full snapshot as the
intended cached read and compaction applies from the second request
on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
pre-fix code) and the existing threshold-crossing test reworked to a
two-request review asserting the warm first request + compacted
second request. 116 tests green across all touched suites; ruff
clean.
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.
Closes#93057
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
Follow-up to @BrunoBza's #93062:
1. Set failed=True only for the new repeated_outer_errors exit reason.
Previously the error exit left failed=False, so finalize_turn reported
completed=True for a turn that actually failed — incorrect.
2. Don't append_message the assistant response at the break. A thinking-
prefill or interim assistant may already be the tail, and appending
would create assistant→assistant role-alternation violation.
finalize_turn (lines 341-353) handles this safely by checking
_tail_role != 'assistant' before appending.
3. Update test to assert failed=True and completed=False for the
repeated_outer_errors exit.
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.
Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.
The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.
Fixes#92450
PR #93269 (kshitijk4poor) landed the outer-handler break for the same
symptom while this branch was in flight. Keep his guard (it covers
shutdown errors from local post-processing and does the resume-hint +
best-effort persist) and keep this branch's inner-retry-handler return
(it fires BEFORE the ⚠️ retry trace, credential rotation, and fallback
attempts that the outer handler never sees). Point his
_is_interpreter_shutdown_error at tools/interpreter_shutdown.py so the
class has exactly one text-matching site, preserving his RuntimeError
type gate and all 7 of his tests.
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
unconditionally computed request_pressure_tokens (zero marginal cost,
covers turn-start AND mid-turn growth) — drops the duplicate every-turn
estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
under the window, so warn -> /compress -> regrow warns again (the dedup
was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
sites (context_compressor always exists; its context_length property
hard-floors positive; the fallback would have been a synchronous
network probe mid-turn that also bypassed config overrides) and the
undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
fake reimplementation left them uncovered) and add dedup, re-arm,
no-rearm-while-over, and multimodal-gate coverage.
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
provider fallback (a deterministic request-shape failure that failover
cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
clamp+compress recovery as plain output-cap 400s.
Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):
- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
tracks the consecutive streak of (tool, canonical args, result-hash); on
the 3rd identical call a compact one-line notice is appended to that tool
RESULT at construction time (cache-safe — tool results are append-only).
Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
changed result, or new turn. Observed on the raw result before the
tool-loop warning suffix so its changing count can't defeat matching.
- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
short reply ENDING on an announced next action ('Let me now…', 'I will
now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
intent-ack continuation path (same interim-assistant + user-nudge
mechanism, same codex_ack_continuations cap of 2), preserving message
alternation — no parallel recovery machinery.
Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.
Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.
Projection-side approach credit: @JoaoMarcos44 (PR #88996).
Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (#88955).
Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused #81841.
Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.
Refs #88955
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).
Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.
New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:
- Deterministic-empty detection: two consecutive empty attempts with
usage present, output_tokens == 0 (reasoning tokens count as output),
and identical (model, provider, finish_reason) skip the remaining
retries and go straight to the fallback chain — a different model may
well answer. Missing usage, nonzero output, or any signature change
keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
included/subscription routes are untouched.
At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.
Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.
Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.
Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.
Refs NS-503.
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.
E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.
- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
cache so new installs appear), persisted so the next turn reuses the
new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
"Conversation started:" date is dropped (timezone kept); no ticking
fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
agent (tool definitions are construction-baked) when the fingerprint
moves, same session id/history, with a user-visible notice
Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).
Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:
- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
status-less paths now attach error_context {billing_unverified,
possible_content_filter}. Reason stays FailoverReason.billing (rotation +
fallback remain the right recovery either way); ClassifiedError grows a
billing_unverified property.
- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
unverified billing exhaustion gets the short transient cooldown instead of
the one-hour bench, regardless of pool size: a content-filter rejection
leaves the credential healthy and fails identically on every key, and the
hour-long sole-credential latch is what replayed the stored error and made
real fixes look ineffective. A true 402 keeps the full bench. The marker
persists with the entry so a restart cannot upgrade it back to a bench.
- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
threads billing_unverified and hands the pool 'billing_unverified' as the
persisted failure_reason.
- agent/conversation_loop.py: the fallback-switch status, max-retries status,
terminal label, and both structured terminal results hedge when the verdict
is unverified. New _billing_terminal_label + _billing_failure_result build
the returned terminal response in one place; the result dict now carries
billing_unverified and the billing_block gains 'unverified': true. The
confirmed-billing path (a real 402 or an API-key credit depletion) keeps
the original assertive wording, so the caveat no longer dilutes it.
Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.
Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.
Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).
Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:
- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
reporter verified returns 200. Meaning, the skill_manage reference, and the
## Skill Safety Rule block are all preserved. The reword is empirically
validated rather than understood, so a comment records the bisect and warns
that any rewrite must be re-verified against an OAuth token, not an API key.
- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
longer asserts exhaustion as fact. It hedges the opening line, names the
content-filter alternative, and gives the operator a way to tell the two apart
(if the usage page still shows quota, suspect a content rejection). It also
points at `hermes auth reset anthropic`, because the credential exhaustion
latch replays the stored error for ~60 min without issuing a request — which
makes a real fix look like it did not work.
- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
not an API key, despite auth_type="api_key". It stays in api_key_env_vars
because that tuple doubles as the credential-discovery list; removing it would
stop Hermes finding a `claude setup-token` credential at all.
Docs updated to match the reworded prompt.
Fixes#82154
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.
- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
adopt only when tip != old_id AND the tip row is live, retry the flush
exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
lookup with the same tip + liveness contract, so gateway transcript
reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
recovery preflight now resolves via get_compression_tip with the same
liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
explanation names compression rotation and tells the client to refresh the
session id instead of blaming a full disk.
Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).
Closes#82001
Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
`_moa_prepared_request` is a private handshake between the conversation
loop and MoAChatCompletions.create. It is attached whenever
agent.provider == "moa", on the assumption that agent.client is still the
in-process MoA facade.
Credential rotation, provider fallback and dead-connection cleanup all
rebuild agent.client from _client_kwargs between attempts, and
pending_moa_prepared_request deliberately carries a prepared request
across exactly that boundary. The rebuilt client is a native OpenAI
client while provider stays "moa", so the key reaches an SDK that has
never heard of it:
TypeError: Completions.create() got an unexpected keyword argument
'_moa_prepared_request'
That error is non-retryable, so every remaining turn on the session
fails. Both dispatch paths are affected: the non-streaming one calls
agent.client directly, and _create_request_openai_client returns
agent.client unchanged for provider "moa".
Re-check the live client at the point the key is attached, which covers
both paths at once. When the facade is gone, send the prepared prompt
without the handshake and log the downgrade.
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.
Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.
Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.
Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.
Rückbau: Commit revertieren; kein Zustand, keine Migration.
(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
The diff-apply salvage introduced stale-base revert hunks — the PR was 1246
commits behind main, and its diff for conversation_loop.py and moa_loop.py
silently dropped symbols added after the PR's base (e.g.
_CODEX_ACK_CONTINUATION_NUDGE, _INTERRUPT_SCAFFOLD_MARKER, cache_ttl plumbing,
finalize_turn import, _restore_user_after_reference_handoff).
Restored both files to origin/main and re-applied only the PR's additive
changes: _moa_reference_metrics_for_hook, _system_prompt_for_hooks, the
system_prompt= and moa_references= hook kwargs, _last_reference_metrics
attribute and accessors, and the slot_metrics population in the fan-out path.
Fixes CI ImportError: cannot import name '_CODEX_ACK_CONTINUATION_NUDGE' from
'agent.conversation_loop'.
Salvaged from PR #83437 by @erosika, with adopted fixes from @bgodlin (#81054),
@aldoeliacim (#82332), @nftpoetrist (#42326), @rodboev (#39653), @FnExpress
(#64292, supersedes #32175 by @db-aeon), @Per0-1 (#61166), @NaMinhyeok (#64797),
and @liuhao1024 (#43130).
Widens the bundled Langfuse plugin from 6 to 11 hooks and fixes two
attribution bugs. Also adopts shutdown/atexit lifecycle fixes and composes
8 prior community PRs with interaction-fix follow-ups.
Model attribution: on_pre_llm_request and on_post_llm_call now prefer the
wire value (request body model, response model) over the agent attribute,
which goes stale after /model switch or provider fallback.
Cost total: both cost paths now send a summed total alongside the per-type
breakdown, since Langfuse does not derive calculatedTotalCost from
cost_details keys. Subscription-included routes send no cost keys at all.
New coverage: api_request_error closes failed generations with ERROR level;
on_session_finalize/on_session_end close dangling traces for tool-only and
interrupted turns; subagent_start/subagent_stop trace delegated children as
spans; MoA advisor fan-out emits one generation per advisor priced at the
advisor's own model.
Capture modes: HERMES_LANGFUSE_CAPTURE=metadata|sanitized|full (default
sanitized). Sanitized mode redacts secret patterns before truncation.
Adopted lifecycle fixes: shutdown client at session finalize when
reason=shutdown (not on session rotation); atexit finalizer ends open root
spans for short-lived processes; root context manager exited to prevent
interpreter-teardown TypeError; TOCTOU on _get_langfuse() fixed with lock;
reasoning_content surfaced in traces; system prompt included in generation
input for Anthropic/Codex/Bedrock; SDK v3 update_trace replaces set_trace_io.
Closes#29482, #43129, #72661.
Supersedes #81054, #82332, #42326, #39653, #64292, #32175, #61166, #64797, #43130.
Partially addresses #67544 (capture modes + secret redaction; user_id remains open).