The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).
- Add optional tier fields to PricingEntry: tier_threshold_tokens,
input/output/cache_read_cost_per_million_above (None = flat, falls
back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
request once usage.prompt_tokens (input + cache read + cache write)
exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
above 200k).
- Flat entries are untouched: no threshold means no behavior change.
Reported and tier-field shape designed by @tornike14 (#93469).
Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.
Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.
- Add context_size and max_input_tokens to both /v1/models probe
candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
handled separately as an output cap (_MAX_COMPLETION_KEYS).
Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.
Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:
max_input_tokens = context window (e.g. 1M for claude-fable-5)
max_tokens = max output (e.g. 128k)
That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).
Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.
Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:
- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
exclusion so an explicit 'Rate limit exceeded' never promotes to
non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
reached. Resets in 6hr 29min.' wrongly read as terminal billing because
main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
signals. (credit @jtstothard, #74785)
Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.
Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.
Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.
Fixes#93188.
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.
bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.
The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.
Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
load_workspace_context() resolves the parent's workspace via the same
_resolve_workspace_hint used for child prompts (explicit sources only —
TERMINAL_CWD / agent cwd hints, never a bare getcwd fallback, so the
#64590 install-tree-leak guard concern doesn't apply) and runs it
through agent.prompt_builder.build_context_files_prompt — the exact
discovery/priority/cap logic the main system prompt uses (.hermes.md >
AGENTS.md chain > CLAUDE.md > .cursorrules; SOUL.md skipped). The
result is embedded in the reviewer briefing as binding review
standards. Subagents are built with skip_context_files=True, so without
this the reviewer judged repo work without the repo's own conventions.
5 new tests incl. real-filesystem AGENTS.md discovery through the real
loader. Docs updated (en + zh-Hans).
The reviewer subagent now inherits the primary agent's working skill
context: collect_parent_loaded_skills() gathers launch-preloaded skills
(from the activation notes in ephemeral_system_prompt) and mid-session
skill_view loads (from assistant tool_calls in history), deduped and
capped at 8, and the briefing instructs the reviewer to skill_view each
and treat their conventions as binding for the assessment.
Reference-file reads (file_path=...) don't count as loads; full-skill
injection was rejected as too costly (a single dev skill can be 40KB+).
Docs: delegation.md /review flow updated (en + zh-Hans).
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.
Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
Follow-up to the ContextVar-leak fix: the autouse fixture now drains on
both sides (drain(); yield; drain()) so earlier files can't pollute this
file's assertions either.
Running `pytest tests/agent/test_prompt_builder.py
tests/agent/test_system_prompt.py` failed
test_build_system_prompt_records_stable_prefix with AttributeError:
'...SimpleNamespace' object has no attribute '_emit_status'
(#93018). A truncation warning recorded by test_prompt_builder.py stays
in the shared thread context under plain pytest, so the later file's
build_system_prompt call drains a warning and forwards it to
agent._emit_status - which the test stub lacked.
Harden both sides:
- tests/agent/test_system_prompt.py: _make_agent() stub gains a no-op
_emit_status, so draining a stray warning is harmless.
- tests/agent/test_prompt_builder.py: autouse fixture drains pending
truncation warnings after every test, leaving the ContextVar clean.
The order-dependent failure no longer reproduces in either ordering.
Two suites still encoded the pre-#93093 contract that the three
structural no-op branches (insufficient_messages, no_compressible_window,
empty_post_handoff_window) increment _ineffective_compression_count:
- tests/agent/test_compaction_anti_thrash.py::
TestMinimumMessagesBranch::test_too_few_messages_records_an_ineffective_pass
- tests/run_agent/test_infinite_compaction_loop.py::
TestCompressNoOpRegistersIneffective::{test_no_op_increments_counter,
test_two_no_ops_block_should_compress}
Structural no-ops are transcript-shape facts, not evidence of an
incompressible floor, so they now arm _structural_no_op_backoff_until
and leave the strike counter untouched. Update the tests to pin the new
contract (count unchanged, backoff armed via time.monotonic(),
should_compress blocked while it holds) and rename accordingly. The
outcome contract of test_two_no_ops_block_should_compress is preserved:
repeated no-ops still block further automatic compression.
Fixes#93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).
Distinguish "nothing eligible right now" from "attempted and
underperformed":
- New transient _structural_no_op_backoff_until (in-memory, 300s)
armed by _record_structural_no_op() at the three structural no-op
sites: insufficient_messages, no_compressible_window,
empty_post_handoff_window. No strikes accumulate; auto-compaction
resumes on its own once the backoff lapses or the transcript outgrows
the protection window.
- The backoff gates should_compress via
_automatic_compression_blocked_locally and surfaces in
_compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
never shrink retries at most once per backoff window instead of
every turn.
- force=True (/compress) clears an active backoff before attempting;
record_completed_compaction() lifts it - both prove the transcript
is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
durable ineffective counter unchanged.
Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
(already wrapped, no 400 path), and $ref without '#/' prefix (stays
structured).
- Document the deliberate structural (false-positive-tolerant) detection and
its O(n) cost in the helper docstring.
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).
tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.
Adds regression tests for the wrap path and the unchanged structured path.
Adversarial-review fixes for the #93057 snapshot-compaction PR:
- Fail-closed detachment: only re-enable compression after
bind_session_state successfully severs the engine's parent binding.
A failed rebind keeps the historical compression_enabled=False
behavior and warns, instead of running compaction against a
compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
preflight + pre-API pressure check) until the fork's first provider
response, so the first request replays the full snapshot as the
intended cached read and compaction applies from the second request
on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
pre-fix code) and the existing threshold-crossing test reworked to a
two-request review asserting the warm first request + compacted
second request. 116 tests green across all touched suites; ruff
clean.
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.
Closes#93057
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.
- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
_AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
of fallback-providers and the delegation /review section missed in
#93339
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).
Two changes:
- image_routing._should_probe_ollama_vision now takes the API key and
forwards it; a new _resolve_inference_api_key mirrors
_resolve_inference_base_url's resolution order (runtime value,
model.api_key, providers blocks) so the key always matches the URL
being probed.
- detect_local_server_type caches a None verdict in memory with a short
failure TTL (5 min, vs 1h for positives) so the next turn is served
from the negative entry instead of re-running the waterfall — while
a transient failure (server starting, key being fixed) recovers in
minutes. Negative verdicts are deliberately not written to the
cross-process disk cache.
Fixes#89863. With a custom: provider pointing at a remote, API-keyed
endpoint (sglang/vLLM/OpenAI-compat), every image turn triggered a 5-request
probe waterfall without Authorization, spraying 401s at the backend.
Two fixes:
1. _should_probe_ollama_vision now takes api_key and forwards it to
detect_local_server_type so keyed local servers don't 401.
2. When provider != 'ollama', remote endpoints (per is_local_endpoint) are
rejected early — server-fingerprint probing is only valid for local
boxes. Non-Ollama remotes expose Ollama-compat endpoints that can
misidentify and trigger unnecessary /api/show probes.
_lookup_supports_vision resolves the runtime api_key via
_runtime_main_value and forwards it to both helpers. New test class
TestShouldProbeOllamaVision covers the contract in both directions.
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.
- agent/review_engine.py: shared engine (snapshot, briefing,
auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
(never model-facing) resolved through the same credential system as
delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
slash_commands handler (binds the approval session key so the
completion routes back), TUI/Desktop live dispatch in
tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
dispatch tests fail without the fix), 4 gateway handler tests
through the real async rail
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.
Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.
New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.
- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
the -900k section.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
CPython interns identifier-like string literals, so 'is' cannot
distinguish an import alias from a copy-pasted literal (verified:
two exec'd namespaces each defining the literal share one object).
The equality assertions three lines above are the full honest guard.
Also reword a comment: raw == is marker-SENSITIVE, not asymmetric.
Review-pass follow-up on the load-time durability stamp:
- hermes_state.py: import the marker from agent.context_compressor instead
of a third synced literal (hermes_state already imports agent.* at module
level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
fill-empty-tail pop site with the shared constant (was outside the drift
guard).
- agent/conversation_compression.py: the no-op progress check now falls back
to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
rows are stamped at materialization time while compress() output is
marker-swept, so a semantically-identical no-op copy on a cold-resumed
session would previously compare unequal and take the progress branch.
Raw == still runs first so engine-returned list subclasses keep their
__eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
assertions; new test_noop_progress_check_is_marker_insensitive
(mutation-checked: fails when the helper is neutered).
Resumed sessions loaded message dicts from state.db WITHOUT the
_DB_PERSISTED_MARKER, so any flush that lost the identity boundary
(compression durable-snapshot adoption, incremental tool-call persists,
rotation preflight on cold resume) re-appended the ENTIRE loaded
transcript as new rows. Compression cycles then doubled the copies:
the incident session grew 998 -> 1995 -> 3990 -> 7981 rows across
three aborted rotations (15,962 active rows, only 472 distinct).
Fix at the architectural chokepoint: SessionDB._rows_to_conversation
(shared by get_messages_as_conversation and get_resume_conversations)
now stamps the marker at row materialization time - a dict built FROM
a durable row is persisted by construction, regardless of which caller
loads it or how the list is later handed to a flush.
Safety:
- Wire-safe: every transport strips underscore-prefixed keys before
the API request (chat_completion_helpers, anthropic_adapter), same
contract as the existing _row_id stamp in the same function.
- Rotation handoffs still write: compression's assembly copies strip
the marker (_fresh_compaction_message_copy + the terminal
_strip_persistence_markers sweep), so compacted transcripts still
flush to the child session (#57491 invariant preserved).
- Branch/seed copies unaffected: /branch and _persist_branch_seed
build fresh field-projected dicts and write via append_messages_batch
directly, not through the marker-gated flush.
Tests: new regression suite (marker sync, load stamping, 3-cycle
amplification repro, new-tail write guard, compaction-copy handoff);
updated the #68454 control test that asserted the old double-write
behavior and the ACP restore shape test.
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.
Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
module scope, before any test can stub sys.modules — later imports are
cache hits that can never re-execute the vendored import under a
poisoned parent. Proven standalone: fake-parent repro fails without
the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
plant fake boto* modules (adapter, integration, model-picker):
snapshot every boto* sys.modules entry before each test, evict+restore
after — no stub window can leak state into a later test regardless of
worker ordering.
- importorskip targets botocore.exceptions (the module the tests
actually need) instead of bare botocore, so a torn install skips
instead of erroring.
148/148 across the four affected suites.
- Drop the 'aborted before its tail' no-op sentence: early aborts are
intercepted by the aborted/no-progress branches and never reach the
would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
armed runway after the refusal — locking in the memory==disk half of
the contract, not just the in-memory value.
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:
- the next eligible iteration's proactive prune fired without the regrowth
interval #79640 introduced — an immediate, unthrottled cache-breaking
rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
from the stale durable row.
The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.
Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).
Fixes the reachable half of the structural asymmetry flagged in #91830.
- credits_tracker: trim inline comment block (duplicated docstring) and
correct its safety claim - a paid model under stealth/ would fail
closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
recognize stealth/ prefix (same bug class as #91843: free_only=true
wrongly skipped the OpenRouter fallback and the paid-lane warning
fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
my-stealth/model not).
Stealth-preview SKUs (e.g. stealth/ox-alpha) are free-tier but carry no
:free suffix, so is_free_tier_model() returned False for them. On gateway
sessions (which never run the model picker's pricing fetch), the free-model
suppression of the credits.depleted banner never engaged, and any response
carrying paid_access:false triggered a false "Credit access paused" notice.
Add stealth/ prefix detection to is_free_tier_model() as a zero-network
signal, same design as the existing :free suffix check. Fail-open to
False (banner still shows) if the prefix changes — recoverable noise,
never a masked depletion on a paid model.
Closes#91843
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.