/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.
- agent/review_engine.py: shared engine (snapshot, briefing,
auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
(never model-facing) resolved through the same credential system as
delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
slash_commands handler (binds the approval session key so the
completion routes back), TUI/Desktop live dispatch in
tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
dispatch tests fail without the fix), 4 gateway handler tests
through the real async rail
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.
Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.
New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.
- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
the -900k section.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
CPython interns identifier-like string literals, so 'is' cannot
distinguish an import alias from a copy-pasted literal (verified:
two exec'd namespaces each defining the literal share one object).
The equality assertions three lines above are the full honest guard.
Also reword a comment: raw == is marker-SENSITIVE, not asymmetric.
Review-pass follow-up on the load-time durability stamp:
- hermes_state.py: import the marker from agent.context_compressor instead
of a third synced literal (hermes_state already imports agent.* at module
level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
fill-empty-tail pop site with the shared constant (was outside the drift
guard).
- agent/conversation_compression.py: the no-op progress check now falls back
to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
rows are stamped at materialization time while compress() output is
marker-swept, so a semantically-identical no-op copy on a cold-resumed
session would previously compare unequal and take the progress branch.
Raw == still runs first so engine-returned list subclasses keep their
__eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
assertions; new test_noop_progress_check_is_marker_insensitive
(mutation-checked: fails when the helper is neutered).
Resumed sessions loaded message dicts from state.db WITHOUT the
_DB_PERSISTED_MARKER, so any flush that lost the identity boundary
(compression durable-snapshot adoption, incremental tool-call persists,
rotation preflight on cold resume) re-appended the ENTIRE loaded
transcript as new rows. Compression cycles then doubled the copies:
the incident session grew 998 -> 1995 -> 3990 -> 7981 rows across
three aborted rotations (15,962 active rows, only 472 distinct).
Fix at the architectural chokepoint: SessionDB._rows_to_conversation
(shared by get_messages_as_conversation and get_resume_conversations)
now stamps the marker at row materialization time - a dict built FROM
a durable row is persisted by construction, regardless of which caller
loads it or how the list is later handed to a flush.
Safety:
- Wire-safe: every transport strips underscore-prefixed keys before
the API request (chat_completion_helpers, anthropic_adapter), same
contract as the existing _row_id stamp in the same function.
- Rotation handoffs still write: compression's assembly copies strip
the marker (_fresh_compaction_message_copy + the terminal
_strip_persistence_markers sweep), so compacted transcripts still
flush to the child session (#57491 invariant preserved).
- Branch/seed copies unaffected: /branch and _persist_branch_seed
build fresh field-projected dicts and write via append_messages_batch
directly, not through the marker-gated flush.
Tests: new regression suite (marker sync, load stamping, 3-cycle
amplification repro, new-tail write guard, compaction-copy handoff);
updated the #68454 control test that asserted the old double-write
behavior and the ACP restore shape test.
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.
Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
module scope, before any test can stub sys.modules — later imports are
cache hits that can never re-execute the vendored import under a
poisoned parent. Proven standalone: fake-parent repro fails without
the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
plant fake boto* modules (adapter, integration, model-picker):
snapshot every boto* sys.modules entry before each test, evict+restore
after — no stub window can leak state into a later test regardless of
worker ordering.
- importorskip targets botocore.exceptions (the module the tests
actually need) instead of bare botocore, so a torn install skips
instead of erroring.
148/148 across the four affected suites.
- Drop the 'aborted before its tail' no-op sentence: early aborts are
intercepted by the aborted/no-progress branches and never reach the
would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
armed runway after the refusal — locking in the memory==disk half of
the contract, not just the in-memory value.
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:
- the next eligible iteration's proactive prune fired without the regrowth
interval #79640 introduced — an immediate, unthrottled cache-breaking
rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
from the stale durable row.
The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.
Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).
Fixes the reachable half of the structural asymmetry flagged in #91830.
- credits_tracker: trim inline comment block (duplicated docstring) and
correct its safety claim - a paid model under stealth/ would fail
closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
recognize stealth/ prefix (same bug class as #91843: free_only=true
wrongly skipped the OpenRouter fallback and the paid-lane warning
fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
my-stealth/model not).
Stealth-preview SKUs (e.g. stealth/ox-alpha) are free-tier but carry no
:free suffix, so is_free_tier_model() returned False for them. On gateway
sessions (which never run the model picker's pricing fetch), the free-model
suppression of the credits.depleted banner never engaged, and any response
carrying paid_access:false triggered a false "Credit access paused" notice.
Add stealth/ prefix detection to is_free_tier_model() as a zero-network
signal, same design as the existing :free suffix check. Fail-open to
False (banner still shows) if the prefix changes — recoverable noise,
never a masked depletion on a paid model.
Closes#91843
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.
Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
Address review feedback on #65076:
- Add resolve_bedrock_runtime_region() to agent/bedrock_adapter.py: the
config-first region resolution (bedrock.region in config.yaml, then
AWS_REGION/AWS_DEFAULT_REGION/botocore profile/us-east-1) that the main
runtime resolver uses, exposed as a shared helper.
- Switch auxiliary client resolution (agent/auxiliary_client.py aws_sdk
branch) to the new helper. Previously it derived its region with bare
resolve_bedrock_region() (env-first), so when config.yaml pinned
bedrock.region to a different region than the ambient AWS env, auxiliary
calls (compression, memory, vision) left the primary runtime's region.
Both the AnthropicBedrock/Converse path and the new Mantle OpenAI
Responses path now resolve identically to the main runtime.
- Add regression tests covering the bedrock.region-vs-AWS_REGION mismatch
for both the Claude auxiliary path and the Mantle auxiliary path.
- Update website/docs/guides/aws-bedrock.md: the guide claimed Hermes never
uses the OpenAI-compatible endpoint, which the Mantle route made stale.
Document the triple routing (AnthropicBedrock / Mantle OpenAI Responses /
Converse), the Mantle auth model (bearer token or SigV4), and add the
GPT-5.5/5.6 model IDs to the models table.
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:
- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
so runtime resolution, auxiliary calls, and MoA slots all take the
SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
additions do not require test surgery; add routing, picker, and
context-length coverage for the 5.6 family.
Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
The Codex OAuth backend (chatgpt.com/backend-api/codex) intermittently
injects prompt_cache_retention into its own upstream call and then rejects
it, returning HTTP 400 invalid_parameter. Hermes never sends that field on
this route (see agent/transports/codex.py::_default_prompt_cache_retention_
for_request, which only sets it for api.meta.ai and bedrock-mantle hosts).
Reproduced live: a minimal 1-message request carrying no cache parameters
at all failed 4/20 (20%) with this error, so the rejection is not
deterministic and retrying the identical request is the correct recovery.
Previously the catch-all in _classify_400 returned format_error/
retryable=False, which tripped the is_client_error abort gate in
conversation_loop and killed the turn on the first attempt - burning an
entire large-context request (~550k tokens) per failure.
Classify these as retryable server_error (should_compress=False - the
request shape was never the problem). The same guard is applied to the
sibling 5xx request-validation branch, where a fronting proxy can surface
the identical rejection.
Deliberately narrow: keyed on parameters we only send on specific routes,
and skipped when the current provider is one that legitimately sends them,
so a genuine client-side bad parameter (max_tokens on GPT-5) still fails
fast as a format_error.
Cron already sets skip_memory=True and denylists the memory toolset.
The default cron toolset still names memory, so init treated that as a
request and built MemoryStore. MEMORY.md then landed in the job prompt.
Treat a denylisted toolset as not requested, and strip memory from the
cron enabled list. Flush agents that actually want the memory tool are
unchanged (#65429).
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
It exercises build_prompt_cache_plan's direct_native_tool_cache fallback,
not repeated apply on pre-decorated input, so it belongs with the other
plan-layout tests rather than in TestApplyIdempotency.
Follow-up to the #90972 salvage:
- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
is copy-on-write on content parts by contract (pops the top-level key,
rebuilds content lists/part dicts fresh), so a shallow top-level copy
preserves the caller-non-mutation guarantee — verified for all four
marker shapes — and removes the redundant second deepcopy the re-mark
path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
tests/agent/test_prompt_caching.py (where this module's tests live) as
TestApplyIdempotency; dropped the three tests that duplicated existing
coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
test_copies_sections_and_keeps_canonical_tools_plain which already
asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
no-tools fallback pins == 3 (marker loss can no longer masquerade as
safety); added the one new _can_carry_marker assertion (native=True
empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
where part-dict aliasing is the risk); fails on pre-fix base with
marker accumulation (9 > 4), passes with the fix.
sol-reviewer round-2 IMPORTANT: relay env vars had no scope
classification, so the two readers disagreed under a multiplexed
profile scope — gateway/config.py (scope-aware getenv) dropped a
process-env GATEWAY_RELAY_URL during the scoped runner reload while
gateway/relay's relay_url()/register_relay_adapter()/self-provision
(direct os.environ reads) still saw it. Result: adapter registered but
Platform.RELAY absent from config, so the connect loop never dialed
and direct adapters stayed up. The inverse split (profile-only stamp:
config enables RELAY, registration finds no URL) was equally dead.
GATEWAY_RELAY_* ROUTING stamps (URL, ENDPOINT, ALLOW_DIRECT_PLATFORMS,
PLATFORMS, BOT_IDS, ROUTE_KEYS, INSTANCE_ID, WAKE_URL, DISPLAY_NAME)
are now in _GLOBAL_ENV_EXACT: deployment config read from os.environ
under any scope, exactly like the API_SERVER listener settings
(#69379), so every reader resolves the same value. Relay AUTH material
(SECRET, ID, DELIVERY_KEY, IDP_*) is deliberately NOT global — it
stays profile-scoped with the fail-closed multiplex guard, mirroring
the non-secret/secret line the terminal env blocklist already draws
(tools/environments/local.py).
The round-1 multiplex regression test asserted the now-rejected
semantic (profile-scoped stamps win); it is inverted to pin the
global-stamp contract: a process-env stamp survives the profile scope
(sweep runs, matching registration), and a profile-only .env stamp
does NOT activate relay.
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
Route the model-supplied target through _bound_error_text so a huge
bogus target can't bloat context, and restore the "Use 'memory' or
'user'" hint. Follow-up to HexLab98's review note on the salvage.
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
Reuse the built-in store predicate during agent initialization and evaluate the config-backed memory tool check immediately after edits instead of applying the generic external-probe TTL.
Adaptive Claude models think by default, so omitting the `thinking`
parameter left thinking ON for users who had turned it off. Send
`thinking: {"type": "disabled"}` instead, and keep the omission for
reasoning-mandatory families that answer a disable with HTTP 400.
Review follow-up on the salvaged #89444:
- Warn fires only from the conversation-loop pre-API site, reusing the
unconditionally computed request_pressure_tokens (zero marginal cost,
covers turn-start AND mid-turn growth) — drops the duplicate every-turn
estimate the turn-context block paid.
- Turn-context block now only RE-ARMS the dedup once the session is back
under the window, so warn -> /compress -> regrow warns again (the dedup
was previously never cleared with compression disabled).
- Char pre-check treats non-string (multimodal) content as over-gate —
len() of a part list defeated the 20k char floor (probe: 10 'chars' vs
~70k real tokens) — and compares against the window, not a flat 20k.
- Deletes the unreachable get_model_context_length fallback from both
sites (context_compressor always exists; its context_length property
hard-floors positive; the fallback would have been a synchronous
network probe mid-turn that also bypassed config overrides) and the
undeduped inline _emit_warning fallback (third copy of the message).
- Tests bind the PRODUCTION warn/clear methods (previously a verbatim
fake reimplementation left them uncovered) and add dedup, re-arm,
no-rearm-while-over, and multimodal-gate coverage.
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.
Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
reduced only as a LAST resort after reasoning/tool/summary shrink ops,
and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
_PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).
Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.
Fixes#88568
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).
Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.
The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
kimi_supported_efforts() used exact/prefix matching and missed Kimi
Coding plan variants like k3-256k, which fell back to the K2-era
low/medium/high set and mistranslated efforts on a K3 wire. Replaced
with the boundary-token regex from #76427 (credit @ruizanthony), which
matches k3/k3-256k/kimi-k3* without matching kimi-k2.6 or mk3000.