Three inline context-manager classes shared the same __enter__/__exit__
boilerplate and differed only in the events yielded before the raise.
Also drop the unreachable 'or agent.base_url' fallback: every
anthropic_messages init path sets _anthropic_base_url, and the two
sibling call sites read it bare.
A tool_use block name recorded by a stream attempt that died before any
visible text survived into the next attempt: only the deltas_were_sent
mid-tool branch cleared result["partial_tool_names"]. When the retry then
streamed plain text and dropped, the partial stub blamed the stale tool
("Stream stalled mid tool-call (old_tool)") and the stale name could make
a later attempt look mid-tool-call when deciding whether the drop is
retryable. Reset it in _start_stream_attempt alongside
provider_tool_in_flight, which already has attempt-local semantics.
Regression: two-attempt stream (tool_use start + parse error, then text +
drop) — stub content and emitted deltas carry no stale tool name.
Follow-up to the cherry-picked fix: keep the classifier widening
("expected value at line" is now a transient stream parse error on the
main turn) but replace the messages.create() fallback with a retry on the
same stream wire.
Why not create(): the fallback ran outside Relay (lost request rewrites),
outside _handle_stream_error (could replace text already shown to the
user with a different generation), ticked no liveness events for the
whole buffered payload, and a bare identical retry still re-emits the same
malformed JSON.
Why not drop the fine-grained-tool-streaming beta (#108583/#109056): live
probe on claude-sonnet-4-5, ~500-line tool call - beta on: max inter-event
gap 1.6 s; beta off: 139 s zero-event gap while Anthropic buffers the
args, which the 180/240 s stale-stream detector kills on larger payloads
(the regression 80a899a8e2 fixed).
Instead, on a parse error the retry sets `eager_input_streaming: false`
on every tool for that request only (the SDK/API per-tool field overrides
the legacy beta header), so Anthropic returns buffered, server-validated
args while the happy path keeps fine-grained streaming. A tool_use that
started streaming is registered in partial_tool_names so the mid-tool
transient retry fires the same way it does on the chat_completions wire.
Under a loaded runner the primary stall plus the fallback retry overran the
0.2s total ceiling, so the retry never started and attempts==1 failed
intermittently (seen once in a 40-worker tests/agent run). Idle stays 0.05s;
the ceiling moves to 2s per the >=2s wall-clock rule in AGENTS.md.
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:
- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
key falls back to the conversation root); the fork now publishes exactly the same instead
of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
`_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
parent's PHYSICAL id, so after a compression rotation the review's Portal
`conversation=` tag fragmented usage attribution across one logical conversation.
Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
build_cache_parity_fork gives the same-model fork the parent's session_id,
cached system prompt, tools[] and session_start — but with
_persist_disabled=True and _session_db=None, BOTH cache-identity resolvers
diverged from the parent on their own: declared_conversation_scope failed
closed on _persist_disabled, and the lineage walk skipped on the missing
DB. The fork's affinity header (set_affinity_scope) and body
prompt_cache_key (cache_scope_id on the OpenAI-wire transports) therefore
keyed a different bucket than the gateway parent, costing one cold
~full-context request per review. Not gateway-only: any parent whose
lineage root != current physical id diverges too (teknium1's triage table).
Fix, per the triage's suggested direction: on the not-routed branch only,
the fork stamps _inherited_cache_scope = resolve_prompt_cache_scope_safe
(parent) — the parent's ALREADY-RESOLVED scope, no DB access from the fork,
persistence fully detached. Both declared_conversation_scope and
resolve_prompt_cache_scope return the inherited scope first when set, so
the header path and the body path are fixed together (fixing only one
leaves the other divergent — Vivamisu's header/body split observation).
Routed (different-model) forks, /branch children, delegate/tool children
and fresh sessions set nothing; the fail-closed default stands untouched.
/btw shares build_cache_parity_fork and gets the repair for free.
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).
Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.
Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
The review fork is a full AIAgent, not an auxiliary_client call, and its
routed branch deliberately skips the parent's reasoning_config (the parent's
effort vocabulary may be invalid for the routed provider). It also never read
the per-task key, so an explicit `auxiliary.background_review.reasoning_effort`
was silently ignored and the routed fork ran at the provider default (#94825).
Routed forks now parse the task key through the shared parse_reasoning_effort
(same levels and `none` alias as every other aux task); unset keeps the
provider default, an unknown level warns and falls through. The same-model
path is untouched: it still inherits the parent's reasoning_config verbatim
for prompt-cache parity.
Salvaged from #94832 (liuhao1024), re-applied on the decomposed
_fork_init_kwargs seam.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
auxiliary.background_review.reasoning_effort was a silent no-op on the
same-model path: the fork inherits the parent's reasoning_config verbatim to
keep prompt-cache parity (#30532), and nothing told the user. Emit a one-time
user-visible warning (parent-scoped, so a nudge-per-turn session warns once,
not per fork) when the key is actually set, document the no-op in the config
reference, and leave the fork-birth request bytes unchanged.
The Python-repr pass ran after the MCP probe header scrub and collapsed
'Authorization': 'Digest ***' to '***', erasing the scheme word that scrub
deliberately keeps (tests/hermes_cli/test_mcp_probe_redaction.py went red
on rebase). Skip values that already carry a mask marker.
Port from OpenHands/software-agent-sdk#4508: their dict-entry secret
redaction was uppercase-only and leaked mixed-case keys (UserPassword,
sessionToken). Apply the same case-insensitive treatment to the Python
mapping-repr pass: a casefolded credential suffix (apikey/token/secret/
password/passwd/credential) now qualifies a key, while metadata names
(TOKEN_COUNT, password_policy, tokenizer) stay untouched.
The streaming accumulator dropped delta.reasoning_details entirely — only
non-streaming responses preserved the OpenRouter unified reasoning replay
data (signatures, encrypted reasoning blocks). Providers that require the
reasoning_details sequence passed back on the next turn (OpenRouter
reasoning models, Anthropic signed thinking via shims) lost continuity on
every streamed turn, which is nearly all turns.
- Accumulate delta.reasoning_details (attr or model_extra) during
streaming and attach the merged list to the final mock message, where
_build_assistant_message's existing passthrough persists it.
- _append_streamed_reasoning_detail merges consecutive reasoning.text /
reasoning.summary fragments into one logical entry (OpenRouter streams
them as word-level deltas; unmerged they bloat the replayed signature
payload) while encrypted/opaque entries stay discrete. Later fragments
backfill signature/id/format/index fields the first fragment omitted.
Ported from earendil-works/pi#8605 (commit c5ad7c1b0), credit
@cristinaponcela for the merge-fragments pattern.
Tests: tests/run_agent/test_streamed_reasoning_details.py (8 tests;
sabotage-verified: disabling accumulation fails the E2E test, disabling
merging fails 4 unit tests). Neighboring suites green (test_streaming.py
40 passed).
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.
Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
vocabulary since #100479, so the "full accounting" row names the serve
unit instead of borrowing a gateway relaunch.
Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
round-trips, absent-registry, sequential "race" re-enactment.
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:
- tests/conformance/test_profile_write_tripwire.py — no writes to the
default profile tree while a profile is active (#88532#92662#89190#89625#92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
table for the Deprecated-.env warning (#88829#89016#89389#90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
strippable injected param x real 400 shapes must strip-and-retry;
unknown params must still fail (#90257#89897#91164#89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
enumerated maintenance ops refuse/degrade under a live writer; copies
of corrupt DBs are refused or flagged (#91839#90806#90613#88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
Bot Chat resolution is idempotent, never mints, unique per profile,
race-safe (#92040#90705#92692#90005#90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
crash never claims success; success requires full fleet accounting;
refusal != failure (#91283#91439#92902#92780)
117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
The prompt template is built from _PROMPT_GOOD_EXAMPLES, so asserting the
template contains them can never fail independently of the code it checks.
The behavioural invariants (echo rejected, greeting allowed, near-miss
passes) stay.
Small title models parroting a prompt example back verbatim produced
sessions named "Fix login button on mobile" with no relation to the
conversation. The example lines in _TITLE_PROMPT_TEMPLATE now render
from _PROMPT_GOOD_EXAMPLES so the guard set and prompt cannot drift,
and generate_title rejects exact (case-insensitive, wrapper-stripped)
echoes so the instant derived title survives instead. 'Friendly
greeting' stays allowed — it is prescribed output for bare greetings.
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).
- hermes_cli/models_catalog_static.py: remove both slugs from the
opencode-free offline floor and the opencode-zen discovery floor;
document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
invariant to cover both.
The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
OpenAI documents credit_balance_exhausted / *_spend_limit_exceeded /
organization_usage_limit_exceeded as HTTP 429 responses, so the invariant
must hold through _status_429 (which short-circuits _by_error_code), not
only on the status-less body path.
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
The streaming choke point was covered only by an identity check on the
imported alias (a change-detector that would pass with a no-op wiring).
Replace it with a real chunk-loop run asserting STOP -> stop and
MAX_TOKENS -> length on the assembled response, and collapse the
duplicate transport cases into one parametrized invariant.
Some OpenAI-compatible gateways fronting Gemini backends emit the native
uppercase finish reasons (STOP, MAX_TOKENS) instead of the lowercase
OpenAI contract values. Every downstream comparison in Hermes uses
lowercase literals, so an uppercase reason silently fell through: a
clean STOP completion missed the stop handling and a MAX_TOKENS
truncation never entered the length-recovery path.
Adds normalize_finish_reason() as the single owner in
agent/message_sanitization.py (case fold + alias map: max_tokens->length,
end->stop, function_call->tool_calls) and wires it at both wire-intake
choke points: ChatCompletionsTransport.normalize_response and the
streaming chunk-capture loop in chat_completion_helpers. Non-string and
empty values pass through unchanged so existing 'or "stop"' defaults
and the Poolside int-reason path keep their behavior.
Adding grep/awk/sed to _FILE_READ_COMMANDS made the PATTERN operand
participate in the secret-file predicate, so `grep .bashrc app.py` or
`grep -n .env src/settings.py` — reads of SOURCE files — ran the
ENV/YAML assignment pass and masked opaque values that main leaves
alone. Skip the first non-flag positional for the pattern-first
readers, as #109369 originally did, so only real file operands gate.
`cat $HOME/.hermes/config.yaml` was ungated because the `$` bail-out
fired before the `.hermes` segment was inspected; strip `$HOME/` and
`${HOME}/` like the HERMES_HOME prefixes.
Review follow-up on #110228.
Fold KoNit-K's `_command_reads_secret_bearing_file` and the pre-existing
`_command_reads_env_file` into a single `_command_reads_secret_file` so the
`code_file` gate in `redact_terminal_output` has one owner: `.env`-style
basenames and shell rc/profile files anywhere, `config.yaml` only under a
`.hermes` directory or `$HERMES_HOME` (arbitrary project YAML stays on the
code_file path). `grep`/`awk`/`sed` join the reader set instead of a second
table with a positional-argument special case: on a file read, any non-flag
operand that names a secret-bearing file is enough — the pattern/program
operand never matches a basename, so the extra rule bought nothing.
Tests: the negative parametrization now uses an opaque credential-shaped value
(the placeholder it used before would never have been masked on either path,
so the "stays unredacted" half proved nothing) and asserts the same value IS
masked under `cat .env` in the same test.
Keep the narrow basename allowlist, but do not treat $HERMES_HOME as an
unresolved path, and split pipelines only on unquoted |;&.
Co-authored-by: Cursor <cursoragent@cursor.com>
The text-only drop guard in _finish_chat_stream required content_parts,
so a stream that died while still emitting delta.reasoning (no
finish_reason, no usage) fell through to the synthesized "stop". With
the reasoning-only clean-stop promotion in finish_text_response that
stamped "stop" turned the truncated thought into the final answer,
where main entered the continuation ladder. Extend the guard with
reasoning_parts so the drop yields the partial-stream stub and the
ladder still runs; a real clean stop carries finish_reason="stop" and
is unaffected.
Review follow-up on #110227.
Follow-up to KoNit-K's commit: rebuild the promotion on the existing
`agent._extract_reasoning` helper (the same reader the ladder terminal and
`build_assistant_message` use) and write the promoted text back onto
`assistant_message.content` so the persisted assistant row carries the answer
as ordinary content. Without that the transcript tail was an assistant row with
empty content and only `reasoning`, which `drop_thinking_only_and_merge_users`
strips from the next request — the model would see its own answer vanish on a
"continue" turn.
Tests: trim to the two invariants (clean stop → one API call, persisted as
content; `finish_reason == "length"` → never promoted, continuation still
owns it) and keep the truly-empty terminal case. The prefill wire-payload
regression test now drives a non-clean-stop reasoning-only reply, which is the
only shape that still reaches the prefill rung.
check_vision_requirements only asked the auxiliary resolver, so a vision-capable
main model on a provider the resolver cannot serve (minimax-oauth, local vLLM,
anything uncatalogued) lost vision_analyze and browser_vision from the tool list
even though both handlers already route to the native fast path and work when
called. The image gate now accepts the native fast path OR an aux client; the
aux-only probe becomes check_video_requirements and stays on video_analyze,
whose handler has no native path.
Fixes#47149.
Twenty-three tests imported helpers from sibling test modules by dotted
path (tests.run_agent.test_run_agent, tests.cli.test_cli_init,
tests.gateway.test_42039_duplicate_user_message); those follow the moves
and renames. The run_agent autouse fixture that zeroes jittered_backoff
now covers all of tests/agent, where one pre-existing test asserts the
real backoff is positive — it opts out via a real_retry_backoff marker,
the same pattern real_concurrent_gate and real_agent_prewarm use.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.
Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.
Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
-> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After: -> _reasoning_config={"enabled": False}.
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
turn prologue now raises instead of silently falling back to per-request wall time, which
would re-create the mid-turn drift the fix removes. Only production caller
(assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
untrustworthy-stamp contract is its own test.
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):
- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
grep hit) was classified as a killed run — the read-only block vanished from the
model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
send/replay parity still held because both consumers made the same mistake. Require
the shape every executor actually produces: the marker is the LAST line of `output`,
inside a JSON envelope with a non-zero exit code (or a bare text result). Real
interruptions from tools/environments (130), managed_modal (130) and
code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
stamp (clock skew, corruption) had negative age and kept the confirmation plus its
api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.
Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:
- Prefix-only on the send path. build_api_messages now canonicalizes only
messages[:current_turn_user_idx]; rows the current turn appended (its own tool
calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
the previous iteration had already sent whenever a tool result matched the
interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
to remove, and it also made the dangling-tail transform order-dependent on when
the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
"exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
`grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
result to an orphan notice (or dropped a read-only block). That heuristic was
tolerable at resume time only; it now runs per request. Match the executors'
bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
live on the send path while replay expired it. _reset_per_turn_agent_state stamps
time.time() once at admission; the three other writes (bind identity, stage
message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
`"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
treat an unknowable age as expired; missing stamps (legacy rows) are still left
alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
_reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
the durable list untouched, live tail preserved; admission-clock freeze across
iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
(send path unpatched, whole-list canonicalization, per-request clock, fail-open,
loose heuristic).
Also widen the two real-subprocess list_models tests from a 2 s to a 30 s response deadline:
spawning a Python interpreter under 16 parallel test workers occasionally exceeded 2 s and the
file flaked (TimeoutError in _request). The deadline only bounds a failure; the happy path
returns as soon as the fake server answers.
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.
Review follow-up on #109610.