`_chat_summary_attempt` now builds through `_build_api_kwargs`, leaving
`_iteration_summary_chat_kwargs` (56 lines mirroring the transport by hand)
and its only consumer `AIAgent._resolve_lmstudio_summary_reasoning_effort`
without a caller. The transport already owns every quirk they re-derived
(fixed temperature, LM Studio `reasoning_effort`, portal tags, provider
preferences, pareto router plugin), so there is nothing to keep in sync.
Three inline context-manager classes shared the same __enter__/__exit__
boilerplate and differed only in the events yielded before the raise.
Also drop the unreachable 'or agent.base_url' fallback: every
anthropic_messages init path sets _anthropic_base_url, and the two
sibling call sites read it bare.
The SDK's malformed-frame ValueError substrings were hand-copied into the
retry classifier and the error summary; a third SDK message would need
both remembered. The mid-tool retry branch no longer clears
partial_tool_names itself — _start_stream_attempt resets it for every
attempt. buffer_anthropic_tool_input's docstring states the trade: the
knob stays on the turn's kwargs and the changed tools block costs one
prompt-cache miss.
A tool_use block name recorded by a stream attempt that died before any
visible text survived into the next attempt: only the deltas_were_sent
mid-tool branch cleared result["partial_tool_names"]. When the retry then
streamed plain text and dropped, the partial stub blamed the stale tool
("Stream stalled mid tool-call (old_tool)") and the stale name could make
a later attempt look mid-tool-call when deciding whether the drop is
retryable. Reset it in _start_stream_attempt alongside
provider_tool_in_flight, which already has attempt-local semantics.
Regression: two-attempt stream (tool_use start + parse error, then text +
drop) — stub content and emitted deltas carry no stale tool name.
Follow-up to the cherry-picked fix: keep the classifier widening
("expected value at line" is now a transient stream parse error on the
main turn) but replace the messages.create() fallback with a retry on the
same stream wire.
Why not create(): the fallback ran outside Relay (lost request rewrites),
outside _handle_stream_error (could replace text already shown to the
user with a different generation), ticked no liveness events for the
whole buffered payload, and a bare identical retry still re-emits the same
malformed JSON.
Why not drop the fine-grained-tool-streaming beta (#108583/#109056): live
probe on claude-sonnet-4-5, ~500-line tool call - beta on: max inter-event
gap 1.6 s; beta off: 139 s zero-event gap while Anthropic buffers the
args, which the 180/240 s stale-stream detector kills on larger payloads
(the regression 80a899a8e2 fixed).
Instead, on a parse error the retry sets `eager_input_streaming: false`
on every tool for that request only (the SDK/API per-tool field overrides
the legacy beta header), so Anthropic returns buffered, server-validated
args while the happy path keeps fine-grained streaming. A tool_use that
started streaming is registered in partial_tool_names so the mid-tool
transient retry fires the same way it does on the chat_completions wire.
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:
- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
key falls back to the conversation root); the fork now publishes exactly the same instead
of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
`_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
parent's PHYSICAL id, so after a compression rotation the review's Portal
`conversation=` tag fragmented usage attribution across one logical conversation.
Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
build_cache_parity_fork gives the same-model fork the parent's session_id,
cached system prompt, tools[] and session_start — but with
_persist_disabled=True and _session_db=None, BOTH cache-identity resolvers
diverged from the parent on their own: declared_conversation_scope failed
closed on _persist_disabled, and the lineage walk skipped on the missing
DB. The fork's affinity header (set_affinity_scope) and body
prompt_cache_key (cache_scope_id on the OpenAI-wire transports) therefore
keyed a different bucket than the gateway parent, costing one cold
~full-context request per review. Not gateway-only: any parent whose
lineage root != current physical id diverges too (teknium1's triage table).
Fix, per the triage's suggested direction: on the not-routed branch only,
the fork stamps _inherited_cache_scope = resolve_prompt_cache_scope_safe
(parent) — the parent's ALREADY-RESOLVED scope, no DB access from the fork,
persistence fully detached. Both declared_conversation_scope and
resolve_prompt_cache_scope return the inherited scope first when set, so
the header path and the body path are fixed together (fixing only one
leaves the other divergent — Vivamisu's header/body split observation).
Routed (different-model) forks, /branch children, delegate/tool children
and fresh sessions set nothing; the fail-closed default stands untouched.
/btw shares build_cache_parity_fork and gets the repair for free.
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).
Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.
Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
The review fork is a full AIAgent, not an auxiliary_client call, and its
routed branch deliberately skips the parent's reasoning_config (the parent's
effort vocabulary may be invalid for the routed provider). It also never read
the per-task key, so an explicit `auxiliary.background_review.reasoning_effort`
was silently ignored and the routed fork ran at the provider default (#94825).
Routed forks now parse the task key through the shared parse_reasoning_effort
(same levels and `none` alias as every other aux task); unset keeps the
provider default, an unknown level warns and falls through. The same-model
path is untouched: it still inherits the parent's reasoning_config verbatim
for prompt-cache parity.
Salvaged from #94832 (liuhao1024), re-applied on the decomposed
_fork_init_kwargs seam.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
auxiliary.background_review.reasoning_effort was a silent no-op on the
same-model path: the fork inherits the parent's reasoning_config verbatim to
keep prompt-cache parity (#30532), and nothing told the user. Emit a one-time
user-visible warning (parent-scoped, so a nudge-per-turn session warns once,
not per fork) when the key is actually set, document the no-op in the config
reference, and leave the fork-birth request bytes unchanged.
The Python-repr pass ran after the MCP probe header scrub and collapsed
'Authorization': 'Digest ***' to '***', erasing the scheme word that scrub
deliberately keeps (tests/hermes_cli/test_mcp_probe_redaction.py went red
on rebase). Skip values that already carry a mask marker.
Port from OpenHands/software-agent-sdk#4508: their dict-entry secret
redaction was uppercase-only and leaked mixed-case keys (UserPassword,
sessionToken). Apply the same case-insensitive treatment to the Python
mapping-repr pass: a casefolded credential suffix (apikey/token/secret/
password/passwd/credential) now qualifies a key, while metadata names
(TOKEN_COUNT, password_policy, tokenizer) stay untouched.
The streaming accumulator dropped delta.reasoning_details entirely — only
non-streaming responses preserved the OpenRouter unified reasoning replay
data (signatures, encrypted reasoning blocks). Providers that require the
reasoning_details sequence passed back on the next turn (OpenRouter
reasoning models, Anthropic signed thinking via shims) lost continuity on
every streamed turn, which is nearly all turns.
- Accumulate delta.reasoning_details (attr or model_extra) during
streaming and attach the merged list to the final mock message, where
_build_assistant_message's existing passthrough persists it.
- _append_streamed_reasoning_detail merges consecutive reasoning.text /
reasoning.summary fragments into one logical entry (OpenRouter streams
them as word-level deltas; unmerged they bloat the replayed signature
payload) while encrypted/opaque entries stay discrete. Later fragments
backfill signature/id/format/index fields the first fragment omitted.
Ported from earendil-works/pi#8605 (commit c5ad7c1b0), credit
@cristinaponcela for the merge-fragments pattern.
Tests: tests/run_agent/test_streamed_reasoning_details.py (8 tests;
sabotage-verified: disabling accumulation fails the E2E test, disabling
merging fails 4 unit tests). Neighboring suites green (test_streaming.py
40 passed).
Small title models parroting a prompt example back verbatim produced
sessions named "Fix login button on mobile" with no relation to the
conversation. The example lines in _TITLE_PROMPT_TEMPLATE now render
from _PROMPT_GOOD_EXAMPLES so the guard set and prompt cannot drift,
and generate_title rejects exact (case-insensitive, wrapper-stripped)
echoes so the instant derived title survives instead. 'Friendly
greeting' stays allowed — it is prescribed output for bare greetings.
The ChatGPT Codex models endpoint interprets client_version as a Codex
CLI compatibility version and filters out any model whose
minimal_client_version is newer than the value sent. Hermes hardcoded
client_version=1.0.0 at both catalog request sites, so model visibility
was accidentally coupled to a version scheme Hermes doesn't follow —
future models gated behind a higher minimal version would silently
vanish from the account catalog.
The backend accepts the exact sentinel 0.0.0 as an ungated request
returning the complete account catalog (verified live: 0.0.0 and
current versions return identical model sets today, while omitting the
parameter is HTTP 400 and out-of-sequence values like 0.0.1 return no
models). Both request sites (hermes_cli/codex_models.py and the
context-length probe in agent/model_metadata.py) now share one
CODEX_UNGATED_CLIENT_VERSION constant.
Clean-room port of the observed behavior in zed-industries/zed#62729;
no GPL code translated.
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
Some OpenAI-compatible gateways fronting Gemini backends emit the native
uppercase finish reasons (STOP, MAX_TOKENS) instead of the lowercase
OpenAI contract values. Every downstream comparison in Hermes uses
lowercase literals, so an uppercase reason silently fell through: a
clean STOP completion missed the stop handling and a MAX_TOKENS
truncation never entered the length-recovery path.
Adds normalize_finish_reason() as the single owner in
agent/message_sanitization.py (case fold + alias map: max_tokens->length,
end->stop, function_call->tool_calls) and wires it at both wire-intake
choke points: ChatCompletionsTransport.normalize_response and the
streaming chunk-capture loop in chat_completion_helpers. Non-string and
empty values pass through unchanged so existing 'or "stop"' defaults
and the Poolside int-reason path keep their behavior.
Adding grep/awk/sed to _FILE_READ_COMMANDS made the PATTERN operand
participate in the secret-file predicate, so `grep .bashrc app.py` or
`grep -n .env src/settings.py` — reads of SOURCE files — ran the
ENV/YAML assignment pass and masked opaque values that main leaves
alone. Skip the first non-flag positional for the pattern-first
readers, as #109369 originally did, so only real file operands gate.
`cat $HOME/.hermes/config.yaml` was ungated because the `$` bail-out
fired before the `.hermes` segment was inspected; strip `$HOME/` and
`${HOME}/` like the HERMES_HOME prefixes.
Review follow-up on #110228.
Fold KoNit-K's `_command_reads_secret_bearing_file` and the pre-existing
`_command_reads_env_file` into a single `_command_reads_secret_file` so the
`code_file` gate in `redact_terminal_output` has one owner: `.env`-style
basenames and shell rc/profile files anywhere, `config.yaml` only under a
`.hermes` directory or `$HERMES_HOME` (arbitrary project YAML stays on the
code_file path). `grep`/`awk`/`sed` join the reader set instead of a second
table with a positional-argument special case: on a file read, any non-flag
operand that names a secret-bearing file is enough — the pattern/program
operand never matches a basename, so the extra rule bought nothing.
Tests: the negative parametrization now uses an opaque credential-shaped value
(the placeholder it used before would never have been masked on either path,
so the "stays unredacted" half proved nothing) and asserts the same value IS
masked under `cat .env` in the same test.
Keep the narrow basename allowlist, but do not treat $HERMES_HOME as an
unresolved path, and split pipelines only on unquoted |;&.
Co-authored-by: Cursor <cursoragent@cursor.com>
The text-only drop guard in _finish_chat_stream required content_parts,
so a stream that died while still emitting delta.reasoning (no
finish_reason, no usage) fell through to the synthesized "stop". With
the reasoning-only clean-stop promotion in finish_text_response that
stamped "stop" turned the truncated thought into the final answer,
where main entered the continuation ladder. Extend the guard with
reasoning_parts so the drop yields the partial-stream stub and the
ladder still runs; a real clean stop carries finish_reason="stop" and
is unaffected.
Review follow-up on #110227.
Follow-up to KoNit-K's commit: rebuild the promotion on the existing
`agent._extract_reasoning` helper (the same reader the ladder terminal and
`build_assistant_message` use) and write the promoted text back onto
`assistant_message.content` so the persisted assistant row carries the answer
as ordinary content. Without that the transcript tail was an assistant row with
empty content and only `reasoning`, which `drop_thinking_only_and_merge_users`
strips from the next request — the model would see its own answer vanish on a
"continue" turn.
Tests: trim to the two invariants (clean stop → one API call, persisted as
content; `finish_reason == "length"` → never promoted, continuation still
owns it) and keep the truly-empty terminal case. The prefill wire-payload
regression test now drives a non-clean-stop reasoning-only reply, which is the
only shape that still reaches the prefill rung.
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.
The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.
Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.
Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.
The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.
Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.
Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
-> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After: -> _reasoning_config={"enabled": False}.
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
turn prologue now raises instead of silently falling back to per-request wall time, which
would re-create the mid-turn drift the fix removes. Only production caller
(assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
untrustworthy-stamp contract is its own test.
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):
- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
grep hit) was classified as a killed run — the read-only block vanished from the
model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
send/replay parity still held because both consumers made the same mistake. Require
the shape every executor actually produces: the marker is the LAST line of `output`,
inside a JSON envelope with a non-zero exit code (or a bare text result). Real
interruptions from tools/environments (130), managed_modal (130) and
code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
stamp (clock skew, corruption) had negative age and kept the confirmation plus its
api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.
Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:
- Prefix-only on the send path. build_api_messages now canonicalizes only
messages[:current_turn_user_idx]; rows the current turn appended (its own tool
calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
the previous iteration had already sent whenever a tool result matched the
interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
to remove, and it also made the dangling-tail transform order-dependent on when
the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
"exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
`grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
result to an orphan notice (or dropped a read-only block). That heuristic was
tolerable at resume time only; it now runs per request. Match the executors'
bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
live on the send path while replay expired it. _reset_per_turn_agent_state stamps
time.time() once at admission; the three other writes (bind identity, stage
message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
`"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
treat an unknowable age as expired; missing stamps (legacy rows) are still left
alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
_reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
the durable list untouched, live tail preserved; admission-clock freeze across
iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
(send path unpatched, whole-list canonicalization, per-request clock, fail-open,
loose heuristic).
Also widen the two real-subprocess list_models tests from a 2 s to a 30 s response deadline:
spawning a Python interpreter under 16 parallel test workers occasionally exceeded 2 s and the
file flaked (TimeoutError in _request). The deadline only bounds a failure; the happy path
returns as soon as the fake server answers.
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.
Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".
gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.
Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.
Review follow-up on #109610.
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.
acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.
`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
Four chat_completions profiles (kimi-coding, deepseek, opencode-go's Kimi K2 and DeepSeek branches, actual) each hand-rolled the same extra_body.thinking / top-level reasoning_effort translation, and the copies had already drifted in small ways (kimi's `.get("enabled", True)`, deepseek's separate effort parsing). agent.reasoning_effort.thinking_toggle_extras is now the single implementation: the Moonshot default emits effort XOR toggle (both is an HTTP 400), and always_emit_toggle=True covers DeepSeek's contract where the toggle must ride on every request to dodge the reasoning_content echo trap. actual keeps its two contract-specific lines (reasoning_config None -> nothing; effort "none" -> disabled toggle plus reasoning_effort="none", which the relay accepts as a real level) and delegates the rest. ox_alpha_reasoning_extras moves alongside so opencode-free imports it like any other helper instead of reaching into the zen plugin's module through sys.modules and swallowing every exception into ({}, {}) - a failure there previously silently dropped the user's effort setting. No wire behavior changes; tests/plugins/model_providers/test_thinking_toggle_parity.py pins the XOR invariant across the matrix and zen/free parity.
Three shims caught UnscopedSecretError and degraded to "" (mem0._scoped_env)
or to os.environ (langfuse._secret, azure_identity_adapter._scoped_env). Under
multiplex os.environ holds the DEFAULT profile's .env, so the langfuse/azure
fallback could ship another profile's keys, and the mem0 fallback silently
routed a mis-spawned turn's memories into the default profile's account. The
exception exists to surface exactly that spawn-site bug (agent/AGENTS.md:
never add environ fallthrough, never swallow it). All three now call
agent.secret_scope.get_secret directly: with a scope installed a miss returns
the default; single-profile deployments (multiplex off) still read the process
env inside get_secret; a scope-less multiplex caller raises.
Behavior change: a mis-spawned child under gateway.multiplex_profiles now
fails loud with UnscopedSecretError instead of running silently unauthenticated
/ on the default profile's identity. The #99121 contract (OSS mode needs no
MEM0_API_KEY in scope) is unchanged and its test now installs an empty profile
scope, which is the situation the issue described; a genuinely scope-less
caller is asserted to raise in a new test.
Tests: tests/plugins/test_scoped_secret_readers_fail_closed.py (scope wins over
environ; scope-less multiplex raises) for langfuse + azure, sabotage red when
the langfuse fallthrough is restored; tests/plugins/memory/test_mem0_v3.py::
test_load_config_fails_closed_without_scope_even_for_identity_settings,
sabotage red with a swallowing wrapper reinstated.
Six of eight memory providers spawned plain threading.Thread for prefetch/sync/
writer work. A plain thread starts with an EMPTY contextvars.Context, so under
multiplex profiles the worker resolved the DEFAULT profile's HERMES_HOME (and
fails closed on scoped secrets). honcho and hindsight had each noticed and
written their own copy_context() wrapper; core had a third in memory_manager.
One canonical pair now lives on the ABC module every provider already imports:
agent/memory_provider.py::ctx_bound / spawn_context_thread. memory_manager,
honcho, hindsight, mem0, retaindb, byterover, supermemory and openviking all use
it; the honcho and hindsight wrappers and memory_manager._ctx_bound are deleted.
Five "json.loads(path.read_text()) or {}" readers (mem0._read_mem0_json,
honcho client/oauth/cli _read_config, hindsight save_config/_load_config) fold
into utils.read_json_or_empty, the read half of every read-merge-atomic_json_write
sidecar store.
holographic.save_config was the only config.yaml writer in the tree that
bypassed hermes_cli.config.save_config: raw open("w") + yaml.dump with no config
lock, no managed-mode refusal, no atomic replace, and a swallowed exception. It
now calls save_config(..., merge_existing=True). Behavior change: a managed
install refuses the write (previously silently rewrote config.yaml); other
sections are deep-merged instead of round-tripped through a raw dump.
openviking._hermes_home_path guarded an impossible ImportError of
hermes_constants (the module already imports agent.*) with a ~/.hermes fallback
that is wrong on Windows and under profile overrides; it is replaced by
get_hermes_home() directly.
Tests: tests/plugins/memory/test_provider_threads_inherit_profile.py drives each
provider's real spawn path with a fake backend and asserts the thread sees the
spawner's HERMES_HOME override (sabotage: retaindb back on threading.Thread ->
red). tests/plugins/memory/test_holographic_save_config.py pins merge-with-
existing-sections and managed-mode refusal (sabotage: raw yaml.dump -> red).
Unifying the credential pool's `_RETRY_DELAY_PATTERNS` into `RETRY_DELAY_PATTERNS`
flipped the pool's precedence: "retry after 30s; resets in 4hr" cooled the
credential for 14400 s where the pool used to take 30. A body carrying both
describes a short throttle inside a long quota window; the explicit retry-after
is the wait the provider actually asks for, so it is tried before "resets in".
Review follow-up on #109539.