Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:
- _is_actionable_user_turn (tail anchor) only checked role/content, so a
notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
operational notices leaked into the compact focus hint.
Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.
Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.
Fixes#92703
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:
Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
(_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
(honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod
Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots
Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.
Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
The review fork's entire job is to emit `memory` / `skill_manage` tool calls,
and by default it inherits the parent's live runtime. A provider that IS an
autonomous agent reaches Hermes through a client shim; if that shim cannot
carry Hermes tool calls back, the fork is a guaranteed no-op that still pays
for a full agent spawn — a whole CLI process, sometimes a JVM — on every
review cadence.
A client declares `SUPPORTS_HERMES_TOOL_CALLS = False` and the fork is
skipped with a warning naming the `auxiliary.background_review.{provider,model}`
override that routes the review to a normal model instead. Anything that says
nothing is assumed capable, so ordinary providers are untouched.
The check runs before the thread-scoped silence so the warning is not
swallowed, and only resolves the review runtime once the cheap capability
test has already failed, so the normal path does not resolve it twice.
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.
Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:
- the self-improvement loop, which distils memories and skills by replaying
`messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
Hermes tool iterations, of which there are none.
So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.
The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
ACP has no OpenAI `tools`/`tool_calls` channel: a prompt is text and a
response is text plus the agent's own tool notifications. Hermes' agentic
surface — memory, todo, skill_manage — is dispatched from OpenAI-shaped
tool_calls, so on an ACP provider it only works if the schemas travel into
the prompt as text and the calls are parsed back out of the response text.
copilot-acp already carried that bridge as private module-level helpers.
Lift it verbatim into `agent/acp_openai_bridge.py` so every ACP client
shares one implementation of the wire contract instead of re-deriving it —
`agent/claude_code_acp_client.py` (#81375) is currently a third copy of the
same four functions, and each copy is a place the `<tool_call>` contract can
drift.
copilot-acp is migrated onto it as the in-tree consumer and loses 176 lines
of duplication; its prompt shape is unchanged, which the new tests pin.
Two things the shared version adds over the copy:
- `render_tool_bridge_sections(..., allowlist=)`. A CLI with no tools of its
own forwards Hermes' whole toolset (copilot, unchanged: no allowlist). A
CLI that *is* an autonomous agent must forward only Hermes' agent-level
tools — re-offering the overlapping read/edit/execute ones makes Hermes
re-run work the agent already finished.
- `StreamChunks`, a list subclass that keeps response-level attributes.
Hermes reads provider extras off the object returned by
`chat.completions.create`; the old plain-list return silently dropped them
whenever a caller asked for `stream=True`.
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.
New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
async completion test)
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
describes the state at timeout; a recycling cleanup never poisons the
healed replacement)
- the async flavor offloads the mark off the event loop
(asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
The reuse reviewer found a third copy of the fc_->call_ synthesis in
agent/transports/codex.py _pair_ids (item_id[3:] spelling, which is why
the len('fc_') grep missed it). All three sites now share
_canonical_call_id_from_fc(), keeping the pairing invariant in one place.
/simplify-code reuse+quality reviewers both flagged the byte-identical
fc_->call_ synthesis blocks in the assistant and tool-result branches as
a correctness coupling — the two sites MUST stay in lockstep or pairing
breaks. Extract _canonical_call_id_from_fc() and route both through it.
Mutation check: pairing regression test fails when the tool-branch call
is stubbed out, green after restore.
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).
Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'
The 400 replays forever until the user manually starts a new session.
Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError). Applied
at both replay sites: the chat-message converter and the preflight
choke-point. Live tool-definition names are left untouched — they must
match the dispatch registry exactly. Pairing is by call_id, so
renaming a replayed function_call is safe.
call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.
Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).
Fixes#31666
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):
1. The parallel batch planner now peels the tool_call bridge wrapper and
decides admission on the underlying tool — supports_parallel_tool_calls
works again when deferral is active. Unparseable wrappers stay
sequential barriers; bridged calls get exactly the admission the same
call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
service-name queries reach tools whose own name omits the service; the
dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).
Salvaged from #92693 by @alt-glitch with authorship preserved.
* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
Independent review caught a compaction authority the gate missed:
post-turn micro-compaction (turn_finalizer -> _micro_compact) absorbs the
oldest exchanges into a rolling summary with no pre-compress checkpoint
hook in its path, and both compression.checkpoint_required and
compression.micro_compact could be enabled together — assistant evidence
could vanish into a summary the checkpoint filter later excludes, without
ever reaching the durable provider.
- agent_init: checkpoint_required forces micro-compaction off (warned),
mirroring the native-compaction suppression
- turn_finalizer: defense-in-depth guard at the call site (attribute is
plain mutable state a future path could flip on a live agent)
- behavioral regression test with a sabotage control (gate off proves the
harness reaches the call site; gate armed proves zero calls)
- docs + config example mention the suppression; stale v1 test header fixed
Per review: existing providers should not be retroactively re-versioned or
handed a changed payload. Version 1 is now the implicit historical
on_pre_compress() contract (best-effort, raw message list) that every
pre-existing provider is already on; the fail-closed checkpoint contract
becomes version 2. MemoryManager routes the raw transcript to v1 providers
unchanged and hands the host-normalized evidence list only to v2+ checkpoint
providers, so the plugin surface contract for shipped providers is
byte-identical with the gate off.
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):
- codex app-server: in "native"/"off" auto-compaction mode (native is the
default) Hermes preflight is skipped and the codex agent compacts its
own thread inside run_turn() — the compress_context() rejection was
unreachable. init_agent now refuses checkpoint_required together with
api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
testable guard), and run_codex_app_server_turn() fails closed as
defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
native_compaction_context_management() now returns None while the gate
is armed, so context_management never goes on the wire and the
checkpoint-aware Hermes compressor stays authoritative. The suppression
is logged once per process, not silently applied.
Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.
Refs #93986
Addresses the review on #93996:
- gateway: hygiene and manual /compress load the memory provider only when
compression.checkpoint_required is enabled (skip_memory=not required).
The historical fast path — no provider init, no best-effort hook — is
back for everyone who did not opt in, so default behavior is truly
unchanged.
- conversation_compression: assistant messages carrying both prose and
tool_calls keep their prose in the checkpoint evidence (the tool_calls
payload is stripped, the original message is not mutated); pure
tool-call wrappers without prose are still dropped.
- tests: legacy-database regression proving the _compressed_summary column
is added by the declarative _reconcile_columns() path on a plain reopen
(no version-gated migration needed — append_message works right after),
plus coverage for the prose-preserving filter.
- docs: providers must implement idempotent, content-keyed checkpoint
writes — a fail-closed block means the next attempt re-runs
on_pre_compress over largely the same transcript.
Refs #93986
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
`_background_review_read_before_write_guard` refuses a background-review
`skill_manage` write when the target file was not loaded via `skill_view` in
the same review turn (patch, edit, write_file over an existing file,
remove_file).
`CURATOR_REVIEW_PROMPT` never says so. It lists `skill_view` only under "read
the current landscape", so the reviewer goes straight to the write and every
mutation is refused. The failure is silent from the outside: the curator run
completes, writes nothing, and reads like a pass that simply found nothing to
consolidate. On our deployment that was 32 of 32 attempted writes rejected over
48h before anyone read the logs.
This adds the missing instruction to the toolset block, plus a test that fails
if a future guarded action is added to `skill_manager_tool` without being named
in the prompt — the guard and the prompt have to drift together or not at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The skill_manage guard (added in #55906) refuses any patch/edit of an
existing SKILL.md, or overwrite/removal of an existing support file,
unless the exact target was loaded via skill_view during the review.
Neither _SKILL_REVIEW_PROMPT nor _COMBINED_REVIEW_PROMPT ever mentioned
this, so models routinely issued the write without the pre-read, got
refused, and burned review iterations (#62397).
Both prompts now carry a Read-before-write section scoped to the
guard's actual contract: existing targets only, exact-path pre-read for
support files, transcript quotes don't count, new skills/new support
files exempt, and a bounded one-view-one-retry recovery instead of a
loop. Direction follows #60331 by @kkwills13 with the scope corrections
requested in review (existing-target-only wording, no delete claim,
bounded retry, contract tests for both prompt variants).
Fixes#62397.
Follow-up to the cherry-picked gate-parity fix: the silent 'return 0'
in inject_memory_provider_tools made #81014 undiagnosable — a
configured provider looked half-on with no hint which config key
suppressed its tools. Now an INFO line names the withheld providers
and the gating keys.
The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provider_tools_enabled()` via platform_toolsets or
disabled_toolsets. Result: the agent received instructions to call
`mnemosyne_remember`, `mnemosyne_recall`, etc., that did not exist in
its tool surface.
Centralize the gating into `memory_provider_tools_exposed(agent)`, use
it from both `inject_memory_provider_tools` and the system prompt
assembly path, and add regression tests covering:
* memory toolset enabled -> both tools and prompt block exposed,
* memory in disabled_toolsets -> neither exposed,
* memory not in enabled_toolsets and not built-in -> neither exposed,
* the built-in "memory" tool present as an opt-in -> both exposed,
* parity between `inject_memory_provider_tools` and
`memory_provider_tools_exposed`.
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.
- hermes-agent skill is now essential: cannot be disabled (config reads
strip it, hermes tools writes drop it), cannot be deleted by
skill_manage, is re-seeded past curator suppression, and is seeded
even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
file+terminal+vision+skills: read_file cannot read images and points
at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
tool isn't loaded.
All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.
- tools/browser_tool.py: remove _extract_relevant_content and
_get_extraction_model; oversized snapshots always truncate at line
boundaries, store the full tree to cache/web, and append a read_file
pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
pattern as session_search/PR #27590), cli.py defaults + env bridge,
gateway/run.py bridged keys, hermes config display, hermes model picker,
dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
LLM path is gone and stored files are secret-redacted
Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:
- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
poisoned state never does I/O; the NEXT caller pays once for a health
probe that clears the suspicion or forces a reconnect. A single
teardown-vs-keepalive race or auth-lock corruption can no longer park
a connection permanently — park stays reserved for genuinely
exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
in-flight RPC races a child-watcher task, so a dead subprocess fails
the call immediately with a retryable timeout instead of riding out
the full 300s. Deliberate teardown/reconnect also fails in-flight
calls now instead of leaving them attached to a dying transport.
Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.
Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.
Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
Widen the salvaged gate-site clamp to the bug class. The gate min() from
the contributor PR capped the gate bound at 360s, which would break the
#79719 contract (gate must extend while a legitimate >360s approval
prompt is answerable) and left the sibling overflow sites live: the CLI
prompt thread.join, the gateway poll deadline, and human_wait_ceiling
all consume the same config value.
Clamp once in _get_approval_timeout() via agent.deadline.MAX_SAFE_TIMEOUT_S
(1 year - semantically unbounded, platform-safe). The gate keeps its
approvals.timeout-tracking behavior above 360s; 7 regression tests pin
lock-acquire/thread-join safety and the gate-extension contract.
Bilateral E2E: with approvals.timeout=1e20 in a real config.yaml, main
crashes every consumer with OverflowError; this branch survives all 5
probes.
The _ConcurrentToolAuthorizationGate uses threading.Lock.acquire(timeout=...)
where timeout comes from human_wait_ceiling() (approvals.timeout + 60s).
When approvals.timeout is set very large (e.g. 999999999999 to effectively
disable timeouts), this overflows macOS timespec and raises:
OverflowError: timestamp out of range for platform time_t
The gate only needs to serialize parallel dispatch — it should never wait
longer than the wedged-holder bound (_AUTHORIZATION_GATE_LOCK_TIMEOUT_S =
360s). Clamp the return value with min() so unbounded approvals.timeout
cannot overflow the lock acquire.
Fixes#83220