ethernet8023: the seeded E2E (turn.start -> deltas -> turn.end with history_version/
message_count) was cut to hello+bogus-frame while ComputeHost._run_real_turn stays live.
New suite drives server._sessions + _run_prompt_submit through the host: turn.started,
message.delta rpc frames, turn.end identity (history_version=1, message_count=2), sid-
required/session-busy turn.error, stale queued-generation -> interrupted, live interrupt
frame ack + interrupted turn.end, and the explicit-stop compat matrix ported off HostSession.
All public on BASE 63279301bc, dropped by the simplify refactor (their tests were deleted or
rewritten to the replacement API). Restore each with BASE signature/body as a thin wrapper over the
surviving implementation, and restore the tests at the original call sites: test_message_reactions
again asserts the role=user contract (a newer assistant message is never the default target);
test_hermes_state / test_watchdog_review_76354 go back to get_session_activity(); toolsets, acp auth,
edit_approval, billing-scope and curated-models tests restored/extended.
ethernet8023: IAM streaming-denial -> converse fallback and stale-ConnectionClosedError
client eviction lost all coverage. Both public helpers restored byte-identical to BASE
(with THROTTLE/OVERLOAD/CONTEXT_OVERFLOW patterns + is_context_overflow_error), the 4
tests re-added verbatim, and TestAgentBedrockStreamRecovery covers the path the agent
actually uses (chat_completion_helpers._bedrock_converse_call / _BedrockStream._fall_back_to_converse).
ethernet8023: TestValidateImageUrl (incl. localhost block) was deleted with the sync
validator. Both names restored with BASE semantics; the async validator (the live
download path) now shares _image_url_shape_ok and gets its own localhost/malformed test.
ethernet8023: the no-leak test was deleted alongside the function it pinned; the
display-only credential status is live in print_credential_summary. credential_summary
restored byte-identical to BASE so the leak contract is CI-pinned again.
ethernet8023: both documented VAD false-trip regression tests were deleted with the
function they covered. listen_for_speech is public API on main (plugins may import it),
so the body is restored byte-identical to BASE; full_duplex_listen/_BargeDetector is
A/B-verified identical to BASE over 400 fuzzed playback/speech scenarios.
BASE (63279301bc) exported agent.agent_init._moa_reference_output_allowed and
_relay_moa_reference_event (salvaged in 3dfe712384 from #67334, plugin-importable,
zero in-tree callers) and covered them with
tests/agent/test_moa_quiet_reference_output.py. The simplification PR deleted both
the helpers and the test. Restore them byte-identical to BASE.
Note: the live build_moa_facade relay in agent/moa_loop.py is unchanged. A/B
harness (/tmp/rf/rev/moa_quiet_ab.py) shows BASE's facade relay already fired
moa.reference under platform=cli/tool_progress_mode=off — the guard only ever
lived in the orphan helper. -Q is protected on both trees by cli.py nulling
agent.tool_progress_callback (_configure_quiet_agent / BASE cli.py:22308).
BASE exposed these public names; the simplify refactor dropped them (and deleted/removed their
tests) although plugins/connectors import them. Restore each with BASE's signature and body
(download() again routes its auth decision through is_relay_media_url), and restore the covering
tests ported to the new layout, plus DB-only/task-cwd coverage for remove_session/cleanup, the
zero->4096 chunking normalization for from_platform_entry, and len() on empty/populated registries.
A/B vs 63279301bc: /tmp/rf/rev/ab_publicapi.py identical output on both trees.
main's 'and len(current_lines) > 1' fence-close guard is dead (the opening
fence is always already in current_lines), so dropping it is byte-identical
for every input (A/B: 11 empty/nested fence shapes). Add a parity test that
asserts main's atoms for empty ``` blocks so the boundary stays pinned.
main guarded scroll's x/y independently (coordinate=[null,100] -> y=100),
while click treats a coordinate without x as no point. The shared _xy()
applied click semantics to scroll; split out _scroll_xy() with main's
per-axis guard and pin both behaviors in a test.
`_write_schema_cache` in the extracted owner tools/mcp_tool_registration.py
read `t.inputSchema` with a bare camelCase getattr. mcp 2.0 renamed the
model field to `input_schema` and kept `inputSchema` only as a serialization
alias, which pydantic does not apply to attribute access, so on SDK 2.x the
cache persisted `"inputSchema": {}` for every tool and a `lazy: true` server
registered from that cache was advertised to the model with all parameters
stripped. Pre-existing on BASE (tools/mcp_tool.py); fixed with the canonical
`mcp_field(t, "input_schema", "inputSchema")` helper, as both PRs do.
Adds the real-SDK regression test from #102129 (genuine `mcp.types.Tool`
driven through `_register_server_tools` -> cache write ->
`_register_from_cache_sync`), ported to the new module layout and isolated
from the module-global registration state.
Co-authored-by: mzkarami <mehrzad.karami@gmail.com>
Co-authored-by: lijinxiao1982 <120761624+lijinxiao1982@users.noreply.github.com>
The hermes_state split added 15 root-level sibling modules but
[tool.setuptools] py-modules still listed only the old set, so the built
wheel/sdist (and the uv2nix sealed venv) failed at 'import hermes_state'
with ModuleNotFoundError. hermes_state_holders and hermes_state_registry
were already missing on main (registry is imported by gateway/run.py and
hermes_state.py).
Add an invariant test that parses pyproject and walks every import in the
packaged root modules and packages, requiring any import that resolves to
a root-level *.py to be declared - so the class of bug is caught by CI
rather than by a broken install.
The retry must land on the same provider cache key as every prior request
in the session. Discard only the one-shot continuation disable and send
agent.reasoning_config verbatim; a config that is itself a disable is
omitted (that session never sent anything else, so nothing warm is lost).
Live: user effort=high, ephemeral disable → 400 → retry carries
{enabled: true, effort: high}.
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.
- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
(configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
mandatory (parity with the Nous profile).
Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
The memory guidance led with 'Save proactively' and the memory tool schema
ranked 'user preferences & corrections' as top priority, while the skills
nudge was a conditional 'offer to save'. In practice that asymmetry made
the agent end sessions writing memory entries (fighting a 2,200-char
budget) and skip updating the skill it had just used, even though the
procedure was the reusable artifact. Both surfaces now state the same
rule with skills first: what you learn doing a task, including the user's
preferences and corrections for that kind of work, goes in the task's
skill; memory is only for facts that apply to every session.
GLM-style models serialize tool calls as XML in the text channel; when the
stream drops mid-serialization with finish_reason=stop, the orphan
<arg_key>/<arg_value> fragment (or a bare unclosed <tool_call> opener)
matched neither the complete-block stripper nor the partial-stream guard
and was stored and displayed as ordinary assistant content.
strip_think_blocks (storage boundary) and the CLI display copy now strip
an unterminated block-boundary tool-call opener, or any line carrying
stray argument markup, to end of text. The response then reads as empty
and flows through the existing empty-retry path. Complete blocks and
inline prose mentions are unchanged.
Drop the loop-side '(empty)' rewrite (the turn-completion explainer already
owns that at delivery, and gateway/desktop match on the sentinel) and the
extra token-count persistence. Keeps: usage-absent empty streaks arm the
deterministic fast-fail after two attempts with no content or reasoning,
and every completed API call logs even when the provider omits usage
(#101898).
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog). A
profiled session with ~130 children was carrying ~1000 threads. All
of these timers now run on a single process-wide daemon thread.
- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
one Condition-driven daemon thread. schedule(fn, interval) -> handle;
handle.cancel(wait=) blocks for an in-flight run like the old join.
A callback returning False stops itself; a raising callback is logged
at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
_lease_refresh_interval; lease-lost / refresh-error interrupt paths
and the stop-event fencing are unchanged; the join(timeout=1.0) is now
cancel(wait=1.0) so the interrupt clear still runs after any in-flight
tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
schedule(); the poll body is _tick(), same sampling state machine.
Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132. At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:
1. bind_subagent_parent() stored the agent strongly in the
`hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
for its own turn, and every asyncio Handle/Future scheduled during
that turn (LSP reader loops, kernel pipe transports) snapshots the
Context — 56 live Contexts held 14 finished children after the bench.
The ContextVar now holds a weakref (non-weakrefable doubles fall back
to a closure); get_active_subagent_parent() dereferences it.
2. AIAgent.close() cleared _session_messages but not the
_db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
every successful DB flush) nor _streamed_assistant_text_parts, so the
agent — kept alive by (1) — retained every message dict. close() now
drops both.
The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.
Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.
What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.
Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).
Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
`limits=` never reached them, so mounts ran on httpx defaults with a 5 s
keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).
Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
Compaction moved run_bang_command's Popen within the env-guard scanner's
proximity window of _bang_env's os.environ.copy(). Route through the single
factory (build_subprocess_env() == _sanitize_subprocess_env(os.environ.copy()))
and allowlist the file for the import-failure fallback copy with justification.
cmd_sessions used 'with db:' which breaks test doubles lacking the context
manager protocol (13 reds in test_sessions_pin/delete/export). Restore the
explicit try/finally db.close() (same semantics for real SessionDB). Add
get_session/count_prune_matches to the FakeDB doubles in test_sessions_delete
instead of re-adding getattr guards to production code.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).
Missing alias identified by @Steve-prog001 in #101905.
Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.
Fixes#101668
Refs #47742
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.
Both gates now go through one _profile_rejects_tool_media() veto.
Refs #89981
(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)
With the picker served from resident caches only, a cold Nous row renders
every model locked (free_tier_pending) until the background prewarm lands.
Surface why on the row's existing warning slot so the user isn't left with
an unexplained greyed-out list; never override an auth warning.
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.
Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
The Bot Mode mention middleware built message_agent targets from
botHandle(), which prefers a roster row's source-qualified UI alias
("default-vera"). Neither resolver accepts that form — the relay matches
canonical handle/profile (± @connection-id) and the local path a bare
profile name or "hermes" — so remote handoffs died with "No teammate
named" before enqueue.
Annotate the canonical form instead: profile@connection-id for remote
rows, canonical bare handle (default→hermes) for local ones. Pin the
profile@connection form on the relay side too, so the emitted target
stays inside the documented resolver contract (#97678).