dingtalk: drop DINGTALK_TYPE_MAPPING/EXT_MAP re-exports. google_chat: card_spec_to_cards_v2 test -> .cards.
matrix: drop module-level MAX_MESSAGE_LENGTH alias (no importers). teams: drop TeamsSummaryWriter re-export
(teams_pipeline/runtime + tests -> summary_writer). wecom: drop WeComStreamExpiredError/STREAM_EXPIRED_ERRCODE/
MAX_INTERMEDIATE_FRAMES re-exports (tests -> .streaming). parallel: drop _get_parallel_client/_get_async_parallel_client
aliases (tests -> _get_sync_client). email: drop stale 'alias' comment (_esecret_int is the only name).
photon: re-remove credential_summary() (shim-only, cb9b7c36f3); its no-leak test now drives print_credential_summary.
BASE's _compress_codex_app_server_session was the single cache-peek site that read
_agent_cache without a lock when _agent_cache_lock was None (tests/gateway/
test_codex_hygiene_compaction.py::_slash_host relies on it). Every other BASE caller returned
None without a lock. Model that as an explicit lockless_fallback=True at that one call site.
Adolanium review §3 (Medium/Low). BASE shape at every steer/redirect/persist/agent-cache
lock site was: lock = getattr(obj, "_x_lock", None); if lock is not None: with lock: <direct
attribute read>; else: <getattr fallback for object.__new__ test stubs>. The refactor rewrote
several of these as 'with getattr(...) or nullcontext(): getattr(slot, None)', which (a) turned a
missing slot under the lock from a loud AttributeError into silent None and (b) in the gateway
peek helper read the cache without any lock when the lock attribute was absent.
Restored BASE semantics at:
- agent/agent_runtime_helpers.py::_requeue_pending_steer
- agent/interrupt_control.py: steer, redirect, clear_interrupt, _has_pending_redirect,
_drain_pending_redirect, _drain_pending_steer (new _ic_slot helper: direct read under lock)
- agent/session_persistence.py::_persist_lock (explicit None check + BASE rationale)
- gateway/slash_commands.py::_cached_agent_for (BASE callers read ONLY under the lock; no lock -> None)
- gateway/run_agent_cache.py::_evict_cached_agent (BASE: self._agent_cache direct under lock)
- agent/client_lifecycle.py::_is_openai_client_closed: BASE body + docstring verbatim (outer
is_closed first; inner _client.is_closed only when _client exists; else False)
- agent/stream_delivery.py::_ensure_stream_writer_state: restore BASE rationale that the lock is
created unconditionally in agent_init (_STREAM_STATE) and the lazy path is stub-only
A/B (/tmp/rf/rev/ab_lock_fallbacks.py, ab_client_closed.py) is byte-identical BASE vs HEAD on
the reviewer's inputs + edge cases. tests/agent/test_lock_fallback_base_semantics.py pins it
(7 of 25 cases fail on the pre-fix tree).
The static py-modules list in pyproject.toml drifted from the source
tree each time the root layout changed. An installed wheel then
raised ModuleNotFoundError on import: hermes_state_common, and every
gateway or CLI start failed. The list missed hermes_state_holders,
mini_swe_runner, and the 15 new hermes_state_* modules from this
branch.
setup.py now derives py_modules from the source tree at build time.
setuptools package discovery sees only directories with an
__init__.py, so root single-file modules need py_modules in every
wheel build. setup() kwargs merge with pyproject.toml, and setup.py
is the only legitimate wheel or sdist builder, so the derived list is
the single source of truth.
The nix build is the only wheel consumer. Its source filter keeps
every root .py file, so the build sandbox derives the same set as the
checkout. Editable installs do not read py_modules: build_editable
never runs bdist_wheel.
Verified: wheel and sdist built with HERMES_NIX_BUILD=1 carry all 38
root modules. The guard still blocks builds without the nix env var.
The sealed uv2nix venv from nix build .#default imports
hermes_state_holders, hermes_state_sessions, mini_swe_runner, and
hermes_startup_watchdog.
(cherry picked from commit e23d467b523d50ad088b89b578017996709589cf)
Invariant test updated to pin the derived list (no static py-modules; every root module packaged code imports ships).
On BASE the #95514 stream-recovery rebound final_response inline, before
the tail-closing / persist-override / micro-compaction / _persist_session
calls inside the guarded persist try-block, so a raise in any of them left
the caller with the streamed text (cleanup_errors reported the failure).
HEAD's _close_transcript_tail returned the recovered value only on normal
completion; an exception in the same block dropped it and the user-visible
answer regressed to the pre-recovery '' (then rewritten by the explainer).
Split the helper into _drop_transcript_scaffolding / _recover_final_from_stream
/ _close_transcript_tail and rebind final_response in the guarded step as soon
as it is computed, restoring BASE ordering. Adds a regression test.
ethernet8023: the seeded E2E (turn.start -> deltas -> turn.end with history_version/
message_count) was cut to hello+bogus-frame while ComputeHost._run_real_turn stays live.
New suite drives server._sessions + _run_prompt_submit through the host: turn.started,
message.delta rpc frames, turn.end identity (history_version=1, message_count=2), sid-
required/session-busy turn.error, stale queued-generation -> interrupted, live interrupt
frame ack + interrupted turn.end, and the explicit-stop compat matrix ported off HostSession.
All public on BASE 63279301bc, dropped by the simplify refactor (their tests were deleted or
rewritten to the replacement API). Restore each with BASE signature/body as a thin wrapper over the
surviving implementation, and restore the tests at the original call sites: test_message_reactions
again asserts the role=user contract (a newer assistant message is never the default target);
test_hermes_state / test_watchdog_review_76354 go back to get_session_activity(); toolsets, acp auth,
edit_approval, billing-scope and curated-models tests restored/extended.
ethernet8023: IAM streaming-denial -> converse fallback and stale-ConnectionClosedError
client eviction lost all coverage. Both public helpers restored byte-identical to BASE
(with THROTTLE/OVERLOAD/CONTEXT_OVERFLOW patterns + is_context_overflow_error), the 4
tests re-added verbatim, and TestAgentBedrockStreamRecovery covers the path the agent
actually uses (chat_completion_helpers._bedrock_converse_call / _BedrockStream._fall_back_to_converse).
ethernet8023: TestValidateImageUrl (incl. localhost block) was deleted with the sync
validator. Both names restored with BASE semantics; the async validator (the live
download path) now shares _image_url_shape_ok and gets its own localhost/malformed test.
ethernet8023: the no-leak test was deleted alongside the function it pinned; the
display-only credential status is live in print_credential_summary. credential_summary
restored byte-identical to BASE so the leak contract is CI-pinned again.
ethernet8023: both documented VAD false-trip regression tests were deleted with the
function they covered. listen_for_speech is public API on main (plugins may import it),
so the body is restored byte-identical to BASE; full_duplex_listen/_BargeDetector is
A/B-verified identical to BASE over 400 fuzzed playback/speech scenarios.
BASE (63279301bc) exported agent.agent_init._moa_reference_output_allowed and
_relay_moa_reference_event (salvaged in 3dfe712384 from #67334, plugin-importable,
zero in-tree callers) and covered them with
tests/agent/test_moa_quiet_reference_output.py. The simplification PR deleted both
the helpers and the test. Restore them byte-identical to BASE.
Note: the live build_moa_facade relay in agent/moa_loop.py is unchanged. A/B
harness (/tmp/rf/rev/moa_quiet_ab.py) shows BASE's facade relay already fired
moa.reference under platform=cli/tool_progress_mode=off — the guard only ever
lived in the orphan helper. -Q is protected on both trees by cli.py nulling
agent.tool_progress_callback (_configure_quiet_agent / BASE cli.py:22308).
BASE exposed these public names; the simplify refactor dropped them (and deleted/removed their
tests) although plugins/connectors import them. Restore each with BASE's signature and body
(download() again routes its auth decision through is_relay_media_url), and restore the covering
tests ported to the new layout, plus DB-only/task-cwd coverage for remove_session/cleanup, the
zero->4096 chunking normalization for from_platform_entry, and len() on empty/populated registries.
A/B vs 63279301bc: /tmp/rf/rev/ab_publicapi.py identical output on both trees.
main's 'and len(current_lines) > 1' fence-close guard is dead (the opening
fence is always already in current_lines), so dropping it is byte-identical
for every input (A/B: 11 empty/nested fence shapes). Add a parity test that
asserts main's atoms for empty ``` blocks so the boundary stays pinned.
main guarded scroll's x/y independently (coordinate=[null,100] -> y=100),
while click treats a coordinate without x as no point. The shared _xy()
applied click semantics to scroll; split out _scroll_xy() with main's
per-axis guard and pin both behaviors in a test.
`_write_schema_cache` in the extracted owner tools/mcp_tool_registration.py
read `t.inputSchema` with a bare camelCase getattr. mcp 2.0 renamed the
model field to `input_schema` and kept `inputSchema` only as a serialization
alias, which pydantic does not apply to attribute access, so on SDK 2.x the
cache persisted `"inputSchema": {}` for every tool and a `lazy: true` server
registered from that cache was advertised to the model with all parameters
stripped. Pre-existing on BASE (tools/mcp_tool.py); fixed with the canonical
`mcp_field(t, "input_schema", "inputSchema")` helper, as both PRs do.
Adds the real-SDK regression test from #102129 (genuine `mcp.types.Tool`
driven through `_register_server_tools` -> cache write ->
`_register_from_cache_sync`), ported to the new module layout and isolated
from the module-global registration state.
Co-authored-by: mzkarami <mehrzad.karami@gmail.com>
Co-authored-by: lijinxiao1982 <120761624+lijinxiao1982@users.noreply.github.com>
The hermes_state split added 15 root-level sibling modules but
[tool.setuptools] py-modules still listed only the old set, so the built
wheel/sdist (and the uv2nix sealed venv) failed at 'import hermes_state'
with ModuleNotFoundError. hermes_state_holders and hermes_state_registry
were already missing on main (registry is imported by gateway/run.py and
hermes_state.py).
Add an invariant test that parses pyproject and walks every import in the
packaged root modules and packages, requiring any import that resolves to
a root-level *.py to be declared - so the class of bug is caught by CI
rather than by a broken install.
The retry must land on the same provider cache key as every prior request
in the session. Discard only the one-shot continuation disable and send
agent.reasoning_config verbatim; a config that is itself a disable is
omitted (that session never sent anything else, so nothing warm is lost).
Live: user effort=high, ephemeral disable → 400 → retry carries
{enabled: true, effort: high}.
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.
- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
(configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
mandatory (parity with the Nous profile).
Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
The memory guidance led with 'Save proactively' and the memory tool schema
ranked 'user preferences & corrections' as top priority, while the skills
nudge was a conditional 'offer to save'. In practice that asymmetry made
the agent end sessions writing memory entries (fighting a 2,200-char
budget) and skip updating the skill it had just used, even though the
procedure was the reusable artifact. Both surfaces now state the same
rule with skills first: what you learn doing a task, including the user's
preferences and corrections for that kind of work, goes in the task's
skill; memory is only for facts that apply to every session.
GLM-style models serialize tool calls as XML in the text channel; when the
stream drops mid-serialization with finish_reason=stop, the orphan
<arg_key>/<arg_value> fragment (or a bare unclosed <tool_call> opener)
matched neither the complete-block stripper nor the partial-stream guard
and was stored and displayed as ordinary assistant content.
strip_think_blocks (storage boundary) and the CLI display copy now strip
an unterminated block-boundary tool-call opener, or any line carrying
stray argument markup, to end of text. The response then reads as empty
and flows through the existing empty-retry path. Complete blocks and
inline prose mentions are unchanged.
Drop the loop-side '(empty)' rewrite (the turn-completion explainer already
owns that at delivery, and gateway/desktop match on the sentinel) and the
extra token-count persistence. Keeps: usage-absent empty streaks arm the
deterministic fast-fail after two attempts with no content or reasoning,
and every completed API call logs even when the provider omits usage
(#101898).
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog). A
profiled session with ~130 children was carrying ~1000 threads. All
of these timers now run on a single process-wide daemon thread.
- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
one Condition-driven daemon thread. schedule(fn, interval) -> handle;
handle.cancel(wait=) blocks for an in-flight run like the old join.
A callback returning False stops itself; a raising callback is logged
at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
_lease_refresh_interval; lease-lost / refresh-error interrupt paths
and the stop-event fencing are unchanged; the join(timeout=1.0) is now
cancel(wait=1.0) so the interrupt clear still runs after any in-flight
tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
schedule(); the poll body is _tick(), same sampling state machine.
Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132. At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:
1. bind_subagent_parent() stored the agent strongly in the
`hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
for its own turn, and every asyncio Handle/Future scheduled during
that turn (LSP reader loops, kernel pipe transports) snapshots the
Context — 56 live Contexts held 14 finished children after the bench.
The ContextVar now holds a weakref (non-weakrefable doubles fall back
to a closure); get_active_subagent_parent() dereferences it.
2. AIAgent.close() cleared _session_messages but not the
_db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
every successful DB flush) nor _streamed_assistant_text_parts, so the
agent — kept alive by (1) — retained every message dict. close() now
drops both.
The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.
Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.
What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.
Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).
Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
`limits=` never reached them, so mounts ran on httpx defaults with a 5 s
keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).
Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
Compaction moved run_bang_command's Popen within the env-guard scanner's
proximity window of _bang_env's os.environ.copy(). Route through the single
factory (build_subprocess_env() == _sanitize_subprocess_env(os.environ.copy()))
and allowlist the file for the import-failure fallback copy with justification.
cmd_sessions used 'with db:' which breaks test doubles lacking the context
manager protocol (13 reds in test_sessions_pin/delete/export). Restore the
explicit try/finally db.close() (same semantics for real SessionDB). Add
get_session/count_prune_matches to the FakeDB doubles in test_sessions_delete
instead of re-adding getattr guards to production code.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).
Missing alias identified by @Steve-prog001 in #101905.
Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.
Fixes#101668
Refs #47742
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.
Both gates now go through one _profile_rejects_tool_media() veto.
Refs #89981
(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)