A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:
1. bind_subagent_parent() stored the agent strongly in the
`hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
for its own turn, and every asyncio Handle/Future scheduled during
that turn (LSP reader loops, kernel pipe transports) snapshots the
Context — 56 live Contexts held 14 finished children after the bench.
The ContextVar now holds a weakref (non-weakrefable doubles fall back
to a closure); get_active_subagent_parent() dereferences it.
2. AIAgent.close() cleared _session_messages but not the
_db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
every successful DB flush) nor _streamed_assistant_text_parts, so the
agent — kept alive by (1) — retained every message dict. close() now
drops both.
The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.
Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.
What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.
Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).
Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
`limits=` never reached them, so mounts ran on httpx defaults with a 5 s
keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).
Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
readPersistedPoolLimits() runs at module evaluation and logs through
rememberLog() on every branch, but hermesLog / desktopLogBuffer /
desktopLogFlushTimer / desktopLogFlushPromise were declared ~110 lines
later. esbuild lowers const/let to var, so the packaged desktop died on
every launch with "Cannot read properties of undefined (reading 'push')"
(#101941, #101960). Moving the four declarations above the read fixes the
crash and keeps the early [pool-limits] line in desktop.log.
Salvaged from #101945 (test dropped: Desktop E2E lane is disabled in CI).
A roster click on a bot whose canonical Bot Chat is already open only
fronted the tile: the pane kept whatever transcript it last painted,
which can predate rows the bot wrote while the user was elsewhere (a
cron delivery, a teammate's message_agent, another bot's turn). The
stale snapshot persisted until the next user turn — #95600's forceResume
only covered the not-yet-open registry path.
Reuse refreshOpenBotChat (the #99393 reclaim mechanism) on the fronted
branch so forceResume re-pulls the latest transcript. Regression test
pins the behavior: fronting an open Bot Chat now requests the canonical
registry open.
The notification action (`NotificationItem`) rendered as
`variant="textStrong" size="xs"` — an 11px underlined muted-grey text link
with a ~44x20px hit target. On the data-training confirm toast raised by
`surfaceModelSwitchConfirm` / `confirmModelWarning` (e.g. picking
`muse-spark-1.2-contributor`) it read as a footnote, not the one action
the toast exists for, and users reported not being able to "press to
accept".
Promote it to the SDK's `default` variant at `size="sm"`: a filled
primary button, larger hit target, obvious affordance. No new styles.
Salvaged from #96562 (toast half only). Refs #96563.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).
Missing alias identified by @Steve-prog001 in #101905.
Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
commandcode (api.commandcode.ai) exposes authoritative
context_length via /models (muse-spark 1M, etc.) but as a
known provider it skipped the custom-endpoint probe at step 2
and has no models.dev entry, so every model fell through to the
256K DEFAULT_FALLBACK. Add a provider-aware branch mirroring
gmi/nous to resolve via _resolve_endpoint_context_length.
Fixes GOAT docs vs status-bar mismatch: muse-spark 1M was shown
as 256K.
Muse Spark 1.2 family (api.meta.ai) ships 1M context (models.dev
opencode/muse-spark-1.2 = 1048576, meta/muse-spark-1.2 = 1048576).
Zen/GO SG /v1/models only returns id (no limit.context), and
models.dev lookup via opencode was missing a hardcoded fallback, so
get_model_context_length fell back to DEFAULT_FALLBACK_CONTEXT=256k.
Banner showed Context: 256,000 for both zen and router-sg lanes.
Add longest-prefix entries 'muse-spark' and 'muse' = 1_048_576 so
all variants (1.1, 1.2, contributor, contributor-free) resolve to 1M
without network.
Add meta/muse-spark-1.3 and meta/muse-spark-1.3-contributor to the
OpenRouter curated list, the meta-ai provider fallback, the
opencode-zen / opencode-free / opencode-go floors, the setup-wizard
shortlist, and regenerate the hosted model catalog.
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.
Fixes#101668
Refs #47742
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.
Both gates now go through one _profile_rejects_tool_media() veto.
Refs #89981
(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)
The naive line scanner's open() regex matched the human-readable phrase
"next picker open (or refresh)." inside the warning string, tripping the
Windows footgun gate. Reword to "next picker open or refresh."
With the picker served from resident caches only, a cold Nous row renders
every model locked (free_tier_pending) until the background prewarm lands.
Surface why on the row's existing warning slot so the user isn't left with
an unexplained greyed-out list; never override an auth warning.
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.
Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
The contributor fix covers remote rows. The reporter's video shows the
sibling shape: with a remote gateway active, the LOCAL twin carries the
'default-this-device' alias, and message_agent's local resolver only
knows bare profile names / 'hermes'. Emit the same target annotation
whenever a local row's alias differs from its resolvable handle.
The Bot Mode mention middleware built message_agent targets from
botHandle(), which prefers a roster row's source-qualified UI alias
("default-vera"). Neither resolver accepts that form — the relay matches
canonical handle/profile (± @connection-id) and the local path a bare
profile name or "hermes" — so remote handoffs died with "No teammate
named" before enqueue.
Annotate the canonical form instead: profile@connection-id for remote
rows, canonical bare handle (default→hermes) for local ones. Pin the
profile@connection form on the relay side too, so the emitted target
stays inside the documented resolver contract (#97678).
Review fold-in on the salvage of #101877:
- `_deliver_result` routed to the durable queue whenever the worker's
`_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
delivering. A worker whose script dispatches another job in-process
(`hermes cron run <other>`) inherits that env and would have queued the
nested job's message under the outer execution id — `INSERT OR IGNORE`
then drops it silently. Match the marker against the delivering job's own
`execution_id`, as `run_one_job` already does. Regression test added
(mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
reported as a delivery error, so `mark_job_run` recorded
`last_status=delivery_failed` for a message the next gateway's drain goes
on to send, and nothing ever corrects the job record. Log and return
success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
of a second hardcoded terminal set.
Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).
Review follow-ups on the final salvage stack:
- `fast_safe_load(f) or {}` collapsed a falsy non-mapping root (`[]`) to `{}`
before the strict isinstance check, so an empty-list config evaded the
raise the previous commit added. Only map a None document (empty file) to
`{}`; every other non-mapping root now hits the strict branch.
- Docstring: the strict flag covers non-mapping roots too, not just parse
failures.
- Tests: assert the tolerant call's return per shape instead of merely
calling it; add the empty-list-root case.
A config.yaml whose top-level value is a list or scalar parses fine, so the
strict check_config_version(raise_on_parse_error=True) from #101778 still
returned (0, latest) and migrate_config() proceeded: sanitize_env_file()
rewrote .env, then save_config()'s fail-closed guard raised RuntimeError.
Raise InvalidUserConfigError up front for that shape too, so the
"no side effect before the invalid config is surfaced" guarantee holds
for both invalid-config shapes. Tolerant callers are unchanged.
Tests: parametrize the two #101778 regression tests over malformed-yaml
and list-root; assert the tolerant call still does not raise. Also make
the test_update_autostash check_config_version mock kwarg-tolerant,
matching the author's fix in test_config.py.
Post-review cleanup on the salvage stack:
- The re-warn-after-reset guard now drives on_session_reset() instead of
the private helper, so a site regressing to a bare rearm-zero fails it.
- One _over_threshold_warnings() helper replaces four inline caplog filters.
- Comment the redundant None check that narrows current_tokens for ty.
Widens the Windows parent-death fix to the class: the kernel inherits
the read end of a pipe whose only write end lives in the host, so host
death by any means (SIGKILL, OOM, crash) is EOF and the kernel exits.
Stdin EOF alone only reaches the runner between cells, so a kernel
SIGKILLed mid-cell used to outlive its host indefinitely. The
integration test now runs on every platform and is trimmed to the
contract (kill host mid-cell, kernel gone), not the handle plumbing.
Same attached-cell guard as the local kernel host (#101861): a remote
kernel mid-cell is never reaped or cap-evicted, so a fan-out never has
its runner killed under a live poll loop.
Local session kernels sweep idle-expired entries and enforce a
process-wide cap (DEFAULT_MAX_SESSION_KERNELS) on every call
(tools/code_kernel.py's _reap_unlocked / _evict_over_cap_unlocked). The
new remote kernel host (#96991) never got the same treatment:
_REMOTE_KERNELS only shrinks lazily when a specific key is revisited and
found dead, so an owner that opens kernels for several distinct
(env_type, task_env_id) combinations (or delegated children) and never
revisits some of them accumulates host-side bookkeeping entries for the
life of the gateway process.
Note this is narrower than the local case: the remote runner already
self-reaps on its own idle timeout, and SSH/Docker connections are
independently bounded by their own transport-level lifecycles (SSH
ControlPersist, Docker's session-scoped idle-timeout in terminal_tool.py)
— so nothing here leaks a live remote connection. What's missing is
purely the host-side dict/cap bookkeeping symmetry with local kernels.
Adds _reap_unlocked/_evict_over_cap_unlocked mirroring the local
implementation, reusing the same max_session_kernels config as an
independent cap on _REMOTE_KERNELS.
Follow-up to the salvaged #101894 (@jwilson411). The over-threshold
"reclamation did not run" warning is deduped on (reason, rearm mark), and
the key was only cleared when a prune committed. Every other path that
zeroes the rearm mark — compress(), on_session_reset/on_session_end,
bind_session_state, update_model — left the key in place, so a lockout
that warned at rearm=0, then a full compaction, then the same lockout
again was silent, contradicting the helper's own "warns again" contract
(and leaking the key across sessions on a rebound compressor).
- ContextCompressor._reset_proactive_prune_rearm(): one helper for the
five rearm-to-zero sites; clears the dedup key alongside the mark.
- _warn_reclamation_no_op(): dropping back under threshold releases the
key (mirrors _clear_context_overflow_warn semantics on the agent side).
- test_proactive_prune_loop_wiring: the attempts_exhausted fixture now
models the only state the real engine can produce for that branch
(should_compress() is should_compress_info()[0]) — budget spent
(max_compression_attempts=0) with the engine saying RUN, instead of
should_compress=False paired with (True, None).
- Two guards: lockout warns again after a rearm reset; dropping under
threshold releases the key. Both fail with the clears removed.
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.
Fixes#101889
Switching to an already-open session remounted the incoming transcript and
re-tokenized every fenced code block from scratch on the main thread (N blocks
x full shiki tokenization per switch, 96-100% CPU for seconds).
- shiki-block: content-keyed LRU cache of highlighted HTML (theme scope +
language + code); remounts of unchanged blocks paint cached markup with
zero highlighter calls; misses debounced, failures degrade to plain text
- use-session-actions: warm resume keeps the session-slice array when the
reconciled content is equivalent (same guard as the cold path), so the
runtime repository and every row keep identity
- transcript-window: per-session sticky window memos; a warm re-visit with an
unchanged transcript reuses the windowed slice by reference (no re-index,
no repository rebuild); sticky cut survives switches for sessions that grew
- perf regression guards: remount must not re-tokenize (codeToHtml called
once per unique block), windowed slice reference preserved across switches,
messageComponents identity stable across session switches
Composition of #92581 on #100985: the hard cap is a constructor constant,
so raising the pool max in Settings would have left new spawns queued
behind the launch-time value. Add setLimit(); slot hand-off now goes
through a single #drain that respects the current cap, which also fixes
the original release path handing a slot to the next waiter even when the
cap had just been lowered (test: lowering never revokes granted slots; new
requests queue until under cap). main.ts constructs from
poolLimits.maxBackends and pushes changes from setPoolLimits(); pinned by
a wiring test.
Hover-intent prewarm sweeps across the Bots rail spawned past the pool cap,
LRU-evicting the backend the user was about to click — an evict/respawn
cascade that made profile switching progressively slower (#91545).
- prewarmProfileBackend skips speculative spawns once every pool slot holds
an open socket; the real click still spawns on demand.
- Pool max/idle become a device preference (Settings -> Advanced), persisted
atomically in userData (pool-limits.json) and applied live over IPC; the
HERMES_DESKTOP_POOL_* env vars remain the initial fallback. Defaults are
unchanged (3 backends / 10 min idle).
Squash of the 3-commit PR #92581 branch (a00dc088c5..783899d12f) applied
via diff onto the spawn-coordinator salvage; import + constant-block
conflicts resolved so the coordinator is constructed from, and follows,
the live preference (setLimit added in the next commit).
Follow-up to the spawn coordinator: the queued ticket waited up to
POOL_IDLE_MS (10 min) for a free local slot, but the renderer gives up on
a backend boot after 45 s. A user clicking a 4th profile with 3 fresh
backends open would see the generic "backend didn't come up" error while
the ticket kept the pool key hostage, so every later click joined the same
stale wait. Cap the wait at 30 s, log the slot pressure when it happens, and
pin the relationship to BACKEND_BOOT_WAIT_TIMEOUT_MS with a wiring test
(fails when the timeout is reverted). Also eslint --fix on the salvaged
files (import order was a lint error).
Desktop could spawn a local hermes serve per profile with no hard cap on
starting+running children: LRU eviction spares keepalive-fresh entries, so
a roster refresh across many profiles became a process wave (40+ backends,
load 30-50 reported).
LocalBackendSpawnCoordinator: at most POOL_MAX_BACKENDS local backends may
be starting or running. Remote descriptors never take a slot. Queue tickets
are per request; a slot is released only after process exit is proven
(exitCode/signalCode). A rejected wait keeps the slot occupied. Pool entries
re-assert ownership at each await so an evicted entry cannot spawn a zombie.
Squash of PR #100985 (6017abbbc4 + merge), applied via diff onto current
main. Original commits were authored as 'Motor (Hermes AI) <ceo@xtremagency.com>';
attributed here to the PR author's GitHub identity.