Commit Graph

27609 Commits

Author SHA1 Message Date
Teknium c96568f66c perf(delegation): finished delegate children no longer pin their transcripts in the parent heap
A parent that fanned out 1,320 subagents over 13h reached 2.6 GB RSS
(1.9 GB anonymous heap). Every closed child AIAgent stayed reachable and
still owned a copy of its full message history. gc.get_referrers on a
finished child (30-child fan-out bench, evals/fanout_resource_bench.py)
showed two retainers:

1. bind_subagent_parent() stored the agent strongly in the
   `hermes_subagent_lifecycle_parent` ContextVar. Each child binds ITSELF
   for its own turn, and every asyncio Handle/Future scheduled during
   that turn (LSP reader loops, kernel pipe transports) snapshots the
   Context — 56 live Contexts held 14 finished children after the bench.
   The ContextVar now holds a weakref (non-weakrefable doubles fall back
   to a closure); get_active_subagent_parent() dereferences it.

2. AIAgent.close() cleared _session_messages but not the
   _db_flush_scan_prefix snapshot (a `messages[:]` shallow copy taken on
   every successful DB flush) nor _streamed_assistant_text_parts, so the
   agent — kept alive by (1) — retained every message dict. close() now
   drops both.

The delegate_task result entry never carried `messages`; a pin test
confirms the per-child result JSON is unchanged.

Bench (30 children / 10 worktrees, ~100 KB final replies so retention is
visible): post-fan-out live child AIAgents 14 -> 0; RSS after fan-out
636 MB -> 556 MB. With the harness' tiny default replies both runs sit at
~192-194 MB (the children's transcripts were never the dominant cost
there; the leaked objects were).
2026-09-03 02:35:37 -07:00
Teknium c3b411dfb7 perf(agents): share one httpx transport pool across every agent's client
A fan-out of 30 delegated children built 183 httpx.HTTPTransport objects
(each with its own httpcore pool + parsed SSL context): 3 per agent x
(primary + aux clients). A profiled session with ~130 children held 107 TLS
sockets to one provider. Peak RSS for the 30-child bench drops 286 -> 195 MB;
live HTTPTransports 183 -> 2, ConnectionPools 183 -> 7.

What is shared: the sync `HTTPTransport` (pool + SSL context) per
(scheme, verify, proxy, happy-eyeballs) identity, in a bounded module dict.
What is NOT shared: the per-agent `httpx.Client` wrapper. Each client mounts
a `_SharedTransport` view whose `close()` marks only that view closed and
never touches the pool, so the #10933 contract (close client A, build client
B, B works) holds unchanged — the pinning tests in
test_create_openai_client_reuse.py / test_sequential_chats_live.py pass as-is.

Safety for cross-thread aborts: `_SharedTransport.handle_request` stamps its
id into `request.extensions`; `_iter_pool_sockets` now only shuts down a
shared pool's in-flight requests carrying the calling client's stamp and
never its idle connections, so interrupting child A cannot sever child B's
stream (#29507 / #72975 walker semantics preserved for unshared pools).

Also:
- `resolve_httpx_verify` caches one SSLContext per CA-bundle path. With
  SSL_CERT_FILE/HERMES_CA_BUNDLE set, every agent used to parse the bundle
  again and — because the share key is context identity — get a private pool.
- The client no longer builds a third, unused default transport; its
  default transport is the https view.
- Mounted transports now actually receive pool limits (Client-level
  `limits=` never reached them, so mounts ran on httpx defaults with a 5 s
  keepalive_expiry). The shared pool uses 50 keepalive / 1000 max so one
  pool covers a whole concurrent fan-out.
- `close_shared_transports()` really closes the pools (tests / shutdown).

Async clients (`async_mode=True`) stay unshared: an httpcore async pool is
bound to the event loop that first uses it. Proxy-backed clients keep
httpx's per-client proxy transport.
2026-09-03 02:35:21 -07:00
Teknium 80fae22bf5 fix(lsp): share one pyright process across git worktrees via workspaceFolders
Multi-root servers (pyright) are keyed by server_id; a file whose resolved
root is new for a running client is attached with
workspace/didChangeWorkspaceFolders instead of spawning another server.
Single-root servers keep the (server_id, workspace_root) key and behavior.
A profiled fan-out across ~30 worktrees ran 30-60 pyright processes
(~8.7 GB); the same fan-out now runs one.
2026-09-03 02:32:15 -07:00
Teknium 9cee679831 bench: explicit utf-8 encoding on text-mode opens (ruff PLW1514) 2026-09-03 02:31:59 -07:00
Teknium 45b0a0ae25 bench: --reply-kb payload padding, transport/live-agent gc counts, after-snapshot rows in --compare 2026-09-03 02:31:59 -07:00
Teknium c77b9d637b bench: fan-out resource harness (threads/RSS/fds/pyright/kernels/httpx clients per N children x W worktrees) 2026-09-03 02:31:59 -07:00
liuhao1024 05f548f35d fix(desktop): declare rememberLog state before the top-level pool-limits read
readPersistedPoolLimits() runs at module evaluation and logs through
rememberLog() on every branch, but hermesLog / desktopLogBuffer /
desktopLogFlushTimer / desktopLogFlushPromise were declared ~110 lines
later. esbuild lowers const/let to var, so the packaged desktop died on
every launch with "Cannot read properties of undefined (reading 'push')"
(#101941, #101960). Moving the four declarations above the read fixes the
crash and keeps the early [pool-limits] line in desktop.log.

Salvaged from #101945 (test dropped: Desktop E2E lane is disabled in CI).
2026-09-03 01:14:37 -07:00
cmyyy 3ea71a47b3 fix(desktop): refresh Bot Chat transcript when a roster click fronts an already-open tab
A roster click on a bot whose canonical Bot Chat is already open only
fronted the tile: the pane kept whatever transcript it last painted,
which can predate rows the bot wrote while the user was elsewhere (a
cron delivery, a teammate's message_agent, another bot's turn). The
stale snapshot persisted until the next user turn — #95600's forceResume
only covered the not-yet-open registry path.

Reuse refreshOpenBotChat (the #99393 reclaim mechanism) on the fronted
branch so forceResume re-pulls the latest transcript. Regression test
pins the behavior: fronting an open Bot Chat now requests the canonical
registry open.
2026-09-03 01:10:16 -07:00
Edder Talmor 37fd6eea97 fix(desktop): toast action is a real button, not a hairline text link
The notification action (`NotificationItem`) rendered as
`variant="textStrong" size="xs"` — an 11px underlined muted-grey text link
with a ~44x20px hit target. On the data-training confirm toast raised by
`surfaceModelSwitchConfirm` / `confirmModelWarning` (e.g. picking
`muse-spark-1.2-contributor`) it read as a footnote, not the one action
the toast exists for, and users reported not being able to "press to
accept".

Promote it to the SDK's `default` variant at `size="sm"`: a filled
primary button, larger hit target, obvious affordance. No new styles.

Salvaged from #96562 (toast half only). Refs #96563.
2026-09-03 00:58:46 -07:00
Teknium d0b7cec0b8 fix(prompt): Muse Spark gets tool-use enforcement + execution guidance on defaults (#96550)
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.

Co-authored-by: Edder Talmor <talmoredder@gmail.com>
2026-09-03 00:58:32 -07:00
Teknium 0b96eaf06b chore(contributors): map sosxradar@gmail.com -> GTHell 2026-09-03 00:58:16 -07:00
Teknium 4359af7705 fix(models_dev): alias opencode-free to the Zen "opencode" catalog; pin Muse Spark 1M invariant
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).

Missing alias identified by @Steve-prog001 in #101905.

Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
2026-09-03 00:58:16 -07:00
Steve-prog001 779aecb62b fix(context): resolve commandcode models via live /models
commandcode (api.commandcode.ai) exposes authoritative
context_length via /models (muse-spark 1M, etc.) but as a
known provider it skipped the custom-endpoint probe at step 2
and has no models.dev entry, so every model fell through to the
256K DEFAULT_FALLBACK. Add a provider-aware branch mirroring
gmi/nous to resolve via _resolve_endpoint_context_length.

Fixes GOAT docs vs status-bar mismatch: muse-spark 1M was shown
as 256K.
2026-09-03 00:58:16 -07:00
GTHell bb8f4afa46 fix(context): add muse-spark 1M fallback (zen/GO SG showed 256k)
Muse Spark 1.2 family (api.meta.ai) ships 1M context (models.dev
opencode/muse-spark-1.2 = 1048576, meta/muse-spark-1.2 = 1048576).

Zen/GO SG /v1/models only returns id (no limit.context), and
models.dev lookup via opencode was missing a hardcoded fallback, so
get_model_context_length fell back to DEFAULT_FALLBACK_CONTEXT=256k.
Banner showed Context: 256,000 for both zen and router-sg lanes.

Add longest-prefix entries 'muse-spark' and 'muse' = 1_048_576 so
all variants (1.1, 1.2, contributor, contributor-free) resolve to 1M
without network.
2026-09-03 00:58:16 -07:00
Teknium fa53e4fedd chore(contributors): map csreyes92@gmail.com -> csreyes (salvage #93073) 2026-09-03 00:57:55 -07:00
mr-r0b0t cfa7e72c9e fix(models): correct contributor guard, 1M context, docs for muse-spark-1.3
- model_data_policy_guard: name the triggering -contributor model instead
  of hardcoded 1.2; per-version verified price tables (1.3 standard
  $1.25/$4.25 via OpenRouter live metadata; cached figures 1.2-only)
- model_metadata: muse-spark-1.3 + muse-spark family at 1048576 (OpenRouter
  verified 2026-09-02) with pre-catalog stale-cache keys so 256K-fallback
  sessions self-heal
- docs: contributor-tier notes cover 1.2 + 1.3
- tests: 1.3 guard regression, muse stale-cache guard, live-catalog mirror
  gains 1.3-contributor-free (confirmed on live relay)

143 tests pass (guard, selection guards, opencode catalog, model_metadata);
ruff clean.
2026-09-03 00:57:55 -07:00
mr-r0b0t 7e60d0c042 feat(models): add Meta Muse Spark 1.3 family to picker
Add meta/muse-spark-1.3 and meta/muse-spark-1.3-contributor to the
OpenRouter curated list, the meta-ai provider fallback, the
opencode-zen / opencode-free / opencode-go floors, the setup-wizard
shortlist, and regenerate the hosted model catalog.
2026-09-03 00:57:55 -07:00
Christian Reyes 7f2aa70add feat(models): add Muse Spark contributor to OpenRouter 2026-09-03 00:57:55 -07:00
Teknium d6bb94a1fb fix(meta-ai): declare supports_vision_tool_messages=False — Muse Spark 400s on image tool results
Muse Spark accepts images on user turns but returns HTTP 400
invalid_request_error 'messages[N].content did not match any supported
type' when the vision_analyze multimodal envelope lands in a role:tool
message. With the profile veto now honored by the vision fast-path gates,
declaring the limitation routes tool-result images through the aux-LLM
text path while user-message vision stays enabled.

Fixes #101668
Refs #47742
2026-09-03 00:57:40 -07:00
liuhao1024 b462989a68 fix(vision): honor supports_vision_tool_messages=False in tool-result media gates
A ProviderProfile that declares supports_vision_tool_messages=False accepts
images in user messages but rejects list-type tool-result content with 400
(xiaomi/MiMo "text is not set"). supports_vision=True alone used to flip
_supports_media_in_tool_results to True, and a vision-capable capability
lookup could re-open _should_use_native_vision_fast_path — so the native
multimodal envelope landed in a role:tool message and 400'd every turn.

Both gates now go through one _profile_rejects_tool_media() veto.

Refs #89981

(cherry picked from commit daed88f940a6a475f12bb435498e48181da60f4d, trimmed)
2026-09-03 00:57:40 -07:00
kshitijk4poor b9dc790332 fix(inventory): reword pending-entitlement warning to avoid Windows footgun false positive
The naive line scanner's open() regex matched the human-readable phrase
"next picker open (or refresh)." inside the warning string, tripping the
Windows footgun gate. Reword to "next picker open or refresh."
2026-09-03 13:08:31 +05:30
kshitijk4poor fb723084be fix(inventory): explain the locked Nous list while entitlement is pending
With the picker served from resident caches only, a cold Nous row renders
every model locked (free_tier_pending) until the background prewarm lands.
Surface why on the row's existing warning slot so the user isn't left with
an unexplained greyed-out list; never override an auth warning.
2026-09-03 13:08:31 +05:30
fangliquanflq 401ac7e1f8 fix(gateway): keep first model picker open responsive on cold pricing cache
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.

Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
2026-09-03 13:08:31 +05:30
Teknium b6041240d1 test(desktop): trim Bot Chat pane-focus regression to the two invariant cases
Salvage follow-up to #101639 (@helix4u): keep the re-adopt-after-overlay and
miss-propagation cases, drop the hidden-pane and healthy-path controls.
2026-09-03 00:36:09 -07:00
Gille b6d549d002 fix(desktop): restore missing Bot Chat panes before claiming focus 2026-09-03 00:36:09 -07:00
hermes-seaeye[bot] e629c900a8 fmt(js): npm run fix on merge (#101952)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-03 07:24:01 +00:00
Teknium bd6cc48b94 fix(desktop): annotate the resolvable target for aliased LOCAL rows too
The contributor fix covers remote rows. The reporter's video shows the
sibling shape: with a remote gateway active, the LOCAL twin carries the
'default-this-device' alias, and message_agent's local resolver only
knows bare profile names / 'hermes'. Emit the same target annotation
whenever a local row's alias differs from its resolvable handle.
2026-09-03 00:18:31 -07:00
liuhao1024 2e542c92e6 fix(desktop): annotate canonical relay targets for remote @mentions
The Bot Mode mention middleware built message_agent targets from
botHandle(), which prefers a roster row's source-qualified UI alias
("default-vera"). Neither resolver accepts that form — the relay matches
canonical handle/profile (± @connection-id) and the local path a bare
profile name or "hermes" — so remote handoffs died with "No teammate
named" before enqueue.

Annotate the canonical form instead: profile@connection-id for remote
rows, canonical bare handle (default→hermes) for local ones. Pin the
profile@connection form on the relay side too, so the emitted target
stays inside the documented resolver contract (#97678).
2026-09-03 00:18:31 -07:00
kshitijk4poor c3e9b28a42 fix(cron): key worker deliveries by the job's own attempt; deferred send is not a failure
Review fold-in on the salvage of #101877:

- `_deliver_result` routed to the durable queue whenever the worker's
  `_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
  delivering. A worker whose script dispatches another job in-process
  (`hermes cron run <other>`) inherits that env and would have queued the
  nested job's message under the outer execution id — `INSERT OR IGNORE`
  then drops it silently. Match the marker against the delivering job's own
  `execution_id`, as `run_one_job` already does. Regression test added
  (mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
  reported as a delivery error, so `mark_job_run` recorded
  `last_status=delivery_failed` for a message the next gateway's drain goes
  on to send, and nothing ever corrects the job record. Log and return
  success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
  of a second hardcoded terminal set.
2026-09-03 12:40:26 +05:30
kshitijk4poor b440a492b3 fix(cron): keep unsent worker deliveries queued, cheapen handoff polling
Follow-up to the salvaged restart-safe worker (#101877):

- delivery_queue: a row still `pending` at the worker's wait timeout was
  marked `failed` and never drained, so any gateway outage longer than the
  300s budget (e.g. a restart that runs `hermes update`) silently lost the
  delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
  queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
  every transaction (each `get_status` poll paid for it; terminalizing
  paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
  (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
  `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
  owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
  ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
  process with a 1s timeout instead — the worker commits its terminal row
  before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
  created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
  runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
  `mark_execution_running -> None`, which now means "ownership lost, return
  before run_job" — the test passed without ever reaching the path it
  names. Mocking `{}` restores it (mutation-checked).
2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 83efdf5e5e [verified] fix(cron): close restart handoff races 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 29e5172487 [verified] fix(cron): harden gateway restart handoff 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson 3373e97693 [verified] fix(cron): preserve active runs across gateway restart 2026-09-03 12:40:26 +05:30
kshitijk4poor 68cbe484a4 refactor(config): close empty-list root gap in strict check; assert tolerant returns
Review follow-ups on the final salvage stack:
- `fast_safe_load(f) or {}` collapsed a falsy non-mapping root (`[]`) to `{}`
  before the strict isinstance check, so an empty-list config evaded the
  raise the previous commit added. Only map a None document (empty file) to
  `{}`; every other non-mapping root now hits the strict branch.
- Docstring: the strict flag covers non-mapping roots too, not just parse
  failures.
- Tests: assert the tolerant call's return per shape instead of merely
  calling it; add the empty-list-root case.
2026-09-03 12:38:46 +05:30
kshitijk4poor 4d24f357a0 fix(config): strict version check also refuses non-mapping config roots
A config.yaml whose top-level value is a list or scalar parses fine, so the
strict check_config_version(raise_on_parse_error=True) from #101778 still
returned (0, latest) and migrate_config() proceeded: sanitize_env_file()
rewrote .env, then save_config()'s fail-closed guard raised RuntimeError.
Raise InvalidUserConfigError up front for that shape too, so the
"no side effect before the invalid config is surfaced" guarantee holds
for both invalid-config shapes. Tolerant callers are unchanged.

Tests: parametrize the two #101778 regression tests over malformed-yaml
and list-root; assert the tolerant call still does not raise. Also make
the test_update_autostash check_config_version mock kwarg-tolerant,
matching the author's fix in test_config.py.
2026-09-03 12:38:46 +05:30
EmanueleCornaggia a3ceee4333 fix(config): refuse migration on malformed YAML
Make validation and migration paths distinguish parse failures from current configs so invalid YAML cannot trigger .env or config-side effects.
2026-09-03 12:38:46 +05:30
kshitijk4poor db40b2c67d chore: add EmanueleCornaggia to contributor map (#101778 salvage) 2026-09-03 12:38:46 +05:30
kshitij de9231a0d7 Merge pull request #101936 from kshitijk4poor/test/101922-rearm-reset-public-boundary
test(compression): pin prune rearm-reset wiring through a public boundary
2026-09-03 12:35:35 +05:30
kshitijk4poor c06f7cf924 test(compression): pin rearm-reset wiring through a public boundary
Post-review cleanup on the salvage stack:
- The re-warn-after-reset guard now drives on_session_reset() instead of
  the private helper, so a site regressing to a bare rearm-zero fails it.
- One _over_threshold_warnings() helper replaces four inline caplog filters.
- Comment the redundant None check that narrows current_tokens for ty.
2026-09-03 12:26:22 +05:30
Teknium 365e2835d4 fix(execute_code): POSIX kernels exit when their host dies mid-cell (death pipe)
Widens the Windows parent-death fix to the class: the kernel inherits
the read end of a pipe whose only write end lives in the host, so host
death by any means (SIGKILL, OOM, crash) is EOF and the kernel exits.
Stdin EOF alone only reaches the runner between cells, so a kernel
SIGKILLed mid-cell used to outlive its host indefinitely. The
integration test now runs on every platform and is trimmed to the
contract (kill host mid-cell, kernel gone), not the handle plumbing.
2026-09-02 23:55:30 -07:00
Dolverin 32d5e9d357 fix(execute_code): stop Windows kernels with backend parent 2026-09-02 23:55:30 -07:00
Teknium 0177c16903 fix(code-execution): remote kernel reap/evict skip kernels with a running cell
Same attached-cell guard as the local kernel host (#101861): a remote
kernel mid-cell is never reaped or cap-evicted, so a fan-out never has
its runner killed under a live poll loop.
2026-09-02 23:55:14 -07:00
nftpoetrist d4126c6f49 fix(code-execution): reap idle and cap remote session kernels
Local session kernels sweep idle-expired entries and enforce a
process-wide cap (DEFAULT_MAX_SESSION_KERNELS) on every call
(tools/code_kernel.py's _reap_unlocked / _evict_over_cap_unlocked). The
new remote kernel host (#96991) never got the same treatment:
_REMOTE_KERNELS only shrinks lazily when a specific key is revisited and
found dead, so an owner that opens kernels for several distinct
(env_type, task_env_id) combinations (or delegated children) and never
revisits some of them accumulates host-side bookkeeping entries for the
life of the gateway process.

Note this is narrower than the local case: the remote runner already
self-reaps on its own idle timeout, and SSH/Docker connections are
independently bounded by their own transport-level lifecycles (SSH
ControlPersist, Docker's session-scoped idle-timeout in terminal_tool.py)
— so nothing here leaks a live remote connection. What's missing is
purely the host-side dict/cap bookkeeping symmetry with local kernels.

Adds _reap_unlocked/_evict_over_cap_unlocked mirroring the local
implementation, reusing the same max_session_kernels config as an
independent cap on _REMOTE_KERNELS.
2026-09-02 23:55:14 -07:00
kshitijk4poor 58a4a11727 fix(compression): release the reclamation no-op dedup key on every rearm reset
Follow-up to the salvaged #101894 (@jwilson411). The over-threshold
"reclamation did not run" warning is deduped on (reason, rearm mark), and
the key was only cleared when a prune committed. Every other path that
zeroes the rearm mark — compress(), on_session_reset/on_session_end,
bind_session_state, update_model — left the key in place, so a lockout
that warned at rearm=0, then a full compaction, then the same lockout
again was silent, contradicting the helper's own "warns again" contract
(and leaking the key across sessions on a rebound compressor).

- ContextCompressor._reset_proactive_prune_rearm(): one helper for the
  five rearm-to-zero sites; clears the dedup key alongside the mark.
- _warn_reclamation_no_op(): dropping back under threshold releases the
  key (mirrors _clear_context_overflow_warn semantics on the agent side).
- test_proactive_prune_loop_wiring: the attempts_exhausted fixture now
  models the only state the real engine can produce for that branch
  (should_compress() is should_compress_info()[0]) — budget spent
  (max_compression_attempts=0) with the engine saying RUN, instead of
  should_compress=False paired with (True, None).
- Two guards: lockout warns again after a rearm reset; dropping under
  threshold releases the key. Both fail with the clears removed.
2026-09-03 12:23:55 +05:30
Justin Wilson 9f12121206 fix(compression): do not let prune rearm lock out over-threshold sessions
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.

Fixes #101889
2026-09-03 12:23:55 +05:30
Finn763 e245e40f73 fix(desktop): warm session switch pegs renderer main thread (#95595)
Switching to an already-open session remounted the incoming transcript and
re-tokenized every fenced code block from scratch on the main thread (N blocks
x full shiki tokenization per switch, 96-100% CPU for seconds).

- shiki-block: content-keyed LRU cache of highlighted HTML (theme scope +
  language + code); remounts of unchanged blocks paint cached markup with
  zero highlighter calls; misses debounced, failures degrade to plain text
- use-session-actions: warm resume keeps the session-slice array when the
  reconciled content is equivalent (same guard as the cold path), so the
  runtime repository and every row keep identity
- transcript-window: per-session sticky window memos; a warm re-visit with an
  unchanged transcript reuses the windowed slice by reference (no re-index,
  no repository rebuild); sticky cut survives switches for sessions that grew
- perf regression guards: remount must not re-tokenize (codeToHtml called
  once per unique block), windowed slice reference preserved across switches,
  messageComponents identity stable across session switches
2026-09-03 12:19:11 +05:30
kshitijk4poor f260ea5347 fix(desktop): let the spawn coordinator follow the live pool max
Composition of #92581 on #100985: the hard cap is a constructor constant,
so raising the pool max in Settings would have left new spawns queued
behind the launch-time value. Add setLimit(); slot hand-off now goes
through a single #drain that respects the current cap, which also fixes
the original release path handing a slot to the next waiter even when the
cap had just been lowered (test: lowering never revokes granted slots; new
requests queue until under cap). main.ts constructs from
poolLimits.maxBackends and pushes changes from setPoolLimits(); pinned by
a wiring test.
2026-09-03 12:19:05 +05:30
ClintonEmok c401756a6a fix(desktop): pool sizing as a live device preference in Settings (#91545)
Hover-intent prewarm sweeps across the Bots rail spawned past the pool cap,
LRU-evicting the backend the user was about to click — an evict/respawn
cascade that made profile switching progressively slower (#91545).

- prewarmProfileBackend skips speculative spawns once every pool slot holds
  an open socket; the real click still spawns on demand.
- Pool max/idle become a device preference (Settings -> Advanced), persisted
  atomically in userData (pool-limits.json) and applied live over IPC; the
  HERMES_DESKTOP_POOL_* env vars remain the initial fallback. Defaults are
  unchanged (3 backends / 10 min idle).

Squash of the 3-commit PR #92581 branch (a00dc088c5..783899d12f) applied
via diff onto the spawn-coordinator salvage; import + constant-block
conflicts resolved so the coordinator is constructed from, and follows,
the live preference (setLimit added in the next commit).
2026-09-03 12:19:05 +05:30
kshitijk4poor 5810172f52 fix(desktop): bound the pool-slot wait below the renderer boot budget
Follow-up to the spawn coordinator: the queued ticket waited up to
POOL_IDLE_MS (10 min) for a free local slot, but the renderer gives up on
a backend boot after 45 s. A user clicking a 4th profile with 3 fresh
backends open would see the generic "backend didn't come up" error while
the ticket kept the pool key hostage, so every later click joined the same
stale wait. Cap the wait at 30 s, log the slot pressure when it happens, and
pin the relationship to BACKEND_BOOT_WAIT_TIMEOUT_MS with a wiring test
(fails when the timeout is reverted). Also eslint --fix on the salvaged
files (import order was a lint error).
2026-09-03 12:19:05 +05:30
Kryptonator e924615bb1 fix(desktop): cap concurrent local profile backend spawns
Desktop could spawn a local hermes serve per profile with no hard cap on
starting+running children: LRU eviction spares keepalive-fresh entries, so
a roster refresh across many profiles became a process wave (40+ backends,
load 30-50 reported).

LocalBackendSpawnCoordinator: at most POOL_MAX_BACKENDS local backends may
be starting or running. Remote descriptors never take a slot. Queue tickets
are per request; a slot is released only after process exit is proven
(exitCode/signalCode). A rejected wait keeps the slot occupied. Pool entries
re-assert ownership at each await so an evicted entry cannot spawn a zombie.

Squash of PR #100985 (6017abbbc4 + merge), applied via diff onto current
main. Original commits were authored as 'Motor (Hermes AI) <ceo@xtremagency.com>';
attributed here to the PR author's GitHub identity.
2026-09-03 12:19:05 +05:30