The statusbar gauge painted only what the backend reported as measured
occupancy, which a session has none of until a turn runs in this process.
Turning the gauge on mid-conversation, or resuming a chat, therefore showed
nothing until the next message.
Fetch session.context_breakdown as soon as the gauge is on screen instead of
when its popover opens. It is the same read-only estimate the popover already
used (chars/4 over the live prompt, tools and transcript — no provider call),
and it reports the measured figure once the backend has one. The popover
becomes presentational and reads the gauge s merged usage, so the bar and the
panel cannot disagree.
delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.
Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
the OLD default 50 -> 250 on update, so existing installs inherit the new
headroom. Any other explicit value (deliberate override) is preserved;
unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default
The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.
Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
ActualProfile.fetch_models() overrides ProviderProfile's default
implementation with its own Actual-specific base_url resolution
(ACTUAL_BASE_URL env var, hosted-vs-local normalization), but called raw
urllib.request.urlopen(req, timeout=timeout) directly instead of the base
class's open_credentialed_url(). Every other provider either uses the
base class default or forwards to it via super() and gets
SafeCredentialRedirectHandler for free — Actual is the only provider that
attaches a Bearer token to its own Request object and opens it with the
stdlib's default redirect handling, which forwards every header,
including Authorization, across a cross-origin redirect.
Actual's own feature surface makes the trigger realistic: ACTUAL_BASE_URL
is a first-class, documented way to point this provider at a self-hosted
or local-offline endpoint (see the local-loopback no-auth path already
handled elsewhere in this provider), so a misconfigured or compromised
endpoint 302-ing to another host leaks ACTUAL_API_KEY to it.
Fix: import and call the same open_credentialed_url() the base class
uses, keeping Actual's own URL-resolution logic unchanged.
Adds an end-to-end regression test using two real local HTTP servers (no
mocking of the security module itself) — one redirects, the other
records the Authorization header it receives — mirroring
test_urllib_security.py's own redirect tests. Also repoints the existing
fetch_models test's mock from urllib.request.urlopen to
hermes_cli.urllib_security.open_credentialed_url, since fetch_models no
longer calls the former. Mutation-verified: the new redirect test fails
on pre-fix code with the Authorization header observed at the redirect
target.
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:
- the requests-based metadata/pricing probe
(agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
(hermes_cli/models.py::probe_api_models)
A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.
This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:
- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
ssl_ca_cert before falling back to the env vars. Callers with no base_url
keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
passes it through open_credentialed_url, which gains an ssl_context seam on
the cloned secure opener. Unmatched or public endpoints pass None and keep
urllib's default policy.
Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.
For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.
Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.
Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
Follow-up on the #83854 salvage: prepend $HERMES_HOME/bin ahead of the
venv and user-local bin dirs, matching the managed-first Browser Use
CLI resolution policy — the worker resolves the same canonical binary
the agent process does.
The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).
Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.
Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
Adds the full MCP setup surface as profile-scoped gateway RPCs so a
desktop client (Bot Mode's bot editor, the core Capabilities tab) can
add/configure/test/authenticate/remove MCP servers for ANY profile, not
just the launch profile:
- mcp.servers.list (profile) -> configured servers (transport, auth,
oauth_tokens_present, enabled, tool names; no secret values)
- mcp.servers.add (profile, name, config|preset, bearer_token?) -> reuses
mcp_config._apply_mcp_preset / _save_mcp_server / _save_bearer_auth_token
- mcp.servers.set_api_key (profile, name, value, env_var?) -> http auth
header template or stdio env ref, via save_env_value
- mcp.servers.test (profile, name) -> _probe_single_server + oauth state
- mcp.servers.remove (profile, name)
- mcp.servers.oauth.start/poll (profile, name[, session_id]) -> mirrors the
PROVIDER oauth session/poll model (not the FastAPI dashboard flow): a
background worker drives the same interactive machinery 'hermes mcp login'
uses, capturing the browser redirect on a local loopback listener. Client
opens auth_url via openExternal and polls until status=='approved'.
All handlers are profile-scoped via set_hermes_home_override in try/finally
(mirrors skills.manage). Shared helpers live in tui_gateway/mcp_rpc_helpers.py
and are aliased onto server.py's namespace so the rebound handler bodies
(HandlerRegistry.install) can resolve them — a plain def in methods_tools is
unreachable post-rebind. Reuses hermes_cli/mcp_config.py throughout; no config
logic duplicated; no raw yaml near config.yaml (config-read-guard safe).
Tests: tests/tui_gateway/test_mcp_profile_rpcs.py, 8 E2E against real temp
HERMES_HOME profiles asserting add/list/set_api_key/remove land in the RIGHT
profile's config.yaml and not the launch profile's. 8/8. Registration +
live mcp.servers.list verified in an imported gateway.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
test_run_prompt_submit_requeues_all_unstarted_notifications_with_real_threading
failed twice in one hour on CI slices for two UNRELATED PRs (86371,
86374) with `assert set() == {proc_batch_2, proc_batch_3}`. Root cause:
session.init/create tests earlier in the file start real per-session
notification poller daemon threads and never stop them. Those pollers
outlive their test and keep polling the PROCESS-GLOBAL
process_registry.completion_queue, stealing-and-requeuing the target
test's events mid-assertion so its bounded drain loop can starve.
Reproduced: with 30 leaked foreign-session pollers injected via a
sabotage conftest, the target test fails standalone ~1 in 3 runs with
the exact CI assertion; with the reap fixture active it passed 8/8
under the same sabotage.
Fix:
- tui_gateway/server.py: _start_notification_poller registers
(stop_event, thread) in module-level _notification_pollers (pruned of
dead threads on each spawn; threads get a stable
tui-notif-poller-<sid> name for debugging).
- tests/test_tui_gateway_server.py: autouse fixture sets every
registered live poller's stop event after each test and joins them
under ONE shared 3s budget (the poller loop wakes at least every
0.5s), so no poller survives into the next test. No per-thread
timeout, no session-dict mutation — a first draft that mutated
session state and joined per-thread hung the file; full-file runtime
with this version is 15.2s vs 13.2s baseline.
* fix(desktop): the main agent's model pick persists as the profile default
Reported: the default bot switches to the OpenAI API account instead of
the user's subscription, and doesn't retain the previous selection.
Root cause: the composer model picker always sent the switch as
--session scope, even for the PRIMARY profile's main agent. So the pick
never wrote config.yaml model.provider — and with model.provider unset,
resolve_provider('auto') falls through to a leftover OPENAI_API_KEY env
var and picks OpenAI/OpenRouter. The subscription the user selected was
only ever a per-session override that evaporated on the next session.
Fix: when the pick targets the primary profile's main agent
(touchesPrimary), send --global so it persists to config.yaml
(model.default + model.provider) via the existing model-switch persist
path. A SET model.provider already outranks the OPENAI_API_KEY env var
in resolve_provider (tier 2 vs tier 3), so the main agent now keeps the
chosen provider across restarts. Secondary chat tiles stay --session so
picking a model in one chat never rewrites the profile default (the
cross-session-contamination guard the old comment protected).
No change to resolve_provider's priority chain, so #29285 (an explicit
env key beating a STALE oauth login) is untouched — we simply make the
user's explicit main-agent selection the config default it always
should have been.
* MoA presets stay session-scoped; update tests for primary-persist intent
Fix CI (ui shard 3of3): the primary main-agent pick now persists via
--global, but MoA (mixture-of-agents) presets must NOT — a transient
orchestration choice can't become the global gateway default. Exclude
provider==='moa' from the persist path (stays --session). Update the
primary-picker test to assert --global (the new intent) and keep the
MoA + secondary-tile tests asserting --session (the guards that prove
the narrowing). 19/19 green locally.
Follow-up to the salvaged #66538 commit: the ZELLIJ env gate fixed
detection, but writeDiffToTerminal still wrapped main-screen frames in
BSU/ESU unconditionally (skipSyncMarkers was only set for alt-screen).
Under Zellij the multiplexer re-chunks the stream, so the markers buy no
atomicity and stale frames leak into main-screen scrollback as the
repeated chrome reported in #66490. The renderer now passes
!SYNC_OUTPUT_SUPPORTED for every write path; supported terminals keep
today's behavior on both screens. Adds emitted-frame regression tests
for both marker modes.
isSynchronizedOutputSupported() only excluded tmux, so running inside
Zellij under an outer terminal that advertises DEC 2026 (e.g. WezTerm
via TERM_PROGRAM) returned true. Zellij, like tmux, sits between us and
the outer terminal and chunks the stream, breaking BSU/ESU atomicity and
pushing old TUI frames into scrollback as repeated output.
Guard on the ZELLIJ env var (set to the session index, e.g. "0") the
same way we already guard on TMUX. Also thread an optional env argument
through the function so the behavior is unit-testable, mirroring
needsAltScreenResizeScrollbackClear() in the same module.
Closes#66490
Per hermes-sweeper review suggestion on #50242: the existing tests only
covered the unset case (override is None after the call). Add a test
where a caller already holds an override — _persist_live_session_system_prompt
must build the prompt under the session's profile and then restore the
caller's override via the reset token, not clear it to None.
Fixes#50233
_persist_live_session_system_prompt rebuilds the system prompt after a
live model switch (/model), but _start_agent_build's finally block has
already reset set_hermes_home_override by then. load_soul_md() and
build_skills_system_prompt() call get_hermes_home() which falls back to
the root ~/.hermes, loading the wrong SOUL.md identity and skills for
the session's profile.
Fix: set_hermes_home_override(session["profile_home"]) before calling
agent._build_system_prompt() in _persist_live_session_system_prompt,
and reset it in a finally block. Also upgrade the failure log from
DEBUG to WARNING so silent profile-wrong-prompt issues are visible.
The first-prompt lazy-build path is unaffected — _run_prompt_submit
already re-sets the override before run_conversation. The /model slash
worker path does not, which is the gap this fixes.
Co-Authored-By: Claude <noreply@anthropic.com>
Follow-up to the salvaged #41484 commit:
- usageChanged() iterates the union of Usage keys generically instead of a
hardcoded field list — the original PR's list omitted active_subagents
(consumed by the status rule's subagent segment and resume hint), which
would have suppressed legitimate updates
- The memo(StatusRule) half of the original PR is intentionally dropped:
main's StatusRule gained battery/subagent/resume-hint segments since,
and the wrapper broke the direct-call test seam. The load-bearing fix is
the stable usage reference: unchanged deltas no longer mint fresh objects,
so $uiState subscribers stop re-rendering per streaming event
- Regression tests: unchanged-reference retention, active_subagents-only
update, key-union asymmetry
The status bar flickers visibly during streaming because every state
patch (thinking.delta, reasoning.delta, tool.*, usage notifications)
creates a new $uiState object, forcing StatusRulePane and StatusRule
to re-render and redo expensive layout calculations on every event.
Three fixes applied:
1. Stabilize usage object references in createGatewayEventHandler.ts
- Add mergeUsageStable() that shallow-compares Usage fields before
creating a new object. When values haven't changed, returns the
existing reference, preventing unnecessary StatusRule re-renders.
2. Memoize expensive computations inside StatusRule (appChrome.tsx)
- statusBarSegments(cols) → useMemo([cols])
- modelLabel() → useMemo([model, effort, fast])
- ctxLabel, bar → useMemo([usage fields, segs])
- Tail segment budget + fits() calculations → single useMemo block
covering all progressive-disclosure logic
3. Wrap StatusRule in React.memo (appChrome.tsx)
- Combined with stable usage references, allows React to skip
re-renders when props haven't actually changed.
Fixes#41480
Normal prompt turns bind session['profile_home'] via set_hermes_home_override
before run_conversation, but the two ephemeral RPC paths (prompt.background,
preview.restart) spawn a fresh AIAgent on a new thread where the HERMES_HOME
ContextVar does not propagate — so a background/preview turn under a
non-default profile ran against the wrong home. Re-bind for the duration of
the ephemeral turn and restore in finally, mirroring the normal prompt turn.
Surgically reapplied from PR #50777 (handlers moved to methods_prompt.py
since the PR was authored; handler bodies rebind onto server.py globals, so
the original pattern transplants verbatim). Includes the contributor's
regression tests unchanged.
The resume follow-scroll fix subscribes term.onScroll and reads
term.buffer.active; the ChatPage test fake predates both and crashed
the suite with 'term.onScroll is not a function' (4 unhandled errors).
Follow-up to the salvaged #59713 commits: main's ChatPage gained the
PTY resume sanitizer + hydration machinery after the PR was authored, so
the write-callback follow-scroll is applied to the sanitized write path
(term.write(rendered, followScroll)) rather than the raw ev.data writes
the original diff targeted. Also adds the contributor email mapping.
Extract the resume-scroll decision into lib/pty-scroll (isViewportPinnedToBottom, shouldFollowPtyOutput) and add focused vitest coverage, per review on #59591.
Two tests asserted resolve+reload events but acquired the lease
immediately (no wait). After gating the reload behind _lease_waited,
these tests must simulate a contended wait via on_wait(0.0) to
exercise the resolve+reload path.
Three follow-up fixes to the cross-process turn lease:
1. Move _clear_durable_turn_lease_interrupt() to after the refresher
thread join in the outer finally. A refresher firing between the
inner stop and the join could set an interrupt that survives the
clear and poisons the next turn on a cached agent.
2. Set self._interrupt_message in the except branch of _interrupt_turn
so _clear_durable_turn_lease_interrupt can match and clear it even
when self.interrupt() itself raised.
3. Gate the conversation_history reload behind _lease_waited so an
immediate acquisition (no contention) does not replace the
in-memory history and cause an unnecessary prompt cache miss.
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
Presence-only _delegate_from/_branched_from checks stopped the lease walk on
continuations that copied a delegate's model_config, so the first refresh
after rotation missed the parent-key lease and hard-interrupted. A failed
get_session probe also skipped acquire entirely. Walk the lineage inside
the write transaction and treat a probe error as contended, not a fresh
session.
Honor interrupts while waiting for admission, stop the turn when refresh
loses the lease, poll once per second under contention, and test dead-PID
reclaim.