Expand user-guide/multi-connection-desktop.md into a complete setup
walkthrough: where to find the pane (settings nav, profile-rail plug,
command palette), the exact add-connection editor fields (Name,
Gateway URL, Authentication: Session token/OAuth, SSH host), Primary /
This device pills, Test semantics, agent roster + profile-rail
switching and per-profile session/cron/messaging scoping, token
storage via Electron safeStorage with the keyring-less Linux plain-
text opt-in, and troubleshooting. All quoted labels match the desktop
i18n strings. Cross-link the rail entry point from desktop.md.
The multi-connection registry (Settings -> Connections) shipped with no
entry point outside the settings nav, and the product-owner report was
blunt: 'I didn't see any obvious way to hook up multiple gateways.'
- Profile rail: a plug pill pinned beside Manage ('Connect another
Hermes gateway...') deep-links to /settings?tab=connections. Always
visible, including for single-profile first-run users.
- Command palette: Settings -> Connections is now a searchable entry
(keywords: add gateway, remote, ssh, cloud, instances, registry).
- i18n: profiles.connectGateway added to en/types/zh; other locales
fall back through defineLocale.
- Tests: profile-rail-connect.test.tsx covers the deep link and the
single-profile visibility guarantee.
The Nous Subscription image_gen row carries imagegen_backend="fal", so
_reconfigure_provider's post-model-picker step ran
img_cfg["use_gateway"] = False
unconditionally at two sites, immediately after the managed branch had
written use_gateway=True. A user who picked Nous Subscription and then
re-entered `hermes tools` to change the model was silently flipped onto
their personal FAL_KEY.
Same bug class as fe63353cb, which fixed the plugin-provider selector
but missed these two legacy-backend sites in the reconfigure flow. Both
now write bool(managed_feature), matching the existing correct site in
_configure_provider.
Adds regression tests driven through the real TOOL_CATEGORIES managed
row; sabotage-verified (tests fail with the old behavior restored).
Consult messaging and cron caches before the by-id fallback so opening a sidebar row neither depends on a redundant network lookup nor duplicates it into regular recents.
Co-authored-by: protas-box <protas.box@icloud.com>
Key resolved platform totals by Desktop profile and source so profile switches neither inherit another profile's count nor discard a count that was already resolved. Keep the full reset for connection configuration changes.
Co-authored-by: frendo <frendo.wu@gmail.com>
Complete the sidebar profile-scope contract across remote Electron routing, older-backend fallbacks, standalone messaging refreshes, and pagination. Reject stale profile responses and keep the explicit all-profiles view unified.
Co-authored-by: 墨綠BG <s5460703@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
MCP server problems (expired OAuth tokens especially) were only
discovered when the user visited the MCP page and a probe ran. Now a
renderer-side background checker (store/mcp-health.ts) sweeps the
active profile's enabled HTTP/SSE MCP servers on gateway connect and
every 30 minutes, and fires an in-app notification with a "Sign in"
action ("<name> MCP needs re-authentication") that navigates to the
MCP page with ?server=<name> so useDeepLinkHighlight focuses the
server and its Authenticate button. Navigation only — OAuth flows are
never auto-launched.
stdio servers are deliberately excluded: probing a stdio server SPAWNS
a local process, so a background timer must never touch them. Only
url-shaped servers (where OAuth expiry lives) are swept, sequentially.
The tab's probeCache/serverFingerprint/probeKey/NEEDS_AUTH_RE moved to
a shared lib/mcp-probe-cache.ts (behavior identical) so the page and
the checker share one probe cache and its 5-minute TTL — neither
surface re-probes what the other just learned.
Notifications fire only on a TRANSITION into needs-auth/error (pure
state machine, unit-tested), hard-capped at one per server per app
session, keyed per profile. Profile switches drop pending timers and
re-arm for the new profile; sweeps never run while the gateway is
disconnected. No new config knobs. i18n keys added across
en/zh/zh-hant/ja/ar + types.
Adds hermes://mcp/install?name=NAME&config=B64 (base64url or standard
base64 JSON), mirroring Cursor's mcp/install deep link, so vendors and
docs can offer an "Add to Hermes" button.
- Electron: the existing generic hermes:// handler already forwards
{kind, name, params}; only its comment is updated (no new handler).
- Renderer: use-desktop-integrations routes kind=mcp/name=install into
a pending-install store; a new confirmation dialog shows the server
name and the FULL pretty-printed config (attacker-controllable input),
with a prominent caution for stdio command entries. Nothing is written
until the user confirms; existing names require a rename or cancel.
On confirm the server is merged over a fresh fetch of the current map
via saveMcpServers, then navigation lands on /skills?tab=mcp&server=…
so useDeepLinkHighlight focuses the new row.
- Validation: name ^[A-Za-z0-9._-]{1,64}$; config must decode to an
object with a string http(s) `url` or a string `command` (never both);
payloads over 32KB rejected; failures surface as a toast.
- Pure parser in src/lib/mcp-deeplink.ts with unit tests (url shape,
command shape, bad base64, non-object, javascript: URL, oversized).
- i18n keys in types + en/zh/zh-hant/ja/ar.
- Docs: "Add to Hermes link" section in the MCP config reference.
Desktop/TUI count full displayed lineage after compression, but
prompt.submit validated truncate ordinals against tip-only history.
Translate via display_history_prefix and recover stale 4018s on Desktop.
Co-authored-by: Cursor <cursoragent@cursor.com>
Each configured server row on the MCP Capabilities page now shows what it
costs and whether it earns its keep:
- ~per-call token estimate of the server's tool schemas, summed over ENABLED
tools only (ceil(schema_chars/4) via the existing include/exclude filter)
- 30-day usage count from getUsageAnalytics(30), cached per scope profile
like the Toolsets tab's toolCallsCache, mapped to servers via the
mcp__<server>__<tool> registry-name convention (tools/mcp_tool.py)
- a subtle muted "unused" pill on enabled, probed-ok servers with nonzero
schema cost and zero 30-day uses — never a dialog
Backend: the /api/mcp/servers/{name}/test probe now fills an additive
per-tool `schema_chars` (length of the SAME converted registry schema the
agent registers). Older backends omit it → renderer shows counts only;
older renderers ignore the extra key. Display-only: nothing changes what
schemas are sent to models, no config knobs.
i18n keys (costTokens/usage30d/unusedPill) added to types/en/zh/zh-hant/ja
(ar inherits en via defineLocale overrides). Pure math lives in
lib/mcp-cost.ts with unit tests; Python wire shape pinned in
tests/hermes_cli/test_web_server_profile_unification.py.
Replace the fixed 500-message REST hydration (getLatestSessionMessages)
with a 120-row newest-first tail page. When the page comes back full, a
new per-session tail store records "possibly truncated + next offset";
"Show earlier" — once the DOM budget and the in-memory store window are
both exhausted — fetches the next older page via the new
getOlderSessionMessages helper (order latest + offset, matching the
backend's back-from-newest paging semantics) and prepends it to the
session store, deduped by durable row id and race-guarded against
session switches. Legacy backends without pagination metadata fall back
to the one-shot full transcript and retire the action.
Tail-page refreshes (background sync, post-turn rehydrate, re-activate,
cold-resume prefetch) graft the refreshed tail onto any backfilled
prefix instead of clobbering it, preserving reference identity on
no-ops. includeCompacted stays on every read — compaction-archived rows
remain part of the durable display history.
PowerShell 5.1 cold starts take 2.4-8s on affected Windows hosts, so the
shared 3s execText timeout hard-failed the parent start-marker probe for
any PID that still needs the PowerShell path (e.g. backend children).
Make execText's timeout overridable and raise the marker probe to 30s.
Fixes#87169
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.
Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.
Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.
Fixes#87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
Review follow-up. Three coverage gaps in the LiteLLM matrix:
- The operator opt-out on a litellm-named provider pointed at an OpenRouter
host. That route previously took the OpenRouter branch and ignored an
explicit per-model `prompt_caching: false`; it is the only cell in the
differential matrix where the salvage REMOVES caching, so pin it as
intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
lookalike host, because the provider id is an independent signal and only
the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
which the token matcher handles but nothing exercised.
Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.
Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.
Self-review follow-up, caught by benchmarking the previous commit.
Widening the custom-provider capability-lookup gate to `is_anthropic_wire or
_is_litellm_route(...)` made EVERY chat_completions route with a litellm-ish
provider/host enter the lookup, including non-Claude models that the grant
branch below can never match. Measured on a route with no config.yaml
(the uncached worst case) that was ~7.5us -> ~1528us per evaluation.
Narrowed the gate to the exact condition the LiteLLM branch grants on
(chat_completions + Claude + litellm route), computed once into a local and
reused by the branch itself so the predicate no longer runs twice.
Measured with a realistic config.yaml present (mtime cache warm), vs
origin/main:
live-agent policy 20.6us -> 61.7us
destination planning 219.3us -> 347.7us
Sub-millisecond and scoped to the routes that actually opted in. The
earlier 1.5ms figures were a tempdir artifact: load_config_readonly's
mtime cache cannot engage when no config.yaml exists, which is never true
of a real install. Non-LiteLLM and non-Claude routes are unaffected
(openrouter Claude measured flat at ~7.9us).
Tests: 82 passed across the policy and TTL-propagation modules.
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.
Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.
Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.
1. Layout. The branch returned the native inner-block layout
(use_native_layout=True) on api_mode == "chat_completions". That layout
writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
messages and depends on the Anthropic adapter to relocate it into the
block — but that adapter only runs for api_mode == "anthropic_messages"
(agent/transports/anthropic.py registers there), and the
chat_completions transport does no relocation. Measured on a 3-tool-turn
transcript: 2 of the 4 available breakpoints landed on markers the
provider never sees. Worse, when LiteLLM itself relocates a top-level
marker for an OpenRouter-backed Claude route
(OpenrouterConfig._move_cache_control_to_content), the marker lands on
an empty assistant turn and produces a cache_control-marked empty text
block — the HTTP 400 "text content blocks must contain" shape already
guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
layout, matching every other OpenAI-wire grant in this function:
4 of 4 breakpoints honored, zero empty blocks.
2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
false-positive class base_url_hostname's own docstring warns against; it
granted Anthropic markers to notlitellm.example.com,
foolitellmbar.example and friends. Replaced with a label-token match in
a named helper, so "litellm" must be a whole dot- or hyphen-delimited
token. All three of the original test hosts still match; a "litellm"
path segment on an unrelated host still does not.
3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
bedrock_converse and codex_app_server. Gated on
api_mode == "chat_completions" explicitly.
4. Operator override. The grant is inferred from a provider/host name, but
the custom-provider capability lookup was gated on is_anthropic_wire, so
an explicit `prompt_caching: false` for the route+model was honored on
/v1/messages and silently ignored on /v1/chat/completions. The lookup
now also runs for a LiteLLM route, and its layout follows the transport
rather than the declaration (an explicit `true` must not promote a
chat_completions request to the native layout).
Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.
Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.
Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
through the same proxy must not receive markers (they may reject the
cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
varies per install (litellm, custom:litellm, or a bare custom alias
pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
excluded -- verified by the existing over-reach regression test.
Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.
Closes#84506. Original diagnosis, patch design, and measurements by
@ottosulin.
The composer's global Esc-to-cancel listener (useComposerEscCancel) fires
whenever the turn is busy and the active composer matches — but overlays
(Settings, Command Center, agents, cron, …) cover the chat while the
composer stays mounted and 'active' beneath them, so pressing Esc on any
of those pages interrupted the session the user wasn't even looking at.
OverlayView's own escape-layer Esc-to-close fired too, but the stream was
already dead.
Stand Esc down with composerFocusBlockedBySurface() — the same signal the
type-to-focus path uses (BLOCKING_OVERLAY includes OverlayView's
[data-overlay-surface] marker). Esc on an overlay now closes the overlay
via its escape layer instead of canceling the stream beneath it.
Fixes#82618
With gateway.multiplex_profiles enabled, the primary Telegram message
handler is the closure returned by _make_default_profile_message_handler(),
so its __self__ is absent. The early intake filter
(_is_user_authorized_from_message) recovered the GatewayRunner via
self._message_handler.__self__ and, finding none, fell back to env-only
authorization — never evaluating the configured chat allowlist through
GatewayRunner._is_user_authorized(). Every non-global sender was then
default-denied in an explicitly allowlisted group.
Prefer the platform-bound authorization callback registered via
set_authorization_check(): it routes through the runner's full auth chain
(platform + group allowlists, pairing store, allow-all) and survives the
closure wrapping, whereas the bound-handler lookup does not. The bound
handler remains the fallback for setups without a registered callback, and
the pairing-passthrough guard for unknown DMs is preserved.
Fixes#87132
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
runuser/su/sudo -u from a root shell leaks XDG_RUNTIME_DIR=/run/user/0 into
the child. _user_systemd_socket_ready() stat-ed sockets under it with a bare
Path.exists(), which only suppresses ENOENT/ENOTDIR/EBADF/ELOOP — EACCES on
the 0700 root-owned dir escaped as a raw PermissionError traceback instead of
the documented UserSystemdUnavailableError remediation path.
- _path_exists_safe(): Path.exists() that treats EACCES as absent; used at
both the readiness and DBUS-detection call sites.
- _ensure_user_systemd_env(): drop an XDG_RUNTIME_DIR that is unset or owned
by another user in favour of our own /run/user/{uid}, so the restart
actually succeeds after su/sudo -u instead of only failing cleanly.
Regression tests cover the EACCES readiness probe, foreign-dir replacement,
and that preflight raises UserSystemdUnavailableError (not PermissionError).
Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.
Refs #83915
session.workspace.move refused a running live session with 4009
(session busy), but the desktop's Move-to-project flow calls exactly
this RPC — so the UI updated its local grouping while state.db kept the
old cwd and the agent's tools kept running in the old workspace. Two
sources of truth disagreed (#86626).
An explicit move now wins: the stored row and the live session re-anchor
together. In-flight tool calls keep the cwd they were launched with; the
next tool call uses the new workspace.