The Nous Subscription image_gen row carries imagegen_backend="fal", so
_reconfigure_provider's post-model-picker step ran
img_cfg["use_gateway"] = False
unconditionally at two sites, immediately after the managed branch had
written use_gateway=True. A user who picked Nous Subscription and then
re-entered `hermes tools` to change the model was silently flipped onto
their personal FAL_KEY.
Same bug class as fe63353cb, which fixed the plugin-provider selector
but missed these two legacy-backend sites in the reconfigure flow. Both
now write bool(managed_feature), matching the existing correct site in
_configure_provider.
Adds regression tests driven through the real TOOL_CATEGORIES managed
row; sabotage-verified (tests fail with the old behavior restored).
Desktop/TUI count full displayed lineage after compression, but
prompt.submit validated truncate ordinals against tip-only history.
Translate via display_history_prefix and recover stale 4018s on Desktop.
Co-authored-by: Cursor <cursoragent@cursor.com>
Each configured server row on the MCP Capabilities page now shows what it
costs and whether it earns its keep:
- ~per-call token estimate of the server's tool schemas, summed over ENABLED
tools only (ceil(schema_chars/4) via the existing include/exclude filter)
- 30-day usage count from getUsageAnalytics(30), cached per scope profile
like the Toolsets tab's toolCallsCache, mapped to servers via the
mcp__<server>__<tool> registry-name convention (tools/mcp_tool.py)
- a subtle muted "unused" pill on enabled, probed-ok servers with nonzero
schema cost and zero 30-day uses — never a dialog
Backend: the /api/mcp/servers/{name}/test probe now fills an additive
per-tool `schema_chars` (length of the SAME converted registry schema the
agent registers). Older backends omit it → renderer shows counts only;
older renderers ignore the extra key. Display-only: nothing changes what
schemas are sent to models, no config knobs.
i18n keys (costTokens/usage30d/unusedPill) added to types/en/zh/zh-hant/ja
(ar inherits en via defineLocale overrides). Pure math lives in
lib/mcp-cost.ts with unit tests; Python wire shape pinned in
tests/hermes_cli/test_web_server_profile_unification.py.
When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.
Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.
Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.
Fixes#87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).
Review follow-up. Three coverage gaps in the LiteLLM matrix:
- The operator opt-out on a litellm-named provider pointed at an OpenRouter
host. That route previously took the OpenRouter branch and ignored an
explicit per-model `prompt_caching: false`; it is the only cell in the
differential matrix where the salvage REMOVES caching, so pin it as
intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
lookalike host, because the provider id is an independent signal and only
the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
which the token matcher handles but nothing exercised.
Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.
Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.
Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.
Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.
1. Layout. The branch returned the native inner-block layout
(use_native_layout=True) on api_mode == "chat_completions". That layout
writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
messages and depends on the Anthropic adapter to relocate it into the
block — but that adapter only runs for api_mode == "anthropic_messages"
(agent/transports/anthropic.py registers there), and the
chat_completions transport does no relocation. Measured on a 3-tool-turn
transcript: 2 of the 4 available breakpoints landed on markers the
provider never sees. Worse, when LiteLLM itself relocates a top-level
marker for an OpenRouter-backed Claude route
(OpenrouterConfig._move_cache_control_to_content), the marker lands on
an empty assistant turn and produces a cache_control-marked empty text
block — the HTTP 400 "text content blocks must contain" shape already
guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
layout, matching every other OpenAI-wire grant in this function:
4 of 4 breakpoints honored, zero empty blocks.
2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
false-positive class base_url_hostname's own docstring warns against; it
granted Anthropic markers to notlitellm.example.com,
foolitellmbar.example and friends. Replaced with a label-token match in
a named helper, so "litellm" must be a whole dot- or hyphen-delimited
token. All three of the original test hosts still match; a "litellm"
path segment on an unrelated host still does not.
3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
bedrock_converse and codex_app_server. Gated on
api_mode == "chat_completions" explicitly.
4. Operator override. The grant is inferred from a provider/host name, but
the custom-provider capability lookup was gated on is_anthropic_wire, so
an explicit `prompt_caching: false` for the route+model was honored on
/v1/messages and silently ignored on /v1/chat/completions. The lookup
now also runs for a LiteLLM route, and its layout follows the transport
rather than the declaration (an explicit `true` must not promote a
chat_completions request to the native layout).
Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.
Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.
Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
through the same proxy must not receive markers (they may reject the
cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
varies per install (litellm, custom:litellm, or a bare custom alias
pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
excluded -- verified by the existing over-reach regression test.
Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.
Closes#84506. Original diagnosis, patch design, and measurements by
@ottosulin.
With gateway.multiplex_profiles enabled, the primary Telegram message
handler is the closure returned by _make_default_profile_message_handler(),
so its __self__ is absent. The early intake filter
(_is_user_authorized_from_message) recovered the GatewayRunner via
self._message_handler.__self__ and, finding none, fell back to env-only
authorization — never evaluating the configured chat allowlist through
GatewayRunner._is_user_authorized(). Every non-global sender was then
default-denied in an explicitly allowlisted group.
Prefer the platform-bound authorization callback registered via
set_authorization_check(): it routes through the runner's full auth chain
(platform + group allowlists, pairing store, allow-all) and survives the
closure wrapping, whereas the bound-handler lookup does not. The bound
handler remains the fallback for setups without a registered callback, and
the pairing-passthrough guard for unknown DMs is preserved.
Fixes#87132
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
runuser/su/sudo -u from a root shell leaks XDG_RUNTIME_DIR=/run/user/0 into
the child. _user_systemd_socket_ready() stat-ed sockets under it with a bare
Path.exists(), which only suppresses ENOENT/ENOTDIR/EBADF/ELOOP — EACCES on
the 0700 root-owned dir escaped as a raw PermissionError traceback instead of
the documented UserSystemdUnavailableError remediation path.
- _path_exists_safe(): Path.exists() that treats EACCES as absent; used at
both the readiness and DBUS-detection call sites.
- _ensure_user_systemd_env(): drop an XDG_RUNTIME_DIR that is unset or owned
by another user in favour of our own /run/user/{uid}, so the restart
actually succeeds after su/sudo -u instead of only failing cleanly.
Regression tests cover the EACCES readiness probe, foreign-dir replacement,
and that preflight raises UserSystemdUnavailableError (not PermissionError).
Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.
Refs #83915
session.workspace.move refused a running live session with 4009
(session busy), but the desktop's Move-to-project flow calls exactly
this RPC — so the UI updated its local grouping while state.db kept the
old cwd and the agent's tools kept running in the old workspace. Two
sources of truth disagreed (#86626).
An explicit move now wins: the stored row and the live session re-anchor
together. In-flight tool calls keep the cwd they were launched with; the
next tool call uses the new workspace.
The "Conversation started:" line carried a bare date (%A, %B %d, %Y). Tools
that accept instants -- nutrition, calendar and similar MCP servers -- reject
naive datetimes and require an explicit UTC offset, so the model had to infer
EST vs EDT from the date alone. Near a DST boundary that is a coin flip, and a
wrong guess does not error: it silently writes the record onto the wrong day.
Append the IANA zone (when configured), the zone abbreviation and the UTC
offset, e.g.:
Conversation started: Saturday, August 15, 2026 (America/New_York, EDT, UTC-04:00)
get_timezone() returns None when no timezone is configured; in that case the
line falls back to the abbreviation and offset of the server-local (still
tz-aware) time, so behaviour is unchanged for users who never set one:
Conversation started: Saturday, August 15, 2026 (EDT, UTC-04:00)
Daily byte-stability is preserved -- the property the date-only format exists
to protect (PR #20451). Zone name, abbreviation and offset are all constant for
the whole day; they shift only at a DST transition, where a change is correct.
The static-prefix reconstruction guard in _restore_plugin_sections matches on
"\n\nConversation started:" and is unaffected by a suffix after the date.
test_datetime_is_date_only_not_minute_precision used `re.search(r":\d{2}")`
over the whole line as a proxy for "no time-of-day". A UTC offset also matches
that pattern, so the check now applies to the date portion (everything before
the zone parenthetical) and the invariant is tightened rather than relaxed:
- test_datetime_includes_utc_offset asserts the offset is present
- test_datetime_line_is_stable_across_rebuilds asserts two rebuilds in the
same day produce a byte-identical line
Fixes#87403
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a bounded fallback path in turn_finalizer.py that always records
a terminal timed_out outcome via _record_task_failure (CAS receipt path)
when the iteration budget is exhausted, regardless of whether the normal
fallback paths (interrupted/failed/anomalous exit_reason) were eligible.
Previously, a kanban worker whose budget was exhausted but whose turn was
interrupted, failed, or exited with an anomalous reason would silently
leave its task in an ambiguous lifecycle state — the dispatcher would
eventually detect it as a crashed or protocol-violation worker, but the
failure was not bounded and could take a full tick cycle to reconcile.
The CAS invariant in _end_run (WHERE ended_at IS NULL) guarantees
idempotence: if another path already closed the run, the call is a no-op.
Extracted the inline kanban-budget-exhausted recording into a shared
helper function (_record_kanban_budget_exhausted) used by both the
existing iteration_limit_fallback path and the new bounded fallback path.
Closes#87096
hermes update decides restart ownership from the live grandchild argv,
not from an env marker. Newly generated plists now include the flag.
The stderr_timestamp wrapper upgrades only historical Hermes gateway
run shapes for stale plists and leaves arbitrary launchd children unmarked.
launchd only stamps XPC_SERVICE_NAME on its direct child. The timestamp
wrapper is that child, so the grandchild gateway sees XPC_SERVICE_NAME=0
and the supervised-conflict guard refuses the service's own spawn.
Forward HERMES_GATEWAY_EXTERNAL_SUPERVISOR=1 when the wrapper itself is
launchd-supervised. Interactive XPC_SERVICE_NAME=0 starts stay unmarked.
Fixes#86893
After the VBS/cmd launcher exits, Task Scheduler marks
Hermes_Gateway_* Ready while the detached gateway keeps running.
The orphan reaper only bailed on Running, then fail-opened the
parent-chain check and killed the live bot on desktop serve start.
Fixes#87001
hermes chat -q sets HERMES_INTERACTIVE=1 (for interactive sudo prompts) but
runs one turn with no user waiting to answer approval prompts. Previously a
dangerous command triggered the interactive gate, waited the full 300s
timeout, then failed closed — and the agent was effectively forced to work
around the block, often silently auto-approving via execute_code (which
auto-approves in non-gateway mode).
Add approvals.single_query_mode (default deny, mirror of cron_mode):
deny — block dangerous commands and execute_code deterministically with
a clear 'no user present' message (no 300s wait)
approve — auto-approve dangerous commands/execute_code in -q mode
cli.py marks the session with HERMES_SINGLE_QUERY_SESSION; the shared gate
(_run_approval_gate, check_all_command_guards, check_execute_code_guard)
treats -q as a deterministic non-interactive context when that marker is set.
execute_code, the -q escape hatch, now honors single_query_mode instead of
auto-approving headlessly. Includes tirith parity in the combined guard and
docs. Fixes#86878.
An MCP server exposing a native tool named read_resource (or
list_resources/list_prompts/get_prompt) collided with the auto-generated
resource/prompt utility of the same name. The registration collision
handler flagged the pair as ambiguous and skipped BOTH entries, so the
server's own tool became silently unavailable on every gateway boot.
Resolve this specific native-vs-utility collision in favour of the native
tool: keep it and drop the shadowed utility, which is only convenience
sugar for servers that expose no such tool of their own. The conservative
skip-everything path still applies to genuinely ambiguous collisions (two
or more native tools normalizing to one name), which we cannot
disambiguate. Add a regression test covering the native-tool-wins path.
Fixes#87112
A named --provider without -m used to swallow lookup failures and
silently keep the global model.default. Log the resolution error so
the fallback is visible.
hermes chat --provider <name> without -m sent the global model.default
to the custom endpoint. Named custom entries already expose
default_model via _get_named_custom_provider(); honor that when the
user selected the provider and did not pass an explicit model.
Fixes#86978
The fire-and-forget handler only caught asyncio.CancelledError and
Exception, so a SystemExit/KeyboardInterrupt escaping a turn (e.g. a
plugin calling sys.exit() in a tool call or summary-LLM path) skipped
the user-facing failure notification and surfaced only as 'Task
exception was never retrieved' — radio silence for the user.
Catch BaseException instead; send the failure notification first, then
re-raise SystemExit/KeyboardInterrupt to preserve shutdown semantics
for the loop's own signal handling. Other BaseExceptions stay contained.
Closes#86651