Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).
LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.
- `_status_429` now honors a structured billing code first (decisive
signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
(429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
limit" was already there; providers use both orders). "terminal billing
limit" text is deliberately NOT matched: substring rules cannot negate
the "non-terminal billing limit" wording — the structured code covers it.
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.
- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
after the read loop; surface EOF-without-[DONE] (warn + error result
when nothing was received); extract _consume_sse_line so line parsing
and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
proven red on origin/main.
Gemini models enable internal thinking/reasoning tokens by default.
When generate_title() called call_llm() with max_tokens=64, Gemini
consumed the entire 64-token budget on internal thought tokens,
truncating the JSON title response before it could complete. The
fallback prose extractor then picked up the opening fence (```json)
or a bare brace as the session title.
Two-part fix:
1. title_generator.py: Pass reasoning_config={"enabled": False} to
call_llm() so thinking is explicitly disabled for title generation.
2. chat_completions.py: In _build_gemini_thinking_config, when
reasoning is disabled (enabled=False or effort="none"), set
thinkingBudget: 0 on Gemini models that support it (2.5+ and 3.x).
includeThoughts: False only hides thought parts from the response
while the model still reasons internally and bills thought tokens
against maxOutputTokens. thinkingBudget: 0 truly disables thinking
so thought tokens do not consume the max_tokens budget.
Fixes#91927
Port from can1357/oh-my-pi#10521: their loop guard only hashed single-call
turns, so a model replaying the same multi-call batch every iteration was
never counted; they widened the hash to the whole batch. Hermes has the
same blind spot in a different shape: observe_call tracks a CONSECUTIVE
identical-call streak, so an A,B,A,B,... cycle of identical (args, result)
pairs resets the streak on every alternation and runs to the iteration
budget unflagged (live-reproduced: 60 calls in a 2-cycle, zero notices,
no hard stop).
Add a period-2..4 cycle detector over a bounded per-turn call history:
notice on the STALL_GUARD_IDENTICAL_CALL_THRESHOLD-th identical lap,
hard stop at no_progress_block_after laps under hard_stop_enabled — the
same thresholds the period-1 streak uses. Cycles whose results change
between laps never fire (real progress); cycles made only of poller-exempt
tools are exempt (legitimate waiting), matching single-call semantics.
Widen the streak-stop propagation seam in run_agent.py to carry the new
decision code.
Review finding: the salvaged rule assumed a lowercase-hex grammar that
AgentMail's docs do not establish (only the `am_` / `am_org_` prefix is
documented), so a non-hex key would have gone unmasked. Discriminate on
what actually separates keys from identifiers: an alphanumeric body with
no `_`/`-` and a 20-char floor. `am_example_identifier_123` still passes.
Review finding: `echo cd backend` injected backend/AGENTS.md, and the
rstrip(";") applied after shlex had removed quoting turned `cd 'backend;'`
into `backend`. Tokenize with punctuation_chars so operators are their own
tokens: a `cd` counts only at a segment start, `backend;ls` splits at the
operator, and a quoted `'backend;'` stays the literal name.
SubdirectoryHintTracker's generic token filter in
_extract_paths_from_command drops any token that does not contain /
or ., so relative directory names in commands like 'cd backend && ls'
were silently ignored. As a result, Hermes missed AGENTS.md /
CLAUDE.md / .cursorrules in the entered subdirectory whenever users
navigated with plain relative names -- the common case.
Add _extract_nav_command_targets, which scans the token stream for
'cd' / 'pushd' and treats the next non-flag token as a path candidate
resolved against working_dir. The generic token pass still runs so
all other shapes (absolute paths, files with extensions, etc.) keep
working. 'cd -' and bare 'cd' are intentionally skipped -- neither
points at a project subdirectory.
Three new tests cover 'cd backend && ls', 'pushd backend', and
'cd backend' appearing after an earlier chained sub-command.
Known limitation called out in review: multi-step chains like
'cd backend && cd src' still resolve each hop against working_dir
rather than simulating the shell's evolving cwd. That's a bigger
change (shell state tracking) and is out of scope for this fix;
the common single-hop case reported in the issue is now covered.
Fixes#11032
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).
- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
request-validation and overflow heuristics, returning format_error with
`retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
verdicts that legitimately reach that branch (policy block, TLS chain, MoA
shape/preset errors) now state `should_fallback=True` explicitly, so the gate
changes behaviour only for the new verdict. Local validation errors keep
their historical fallback.
Fixes#12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.
Co-authored-by: cuyua9 <2114364329@qq.com>
`re.sub` with the home path as a template string parsed backslashes as
escapes (re.error dropped the whole injected config block); use a callable.
The early `~/…` return also skipped `expandvars`, leaving `~/$LEAF` half
resolved. Prefix-substitute and fall through to normal expansion instead.
The _SKILL_INVALID_CHARS regex stripped all non-ASCII characters from
skill names, so skills with CJK, Cyrillic, or other Unicode names
(e.g. "小说拆条") produced an empty slug and were silently skipped.
Change the regex from [^a-z0-9-] to [\w-] so Unicode word characters
are preserved. Platform-specific sanitizers (Telegram, Discord) already
handle their own character restrictions downstream.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
`skip_context_files=True` kept AGENTS.md/CLAUDE.md out of the system prompt,
but `SubdirectoryHintTracker` was always on, so the first tool call touching
a directory with such a file spliced its full text onto the tool result.
Cron jobs without a workdir (which set skip_context_files) that relay exact
stdout then delivered `[Subdirectory context discovered: ...]` plus the file
body to Telegram/Discord.
The tracker now takes `enabled=` and agent init wires it to
`not skip_context_files`: one flag, both injection paths. Interactive
sessions and cron jobs with a workdir are unchanged.
Direction proposed in PR #9434 (@zhitiao), which gated on platform == "cron";
gating on the existing skip flag covers the same case without a platform
special-case.
Fixes#9441
Gateway sessions and batch workers append to the same default
trajectory_samples.jsonl / failed_trajectories.jsonl with a plain open("a") +
write(); concurrent writers interleaved mid-object and the file stopped
parsing (#12684). The append now holds an exclusive lock for write+flush:
flock on POSIX, a 1-byte msvcrt.locking range on Windows.
Tests: a foreign process holding the lock must block the append (red on
base); six processes appending oversized entries all land parseable.
Fixes#12684. Salvaged from #12685 by @shafdev; Windows arm and test trim ours.
`build_tool_preview()`'s generic-key fallback and the cute-message helpers
still truncated with a bare `text[:max_len - 3] + "..."`; for max_len 1-3 the
slice goes negative and returns almost the whole string (27 chars for
max_len=1). `_truncate_preview` already had the guard, so the two code paths
disagreed.
One truncation helper (`_tail_trunc`) with the guard, used everywhere; the
head-truncating `_cute_path` gets the same clamp.
Salvage of PR #48483 by @HeLLGURD (current-code fix); the earliest reports and
patches were #9464 (@LarHope), #9477 (@kagura-agent) and #9497.
Co-authored-by: LarHope <12761142+LarHope@users.noreply.github.com>
Fixes#9439
`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.
Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.
Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.
Fixes#9814
`add_provider()` flipped `_has_external` and appended the provider before
calling `get_tool_schemas()`. When schema loading raised, the broken provider
stayed registered and the single-external slot was poisoned for the rest of
the process: every later provider was rejected as "already registered".
Materialize the schema list first; state changes only after it succeeds.
Exception propagation is unchanged.
Hand-port of PR #9997 by @zhouhe-xydt onto the current add_provider() (the
reserved-core-tool filter landed in between); one invariant test.
Fixes#9948
An unquoted redirect target breaks if the pytest tmp dir contains a
space: the hook then never writes its pid and the test waits out the
300 s hook instead of failing fast. The neighbouring tests in this file
are unquoted, so only the new test is changed here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_spawn puts the hook child in its own process group on POSIX, so the
terminal's SIGINT never reaches it and only _spawn's own cleanup can stop
it. That cleanup sat in `except Exception`, and KeyboardInterrupt is a
BaseException — Ctrl+C during a hook skipped kill_process_tree entirely
and left the hook (and anything it forked) running.
Catch BaseException, run the existing tree kill and drain, then re-raise
anything that is not an Exception subclass so KeyboardInterrupt and
SystemExit still reach the caller unchanged. Every other return path is
untouched.
Follow-up to #84901, which added the same cleanup for the timeout path.
The regression test mirrors the existing real-subprocess tests in
tests/agent/test_shell_hooks_tree_kill.py: it interrupts the process once
the hook is confirmed running, then asserts both that KeyboardInterrupt
propagates and that the hook pid is gone. Verified RED on the unfixed
tree (that test alone fails, the other five pass) and green with the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the #106769 pick: the 8-line dated bugfix narrative in
parse_context_limit_from_error becomes a 2-line WHY, and one invariant
test pins the contract: an output-cap message parses to None as a
context limit, 16384 as the available output tokens, and classifies as
an output-cap error, while a genuine "maximum context length" message
still parses. Red on origin/main (returned 16384 as the context limit).
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
Two LSP freshness bugs reported by @tobific (#108882, #108881):
- `_current_diags_async()` keyed the client lookup by the enclosing
workspace root while `_get_or_spawn()` stores single-root servers under
`srv.resolve_root(...)` (a nested package.json project). The lookup
returned [] for a live client with diagnostics, so the delta baseline was
refreshed from nothing. Use the same resolved-root key.
- `open_or_change()` published `_DocState.version` only after awaiting the
didChange write. A versionless publishDiagnostics read during that await
was credited with the OLD version and judged stale once the send resumed.
Bump the version before the send; a failed send (swallowed by
`_send_notification`) leaves a version nothing satisfies, i.e. "no
verdict", which is the existing contract.
The mock server gains a push-only `versionless` script so the race is
reproducible without a real language server.
boto3 and azure-identity freeze the credential chain into the client at
construction, and the process env under a multiplexed turn belongs to the launch
profile. A region-only (bedrock) / config-only (lru_cache) slot therefore signed a
served profile's calls with the launch profile's keys and served its account's
model list to everyone.
Under a HERMES_HOME override the clients are built from the profile's secret
scope (AWS_* / AZURE_* from its .env) and cached per (home, service, region) /
(home, config); the discovery cache key carries the home. The unscoped path keeps
the region slot and the maxsize=1 lru byte-for-byte.
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift
The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.
* refactor(desktop): one consent card for connectors and MCP setup
McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.
* feat(desktop): connector card drives the agent through manage_connections wait
The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.
Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.
* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them
The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.
* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate
A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.
* fix(agent): name a provider retry backoff on the live status line
The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.
* test(desktop): connector rehearsal launcher and flagged connector spec
connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.
* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout
The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.
* style(desktop): blank lines in connector-flow test per lint
* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film
The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.
* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds
The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.
* fix(desktop): no provider picker or free-tier chip over the guided first launch
Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.
* fix(desktop): onboarding connector picks are real catalog slugs
The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.
* fix(desktop): the free-tier ready screen never interrupts the guided chat
A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.
* feat(desktop): tour options that lead to building, and a fork that follows the tour
"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.
* feat(desktop): the onboarding connector picker reads the live catalog
A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.
* test(desktop): the guided first launch never forces a sign-in
The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).
* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape
Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.
* style(desktop): one answeredAfter helper for the ask card and first-build chips
* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows
Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.
* style(desktop): the 'nothing connects yet' line reads first on the connectors card
Extends the routing-variant fix to the sibling catalog paths, so the whole
bug class is covered rather than just context length (#97820).
_find_model_entry(), lookup_models_dev_context(), get_model_info(), and
get_model_capabilities() all keyed on the full suffixed id, so a routed id
such as z-ai/glm-5.3-flash:floor missed models.dev and fell through to the
generic "glm" family default (202,752 instead of 1,310,720).
The retry with the base id runs LAST — after exact and case-insensitive
matching — so a real catalog SKU always wins over its base.
:free and :batch are deliberately NOT stripped. They are real catalog SKUs,
not routing modifiers: OpenRouter's /models currently lists 18 :free and 65
:batch entries, and 13 of those carry a context window different from their
base (z-ai/glm-5.2:free is 256K vs the base's 1.05M). Stripping them would
report a window LARGER than the model has, so the request fails at the API
instead of merely compacting early — a worse failure than the under-report
this fixes. An absent SKU must also miss rather than inherit the base, so
model_overrides _default fill-gap semantics keep working.
The suffix set is shared from hermes_constants, so validation, context
resolution, and catalog lookup all agree on one definition.
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.
`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:
openai/gpt-5.5:nitro -> 256K (real 1.05M)
x-ai/grok-4.6:nitro -> 131K (generic "grok" catch-all)
anthropic/claude-opus-4.6:nitro-> 200K (generic "claude" catch-all)
The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.
Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.
`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.
The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
One multiplexed gateway process serves every profile, but several per-turn
reads still went through state frozen from the LAUNCH profile:
- `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the
module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS,
so every secondary ran with the default profile's turn budget. A routed turn
(HERMES_HOME override) now resolves `agent.max_turns` from its own config.
- `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide
slot, so secondaries fell back through the default's provider/model with their
own keys. It now reads the active gateway home and keeps a last-known-good
chain per home.
- `_load_prefill_messages` resolved relative paths against the launch home.
- `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a
secondary's compression/title/vision calls authenticated to Nous with the
default profile's token when it had no pool entry. Resolved per call via
`hermes_cli.auth._auth_file_path()` (patched constant still wins in tests).
- `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was
loaded outside any profile scope, so secondaries' `hooks/` never ran and the
default profile's handlers received every profile's messages, responses and
user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the
runner holds one registry per served home, picked from the active scope at
emit time and front-loaded under each secondary's startup scope.
- Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME
and the default profile's secrets). They now get the routed HERMES_HOME via
`build_subprocess_env`, scrubbed under multiplexing, and the stdin payload
carries `profile` so one script can tell which profile fired it.
- Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`,
`trust_recent_files*`) was bridged once into env at startup and read from env
per delivery; under a HERMES_HOME override the validator now reads the routed
profile's config. Single-profile runs keep the env-bridge contract.
Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro
(temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7,
fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values.
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.
Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.
Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.
Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.
Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.
Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.
Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.
session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.
Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.
Fixes#82903Fixes#65941Fixes#99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.
Two catalog gaps still produced the reported symptoms:
- A leftover context_length_cache.yaml entry of 128K (the old
``deepseek`` catch-all) outlived the 1M catalog keys because
deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
image routing falls through to lossy text. Vendor docs (2026-09-10)
mark deepseek-flash as vision-capable and deepseek-v4-pro as not.
Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.
Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
bedrock.guardrail was only attached on the Converse route (guardrailConfig in the
body). Claude on Bedrock goes through the AnthropicBedrock SDK, i.e. InvokeModel,
whose body has no guardrailConfig, so the default Claude route ran with no guardrail
at all (#52179; live-verified by JiaDe-Wu: the blocked word came back through Hermes).
Bedrock reads the guardrail for InvokeModel from X-Amzn-Bedrock-GuardrailIdentifier /
-GuardrailVersion / -Trace headers. Attach them as default_headers in
build_anthropic_bedrock_client so every AnthropicBedrock client Hermes builds
(primary init, /model switch, fallback, per-request rebuild, auxiliary) enforces the
same guardrail, with prompt caching / thinking / 1M context kept (the reason Claude is
not routed through Converse).
InvokeModel blocks do NOT change stop_reason (stays end_turn) and return the guardrail's
canned text as an ordinary assistant reply, flagged only by
amazon-bedrock-guardrailAction=INTERVENED in the body (SDK: response.model_extra).
AnthropicTransport.response_finish_reason maps that to content_filter so the loop runs
its refusal handling instead of reasoning over the canned text; _derive_finish_reason
uses it for the anthropic_messages branch.
Mantle (openai.gpt-5.x) is documented by AWS as not supporting Guardrails on the
Responses endpoint; the docs now say so instead of promising "all model invocations".
Header mechanism proposed in #52312 by @JoaoMarcos44 (stale base, 7-file conflict,
detection keyed on a Converse-only stopReason); reimplemented on current main.
Live probe (local sink, SigV4 fake creds): before, no X-Amzn-Bedrock-* header on the
InvokeModel request; after, headers present, SigV4 intact, INTERVENED → content_filter.
The SDK-private _build_headers probe was spelled three times across two files; the api-key shard also re-asserted the bearer invariant the portal test owns. dict(headers or {}) → dict(headers): every caller passes _beta_header's dict.
The bearer guard is now an Omit() default header (copy-safe), so the SDK
attribute may still hold the env key while nothing reaches the wire. Assert
on _build_headers for the client and a with_options() copy instead.
The api-key branch now uses an Omit() default header because an attribute
clear does not survive with_options(); the mirror bearer branch still relied
on `client.api_key = None`. Probe on anthropic 0.87.0: a with_options() copy
of a bearer client re-read ANTHROPIC_API_KEY and sent x-api-key alongside
the Bearer (#26970's dual-auth leak, on copies). Both branches now set one
Omit() header, built once from the beta headers. Wire test asserts the
Authorization header is absent outright, not merely sentinel-free.
The pinned anthropic==0.87.0 exports `Omit` at package level, so the
import-probe plus httpx header-stripping fallback added for "SDK too old"
never runs; drop both and set the default header directly. Tests trimmed
to two invariants: the header-builder contract on the client and a
with_options() copy, and the wire-level end-to-end through
build_anthropic_client. The bearer-style mirror case stays covered.
The auth-boundary group (env bearer isolation, fail-closed Omit
fallback, wire-level header captures) moves to
test_anthropic_client_auth.py: the main adapter suite was pushed past
the 2,000-line ceiling by the group, and the shard keeps every test —
including both local HTTP round-trips — with explicit imports and a
shared capturing-server fixture. Both suites run together; coverage
is unchanged.
The api-key credential-isolation guard in this PR only removes the
env-derived Bearer when `anthropic._types.Omit` can be imported. When the
sentinel is unavailable (older/exotic SDK layout), `_new_sdk_client` fell
through to a bare api-key client whose SDK env fallback re-reads
ANTHROPIC_AUTH_TOKEN and ships `Authorization: Bearer ***` to third-party
endpoints — the compatibility path failed open (P2, reported by @egilewski).
Fail closed instead: when Omit is unavailable, build the client with a
copy-safe httpx request hook that strips Authorization on every request.
`http_client` propagates through `with_options()`/`copy()`, so the original
client and every copy are covered without depending on the SDK's
header-omission internals — the same custom-client shape already used by
`_build_anthropic_client_with_bearer_hook`.
Adds a real-loopback e2e test that simulates Omit unavailable and asserts
no Bearer leak on the original client and on a `with_options()` copy, while
x-api-key and request success are preserved.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>