Commit Graph

2999 Commits

Author SHA1 Message Date
teknium1 02b398bda5 fix(mcp): recovered application errors keep the breaker strike
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (3ff18ffe14,
#10447: a server answering errors made the model hammer it 8x in 10s).
Route the recovered result through _record_call_outcome so the caller
sees the tool's answer and the counter still moves the right way.
2026-09-12 08:23:28 -07:00
Dennis Hermsmeier 5dca47a651 fix(mcp): preserve application results after transport recovery 2026-09-12 08:23:28 -07:00
Teknium 850c48cd84 feat(skills-hub): tap K-Dense and OpenScience scientific skills under one "science" bucket
~480 scientific research skills become searchable/installable through the Skills Hub
with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and
synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0).

A new optional tap-level `bucket` key stamps extra["category"] on every skill from a
tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub
category; a sidecar grouping still wins when present. Both repos stay at community
trust (not in TRUSTED_REPOS) so the guard scans every install.

Re-grafted from #60559 onto the post-split tools/skills_hub_github.py.
2026-09-12 07:54:04 -07:00
teknium1 e7794124da fix(skills-index): bound ClawHub owner enrichment so the scheduled index build finishes
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.

Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.

- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
  without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
  (clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
2026-09-12 07:54:04 -07:00
Teknium e705dde6e4 chore(tests): pass encoding=utf-8 in test_skills_guard file writes (windows-footgun ratchet) 2026-09-12 05:09:26 -07:00
Teknium 596bd8fec6 fix(skills_guard): dns_exfil no longer fires on the English noun "host" in prose
`(dig|nslookup|host)\s+[^\n]*\$` matched any line where the word "host"
was followed, anywhere later, by a `$` -- "Set the host value and run
`${SKILL_DIR}/scripts/check.py`" was a CRITICAL DNS-exfiltration finding
that blocked a one-file community skill from installing (#108873).

DNS exfiltration puts the data in the queried NAME, so the pattern now
requires the interpolation in the first positional argument (after
optional -flags with values, +opts and @server). Real `host $SECRET.x`,
`dig @1.2.3.4 +short $TOKEN.x`, `nslookup -type=txt "$KEY".x` and
`host -t txt ${API_KEY}.x` still flag; the llama.cpp `--host ... $PORT`
exemption is preserved.
2026-09-12 05:09:26 -07:00
Teknium 4b8c01f691 fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).

Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
2026-09-12 01:35:05 -07:00
brooklyn! 4f1966edac feat(desktop): connector cards that wait for the sign-in, and a guided first launch that holds together (#108292)
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift

The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.

* refactor(desktop): one consent card for connectors and MCP setup

McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.

* feat(desktop): connector card drives the agent through manage_connections wait

The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.

Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.

* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them

The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.

* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate

A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.

* fix(agent): name a provider retry backoff on the live status line

The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.

* test(desktop): connector rehearsal launcher and flagged connector spec

connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.

* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout

The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.

* style(desktop): blank lines in connector-flow test per lint

* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film

The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.

* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds

The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.

* fix(desktop): no provider picker or free-tier chip over the guided first launch

Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.

* fix(desktop): onboarding connector picks are real catalog slugs

The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.

* fix(desktop): the free-tier ready screen never interrupts the guided chat

A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.

* feat(desktop): tour options that lead to building, and a fork that follows the tour

"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.

* feat(desktop): the onboarding connector picker reads the live catalog

A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.

* test(desktop): the guided first launch never forces a sign-in

The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).

* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape

Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.

* style(desktop): one answeredAfter helper for the ask card and first-build chips

* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows

Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.

* style(desktop): the 'nothing connects yet' line reads first on the connectors card
2026-09-12 10:00:06 +05:30
Teknium fc71fb63e5 test(multiplex): invariants for the residue fixes; retire allowlist tests
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own
  sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the
  routed scope (and does not borrow the default's on a miss); a stale served
  turn never recreates an archived profile; MCP discovery runs per profile home.
- test_config.py: v43 migration removes multiplex_profile_allowlist.
- Allowlist tests deleted (feature removed) or rewritten to "served set = all
  live profiles"; the unserved cases now use a tombstone / missing dir.
- MCP discovery fixtures converted to the per-home set/dict slot.
2026-09-11 19:50:46 -07:00
Teknium 9c9e7ab6e5 fix(multiplex): a served profile's turn sees its own cwd, approvals, redaction and tool policy
Under gateway.multiplex_profiles a secondary profile's turn ran with the LAUNCH
profile's working directory, command allowlist, redact_secrets switch, credential
file mounts, browser engine/headed flags, LSP service, auxiliary-provider health
marks and MCP stderr log, and several TERMINAL_ENV consumers read the process env
instead of the routed profile's terminal scope. A standalone `hermes -p X gateway
run` never behaved that way.

- tools/terminal_scope.py: resolve the terminal.cwd placeholder inside the
  profile scope with the same rule gateway/run.py applies at import (local ->
  $HOME, sandbox default otherwise) so the system prompt, context files and the
  terminal of a routed turn start where the profile's standalone gateway would.
- tools/image_source.py, credential_files.py, image_generation_tool.py,
  skills_tool.py, delegate_tool_progress.py, agent/tool_executor.py: read
  TERMINAL_ENV / TERMINAL_CWD through the terminal scope.
- tools/approval.py (+ approval_floors.py): one permanent allowlist per routed
  profile home; the unscoped module set stays for single-profile processes.
- agent/redact.py: `_redact_enabled()` resolves security.redact_secrets for the
  routed profile (scope .env, then config); launch snapshot kept when unscoped.
- tools/credential_files.py, agent/auxiliary_health.py, agent/lsp/__init__.py,
  tools/browser_tool_cloud.py, tools/mcp_tool_config.py,
  tools/tool_result_storage.py: key process caches by profile home (or bypass
  the slot under an override).

Tests: tests/tools/test_multiplex_turn_parity.py (4, red on base).
Docs: multi-profile-gateways.md isolation table.
2026-09-11 19:39:12 -07:00
Teknium adf23550f5 fix(tools): profile-scoped checkpoint/snapshot paths, tool caches, TZ and schema paths under multiplex
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:

- tools/process_registry.py, tools/environments/{modal,singularity}.py:
  `_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
  call time (same seam as `tools/skills_tool._skills_dir`, so the existing
  `monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
  the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
  path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
  per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
  resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
  the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
  the `_X_resolved` flags and the lifecycle reset keep their shape.
  tools/file_tools.py drops its private `file_read_max_chars` memo and reads
  the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
  env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
  startup) is ignored in favour of the routed profile's config.yaml. Both
  sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
  the static schema text is profile-neutral and `dynamic_schema_overrides=`
  rebuilds the `display_hermes_home()` / create-dir hint per
  `get_definitions()`, so a routed profile's model sees its own paths.

Refs #95685.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
2026-09-11 15:44:00 -07:00
Teknium 72cc96578a test: trim salvaged multiplex tests to the invariant pair per fix 2026-09-11 15:29:15 -07:00
Teknium 388b881b33 fix(gateway,tools): per-profile Yuanbao home, env_passthrough allowlist and Slack ignored-channel guard under multiplex
- gateway/platforms/yuanbao.py::AutoSetHomeMiddleware: the first authorized DM
  to a SECONDARY Yuanbao bot wrote YUANBAO_HOME_CHANNEL into os.environ, making
  that tenant's chat the default profile's cron/notification home. The write
  now only happens unscoped; reads go through the scoped reader + config.
- tools/env_passthrough.py::_config_passthrough: one module slot froze the
  first profile's terminal.env_passthrough for every profile's sandbox children;
  keyed by hermes_home_key().
- gateway/run.py::_slack_ignored_channels_from_gateway_config: the runner-level
  fail-safe only had the DEFAULT profile's GatewayConfig, so a secondary Slack
  bot's traffic was judged by the default's ignored list. It now takes the
  source's routed adapter (whose extra is the secondary's own config) and reads
  the env fallback through the scoped gate reader.
2026-09-11 15:29:15 -07:00
PRATHAMESH75 7af5006b24 fix(file-safety): bind the write-guard resolver fallback to the active profile
The per-call home/config getters fell back to `_expand_tilde("~/.hermes...")`
when the primary resolver raised. `_expand_tilde` follows the subprocess-HOME
contract, which under host `auto` mode can be the real/default user home rather
than the active multiplex `HERMES_HOME`. So on the exception path the guards
recreated the very cross-profile authority bug the happy path fixed: beta's
`config.yaml` was compared against the default/root config (hard-block fails
open), and the protected-instruction exemption resolved against the wrong home.

Re-derive both fallbacks from the same `get_hermes_home()` key the happy path
uses (`Path(home)/config.yaml`, `realpath(home)`), and substitute no unrelated
home if the active security path cannot be established — a `None` fails closed
at the protected-instruction consumer (exemption skipped, gate runs).

Adds opposite-side regressions: forcing the primary config resolver to raise
keeps beta's own config refused (and does not spuriously protect alpha's under
beta's scope); forcing the primary home resolver to raise keeps beta's
instruction-file exemption resolved against beta.

Addresses the fallback-authority review on #107335 (thanks @andrexibiza).

(cherry picked from commit 119d88b46f745ba081f12adf3f6ebba457d95d68)
2026-09-11 15:29:15 -07:00
PRATHAMESH75 1271622e4b fix(file-safety): resolve HERMES_HOME/config per call so multiplex profiles don't poison the write guards (#107327)
In a multiplexed gateway (`gateway.multiplex_profiles: true`) each profile turn
scopes `HERMES_HOME` through a per-turn contextvar. But
`tools/file_tools_write_guards.py` memoised the resolved home and config path in
process-global module state, filled once by whichever profile ran first. Both
the protected agent-instruction approval gate (`_get_real_hermes_home` →
exemption for a profile's own home) and the `config.yaml` hard-block
(`_get_hermes_config_resolved`) therefore became order-dependent: a later
profile's own `workspace/AGENTS.md` was gated against a *sibling* profile's home,
and — worse — its own `config.yaml` stopped matching the block, so a
prompt-injected agent could rewrite the very file the block exists to protect
(reproduced end-to-end in #107327).

Resolve both values per call instead. `get_hermes_home()` / `get_config_path()`
are contextvar-scoped, so the getters now track the active profile; the guard
already pays a `realpath` per call, so the extra cost is negligible. The two
module slots are kept purely as a test-override surface (set the slot + its
`_loaded` flag to pin a value); production leaves them unset and resolves live,
which also removes the cross-test poisoning the process memo could cause.

Adds regression coverage: both getters track the active profile after a prior
profile's scope, and the `config.yaml` hard-block fires for beta's own config
even after an alpha turn ran first.

(cherry picked from commit 36b257391da497ac31e6c440727dc57ecd584e71)
2026-09-11 15:29:15 -07:00
infinitycrew39 606903badc fix(tui): bind launch-profile terminal scope once multiplexing is active
After any secondary profile home is served, launch-profile turns used to stay
unscoped and fall back to ambient os.environ. Bind the launch home's own
terminal policy in that case so a poisoned ambient bridge can never become
the launch turn's authority (#107422 residual of #68559).

(cherry picked from commit f81147c1e5d283837e5e27f4da79710da7025235)
2026-09-11 15:29:15 -07:00
infinitycrew39 a5c801c8dc fix(tools): never ambient-bridge TERMINAL_* under a profile home override
A multiplexed dashboard can call _ensure_terminal_env_bridged while a
secondary profile's HERMES_HOME override is active. The one-shot latch then
wrote that profile's docker policy into process-global os.environ and poisoned
later unscoped launch-profile tool calls (#107422).

Skip the ambient bridge whenever a context-local home override is set —
ambient env is launch-profile authority only; routed profiles must use
terminal_scope (same rule as env_loader._reapply_terminal_config_bridge).

(cherry picked from commit 2050efb24fdc9d54a282b24d0042b90f47486c5a)
2026-09-11 15:29:15 -07:00
Teknium ceaf622c6d fix(mcp): same-named MCP servers with different credentials connect per profile; owner /reload-mcp keeps adopters' tools
Under gateway.multiplex_profiles every connection ledger in tools/mcp_tool.py
(_servers, _server_scope_keys/_server_tool_scopes, connecting/error/cooldown
maps, the circuit breaker, lazy schema-cache configs, trust metadata) was keyed
by the bare server NAME. The common per-tenant layout — each profile names its
server `github`/`notion` with its own token — gave only the first profile a
connection: the second profile's register_mcp_servers saw the name as "already
connected", refused to adopt it (different credentials, 4ddbcbd35e), and left
the profile silently tool-less with a healthy-looking `configured` status
(#106005 Bug 1/2, #91654). Siblings of the same bug: profile A's failing `x`
put profile B's healthy `x` into A's 10-minute connect cooldown and A's open
circuit breaker short-circuited B's calls; toolsets._resolve_toolset_memo was
not scope-keyed, so B resolved A's `mcp-<server>` tool names.

Keys are now the connection key from the new tools/mcp_tool_scope.py: the bare
name outside a multiplexer (single-profile processes are unchanged) and
(owner_scope, name) under one. Call-time lookups (_resolve_server_key) prefer
the calling scope's own connection, then a shared connection it adopted, so
identical-route profiles still share one subprocess. _select_new_servers,
the cooldown/breaker/trust maps, lazy registration and get_mcp_status all read
and write through the composite key; teardown resolves a task's key by
identity (the MCP loop has no profile context). The toolset memo key includes
registry.current_scope_key().

An owner's scoped /reload-mcp tore down its connection and, with it, every
adopting profile's tool overlay; nothing re-ran the adopters' discovery until
they reloaded. shutdown_mcp_servers(scope=) now records the orphaned adopters
and register_mcp_servers re-registers them under their own home + secret scope
once the owner's rediscovery pass completes.

Docs: multi-profile-gateways.md states the per-profile connection rule.

Fixes #106005
Fixes #91654
Co-authored-by: Bergmann89 <info@bergmann89.de>
Co-authored-by: Izzy-Gottz <srulynj@gmail.com>
2026-09-11 15:27:23 -07:00
Teknium a9838c2100 fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.

Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.

Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.

Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.

Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.

Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.

Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.

session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.

Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.

Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
2026-09-11 15:26:46 -07:00
Teknium 0dcadf6f41 revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.

Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
2026-09-11 11:54:49 -07:00
liuhao1024 df1388e77b test(tools): extend the delegation fd-leak injection to the schema replay
The schema initializer now replays SCHEMA_SQL through executescript (the
single-authority path from #94701's follow-up), which bypasses the
execute()-level DDL failure injection — the regression stopped raising.
Bind the same simulated failure onto the executescript path so the
connect-close-on-init-failure contract stays pinned for both replay
mechanisms.
2026-09-11 06:24:54 -07:00
teknium1 df0eed4f6b fix(session_search): a bare session id never reads another profile's state.db
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.

A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.

Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
2026-09-11 06:24:11 -07:00
shannonsands a6ee31f55a feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation

* feat(wisdom): add private contribution loop

* feat(wisdom): add managed consumption workflows

* fix(wisdom): close cross-repository safety gaps

* fix(wisdom): align local package and lifecycle policy

* fix(wisdom): require explicit profile setup

* docs(wisdom): repin reconciled gateway head

* fix(wisdom): fence content downloads and approval receipts

* docs(wisdom): record generation-fenced downloads

* docs(wisdom): record unified delivery PR

* fix(ci): stop passing invalid classifier inputs

* docs(wisdom): remove internal requirements ledger

* feat(wisdom): localize dashboard and desktop copy

* feat(wisdom): complete local contribution and consumption UX

* style(wisdom): satisfy desktop lint

* chore(wisdom): refresh requirements pin

* test(dashboard): allow formatted profile copy

* test(wisdom): stabilize desktop interaction coverage

* fix(wisdom): surface dashboard action failures

* fix(wisdom): add repeatable Portal demo login

* feat(wisdom): add actionable skill notifications

* feat(wisdom): add notification install and update actions

* fix(wisdom): make Telegram skill alerts actionable

* fix(wisdom): always refresh demo Agent login

* feat(wisdom): embed Telegram notification actions

* fix(wisdom): preserve Telegram notifications after actions

* fix(wisdom): keep Telegram notification cards readable

* feat(wisdom): add Telegram candidate approval flow

* feat(wisdom): explain Telegram qualification reasons

* fix(wisdom): reconcile cross-surface candidate actions

* feat(telegram): add Collective Wisdom management command

* chore(wisdom): refresh Gateway contract pin

* chore(wisdom): advance Gateway contract pin

* feat(wisdom): align command UX across clients

* feat(slack): add Collective Wisdom management parity

* feat(wisdom): add security and professionalism reviews

* feat(wisdom): add first-time qualification guidance

* feat(wisdom): simplify qualification sharing choices

* feat(skills): add optional editorial metadata

* feat(wisdom): enrich legacy skill presentation

* fix(wisdom): harden review and update boundaries

* fix(wisdom): emit canonical review timestamps

* fix(wisdom): align with merged gateway and main

* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)

- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
  7-day evidence builder that excludes bundled/hub/managed skills and
  dismissed/handled/recently-suggested content hashes, strict pydantic
  schemas for agent output with repair-or-reject, fixed copy templates
  (Share / Teammate / Published / Update / Mute), idempotent retried
  delivery ledger with stale-action resolution, weekly review job,
  resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.

* wisdom: agent-led renderers and button action dispatcher

- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
  is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
  ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
  packaging flow, Install/Update -> plan command. Never publishes/installs.

* wisdom: CLI verbs, agent_led config default, conversational catalog skill

- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
  verbs, share/install flows and fixed notification templates.

* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons

- gateway housekeeping tick calls maybe_run_weekly_review with a home
  channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
  duration keyboard, send_wisdom_agent_recommendation rich card + fallback.

* fix(wisdom): integrate local mediation and harden model and setup boundaries

* fix(wisdom): honor authoritative recommendation policy and defer on failure

* fix(wisdom): synchronize opaque suppression and recheck delivery preferences

* feat(wisdom): route weekly selection through the session-owned assessment queue

* fix(wisdom): prepare and submit the reviewed generated share package

* feat(wisdom): separate native Share preparation from publication consent

* feat(wisdom): sync native mute choices through a leased preference outbox

* feat(wisdom): bind native mute controls to durable preference choices

* feat(wisdom): add scoped desktop and dashboard notification settings

* fix(wisdom): revalidate feed recommendations before assessment and delivery

* fix(wisdom): persist validated delivery receipts before completing notices

* feat(wisdom): add private notification claim and receipt client

* Persist Wisdom send reservations and recover delivery acknowledgements

* Route legacy Wisdom controls through current native review

* Add typed private Wisdom operation outcome client

* fix(wisdom): make agent-led advice usable in the local demo

* fix(wisdom): keep requested consent outside proactive limits

* fix(wisdom): distinguish unavailable assessments and preserve digest text

* fix(wisdom): assess ongoing usefulness beyond the current task

* fix(wisdom): restore immediate qualification sharing controls

* fix(wisdom): separate qualification review from installation advice

* fix(wisdom): collapse review checklists and simplify sharing copy

* fix(wisdom): show compact sharing progress and publication receipts

* fix(wisdom): require credential prefixes rather than matching skill names

* fix(wisdom): finish package checks before presenting sharing consent

* fix(wisdom): scan local skills before qualification cards

* fix(wisdom): update moderation results on existing sharing cards

* fix(wisdom): keep sharing review accessible from receipt cards

* fix(wisdom): align mediated review cards and collapsible checks

* fix(wisdom): clarify clean security summary wording

* fix(wisdom): normalize consent plans and add explicit recheck

* fix(wisdom): keep install and update receipts concise

* fix(wisdom): collapse assessments and deduplicate operation cards

* fix(wisdom): restore private Portal review from native cards

* fix(wisdom): sync Portal publication to original consent card

* fix(wisdom): show local skill version on sharing cards

* fix(wisdom): skip agent recommendations for self-published versions

* fix(wisdom): simplify candidate notices and local-edit recovery copy

* feat(wisdom): submit locally reviewed packages with one confirmation

* feat(wisdom): expose safe receipt and outcome sync recovery

* wisdom: onboarding notice says detect and share, names the user's own skill

Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark

Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.

* wisdom: one opener, no approval line, ask to share after the skill is shown

Product owner review of the candidate card.

- The Hermes written card now opens with the same sentence as the fixed card
  ("Your organisation has enabled Collective Wisdom, a feature designed to
  automatically detect and share useful skills across all team members.")
  instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
  and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
  It is now the last line, after the skill name, description, why suggested
  and the checks, and reads "Would you like to share it?" (matching the
  agent led template wording).

Tests updated for the new order; proposalNotice removed from all desktop locales.

* wisdom: American spelling, organization

Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.

* wisdom: candidate card copy round 4 (owner review)

Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:

1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
   the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
   card (Telegram rich card and plain fallback, legacy agent-led share
   template).
3. The skill name and description are labelled: "Skill name: <name>" and
   "What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
   inappropriate content found)" with no per-check bullets and no "Pass";
   a failed review reads "Needs a look before sharing at work (possible
   inappropriate content)" and lists only the checks that flagged
   something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
   details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
   days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
   editorial_name, a simple one_line_description and a compelling
   why_coworkers_benefit under 300 characters; "Be concise and
   convincing." becomes "Be concise and compelling: the goal is that the
   user wants to share it."

Tests updated for the new strings; review_text() gains direct coverage.

* wisdom: re-apply owner copy after rebase

- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice

* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors

Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.

* fix(wisdom): reconcile optional SDK tests and frontend lint

* fix(wisdom): default to agent-written notification summaries

* fix(wisdom): restore deferred install review and browse controls

* feat(wisdom): inspect installed setup with exact package provenance

* feat(wisdom): run native-approved installed setup steps with durable evidence

* fix(wisdom): recover interrupted setup with explicit native consent

* feat(wisdom): hand native installs into guided setup review

* fix(wisdom): continue requested setup with fixed notification copy

* fix(wisdom): preserve setup while waiting for a session model

* fix(wisdom): expose canonical setup review controls on desktop

* fix(wisdom): resume setup after recorded automatic updates

* fix(wisdom): make missing setup prerequisites recheckable

* chore(wisdom): align Agent with verified Gateway contract

* fix(wisdom): stop guessing team slugs in portal links

* fix(wisdom): retire pending advice on account sign-out

* fix(wisdom): cancel advice after terminal account revocation

* fix(wisdom): fence feed responses across account sign-out

* fix(wisdom): checkpoint signed-out feed before reactivation

* fix(wisdom): link proactive advice to scoped notification settings

* fix(wisdom): coalesce queued publication recommendations by version

* fix(wisdom): keep package review navigation local and deferable

* fix(wisdom): reflect installed state in discovery controls

* fix(wisdom): show exact checks before command confirmation

* chore(wisdom): pin bounded analytics privacy contract

* chore(wisdom): pin retired legacy notification contract

* feat(wisdom): review publisher usage with exact sharing copy

* fix(wisdom): align discovery and review check summaries

* fix(wisdom): show expired consent before confirmation

* fix(wisdom): require fresh review for legacy install controls

* fix(wisdom): preserve review expiry across check toggles

* fix(wisdom): retain update policy in native install reviews

* fix(wisdom): surface failed native card edits

* fix(wisdom): persist local command approval reviews

* fix(wisdom): use saved approvals for messaging commands

* test(wisdom): provide scan result in setup handoff fixture

* test(wisdom): exercise Telegram approvals with saved review state

* fix(wisdom): retain suppression policy for offline deferral

* fix(wisdom): reconsider candidates after deferred suppression expires

* fix(wisdom): bind review checks and report verified readiness separately

* fix(wisdom): persist accepted publication intent and recover exact outcomes

* fix(sync): pin UTF-8 tree ordering across writers

* chore(wisdom): pin organisation-scoped Gateway authorization

* fix(wisdom): restrict consent delivery to user-facing sessions

* chore(wisdom): refresh reviewed Gateway contract pin

* fix(wisdom): preserve kept tools in Blank Slate exclusions

* test(auth): reset anonymous fixture with a profile-scoped cache

* fix(wisdom): gate local surfaces and work on current profile entitlement

* fix(wisdom): invalidate quiet tool cache on entitlement changes

* test(wisdom): authorize local consent gateway fixtures

* fix(wisdom): keep entitlement decoding free of native crypto imports

* test(wisdom): provide local entitlement to demo CLI subprocess

* ci: leave upstream workflow unchanged in Wisdom PR

* fix(wisdom): ship package and contracts in Nix wheels

---------

Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
2026-09-11 19:04:06 +10:00
kshitijk4poor bffa5f75e4 test(terminal): trim NOPASSWD probe tests to behaviour contracts
Six tests became three: the _prepare_command wiring test, one parametrized
probe test (rc 0 -> True, rc 1 -> False, unsupported backend -> never runs),
and a headless guard proving the probe is not paid when no prompt can fire.
The exact-kwargs assertion on `_run_bash(..., login=False, timeout=3,
stdin_data=None)` froze incidental defaults and is gone.
2026-09-11 11:19:53 +05:30
kshitijk4poor 4fc64b6730 refactor(terminal): probe NOPASSWD only when a prompt would fire; drop the host-only probe
With BaseEnvironment supplying a backend-scoped probe to every production
caller, the module-level host-only `_sudo_nopasswd_works` (and its
TERMINAL_ENV gate) had no callers left; the `or` fallback and the outer
try/except around a callback that already fails closed were dead too.

Move the probe under `should_prompt_for_sudo`: headless callers (gateway,
cron, delegated children) reach `(command, None)` whether or not the probe
runs, so the extra backend round trip — an ssh exec on the SSH backend —
was pure waste on every headless sudo command.

test_subagent_sudo_prompt no longer needs to patch the host probe out:
bare `_transform_sudo_command(cmd)` calls have no probe by construction.
2026-09-11 11:19:53 +05:30
fangliquanflq 39abca492d fix(terminal): gate sudo probes by backend cancellation safety
Opt in only Local, Docker, SSH and Singularity: a timed-out `sudo -n true`
probe on those backends kills one process, while SDK adapters (Modal,
Daytona, Vercel) cancel by terminating the whole sandbox.

Net of the original PR's commits f7db10ef + 9965b1b9 (the intermediate
_ThreadedProcessHandle special-case was superseded by this gate).
2026-09-11 11:19:53 +05:30
fangliquanflq d38064e3c1 fix(terminal): probe remote passwordless sudo 2026-09-11 11:19:53 +05:30
Teknium 45a6101f36 fix(gateway): secondary-profile send_message, notices and /loop wakeups go out via their own bot
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved
its live adapter by bare platform from runner.adapters — the DEFAULT profile's map —
so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy,
every plugin platform), its "Gateway shutting down/restarted" and /update notices,
its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left
through the default bot (Telegram DMs landed in the user's chat with the other bot).

Every such door now resolves through the profile-aware, fail-closed resolver already
used by the inbound reply path (authz_mixin: _adapters_for_profile /
_authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear
error, never the default bot.

- tools/send_message_senders.py::_live_adapter — resolve via
  runner._authorization_adapter(platform, get_active_profile_name()); shared by
  _send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and
  the WeCom standalone sender.
- gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay-
  aware resolve_delivery_transport callers); _authorization_adapter reuses it.
- gateway/run_shutdown.py — shutdown/restart notice for a running session uses the
  session's source transport / agent:<profile>: key lane, never self.adapters.
- gateway/slash_commands.py + run_notifications.py — /restart and /update markers
  persist `profile`; the restart notice, update result and update prompt resolve the
  requester's own adapter (legacy markers fall back to the session_key lane).
  /goal, /heartbeat, /approve, /deny confirmations use the source's own transport.
- gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its
  route; the wakeup watcher scans every served profile's store under its own scope
  (same shape as _handoff_watcher) and fires through that profile's adapter map.
- plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in
  the owning profile (its adapters and its home channels).
- plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter.
- docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity).

Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py,
tests/gateway/test_multiplex_notice_egress_profile_adapter.py.

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-09-10 18:17:22 -07:00
Teknium 580322ef1e fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for
every name tagged in the process-global `_SECRET_SOURCES` map. That map is
filled by EVERY served profile's secret-source hydration, while `os.environ`
only ever holds the LAUNCH (default) profile's values — so once any profile's
1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio
MCP server was started with the default profile's token.

Resolve those names through the active profile's secret scope (`get_secret`)
instead: the routed profile's value, or omitted when that profile has none.
Under multiplex `get_secret` never falls through to environ; single-profile
runs keep the .env overlay + environ behaviour, so the existing "vault vars
reach MCP subprocesses" contract still holds there. `secret_source_names()`
exposes the tagged NAMES only — values are never read from the shared map.

Docs: the multi-profile guide's "MCP subprocesses only see their own profile's
secrets" claim is now true for source-injected names too; say so explicitly.
2026-09-10 18:12:39 -07:00
Teknium 1bd4e33d36 fix(vault): make browser_vault_fill work on the default Browser Use backend
On the default backend (browser.backend unset → browser_exec) the vault tools were
advertised but could never fill: the CDP supervisor that carries the secret-bearing
eval is started only by the built-in browser_* session path, so _eval_js_secret
failed closed with supervisor_required and the origin pre-check fell back to an
agent-browser CLI eval against a browser browser_exec never touched.

- browser_exec now attaches SUPERVISOR_REGISTRY to the CDP endpoint it just routed
  the harness to (BU_CDP_WS/BU_CDP_URL), so the fill talks to the SAME browser over
  the same secret-capable WebSocket. BU direct-cloud (BU_AUTOSPAWN) exposes no
  endpoint and keeps the supervisor_required refusal.
- CDPSupervisor.focus_page(origin, accept=<js>) (used by browser_vault_fill, next commits) re-attaches the page session to the
  open tab on the item's origin whose DOM holds the form being filled (browser_exec
  opens its own tabs; the supervisor's initial attach picks the first page target,
  which is chrome://new-tab-page). browser_vault_fill uses it before the origin
  pre-check with a per-kind probe (password input / card fields / address fields).

Live: evals/vault_fill_live_e2e.py drives the real browser_exec tool against Hermes'
packaged Chromium with the login page in the third tab; A/B with the attach line
disabled fails at "did not attach a supervisor", enabled fills the password into the
/login tab and card fields into the /checkout tab with every model-facing read
scrubbed.

Also: browser_vault_list/fill described the workflow as "type the identifier with
fill_input", a helper that exists only inside browser_exec code (toolset browser-use) and is
a ghost on the built-in stack. model_tools._rewrite_browser_vault substitutes the concrete
name from the session's actual tool set (`fill_input` inside browser_exec, or browser_type),
the same dynamic cross-reference pattern browser_navigate uses for web_search.
2026-09-10 10:35:07 -07:00
Teknium 4140901d15 fix(browser): stop refusing credential-named query params on cloud browser/extract backends
browser_navigate / browser_exec / web_extract refused any URL whose query carried a
credential-NAMED parameter (token, signature, access_token, ...) when the backend was
a cloud provider. That is exactly the shape of magic links, OAuth callbacks and signed
CDN assets, so on Browserbase/Browser Use the agent could not finish a sign-in flow or
open an X video asset ("Blocked: URL contains a credential-like query parameter").

The floor protected nothing: the cloud browser already sees every cookie and typed
password of the session, and with the credential vault it receives the real password at
fill time. Hermes' own secrets leaking into a URL stay blocked by the value-shaped
_PREFIX_RE check (_secret_url_error), which is backend-independent. IMDS and
private-address floors are unchanged.
2026-09-10 10:35:07 -07:00
Teknium b51da65258 fix: adapt execute_code cell authority to the widened prompt-callback table
_callback_api() now yields (getter, setter) pairs for every per-thread prompt
(approval, sudo, vault unlock); the kernel cell captured and restored the old
fixed 4-tuple. Iterate the table so a cell carries every callback and a future
addition needs no change here. Test recorder unpacks the new shape.

Also: perfectionist import order in ui-tui interfaces.ts (CI lint).
2026-09-10 10:35:07 -07:00
Teknium ff90eb28ff test(bot-mode): trim the author and loop-guard suites to their invariants
Keep one or two behaviour tests per seam (author reset on a cached agent, forged
_turn_author refused, guard trips and cools, single charge on the busy path) and
drop the parser/setting enumerations. a2a_key goes with them: nothing in this PR
reads it; the honcho follow-up that does can bring it back with its consumer.
2026-09-10 10:27:07 -07:00
Erosika 091ac34ba9 fix(bot-mode): qualify the peer-dm author id with the sender's hostname
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id.
2026-09-10 10:27:07 -07:00
Erosika 3db4defcc1 fix(bot-mode): qualify a relayed author from the local connection too
delivery_turn_author kept the bare bot:<profile> id when the sender's connection was the Desktop's own "local", so a DM relayed from that machine collided with the recipient's profile of the same name. A relayed DM always crosses gateways, so the connection id is now part of the id whenever the Desktop sends one, and only the direct message_agent path in tools/bot_mode_dm.py stays bare. The relay.ts and session_auto_continue.py comments added earlier are cut to one line each.
2026-09-10 10:27:07 -07:00
Erosika d3c8bacbe3 fix(bot-mode): qualify relayed authors with the sender's connection id
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
2026-09-10 10:27:07 -07:00
Erosika 6881e4d3fc fix(bot-mode): carry the relay sender into a live Bot Chat turn
When the target Bot Chat is already open on this gateway, the relay handler delivers through
`prompt.submit` with `queued: true`, and that branch dropped the envelope's sender. The model
still saw the text prefix, but the turn reached the agent unattributed, the exact case the
subprocess branch fixes.

The relay handler now stamps the author on the submit as a `DeliveryAuthor`, an in-process object
a JSON client cannot build, so `prompt.submit` accepts it the way it accepts a hosted-room callback
and refuses a dict with error 4124. The busy queue keeps an authored envelope in its own slot, the
drain hands the author to the turn runner, and the runner passes it to an agent that declares the
keyword. A plain prompt after an authored dm carries no author.

Local deliveries to a desktop-owned Bot Chat take the live-owner mailbox instead. The admission
intent and the mailbox record now carry the author, a retry under the same id with a different
author is refused, and the owner gateway hands the author to the turn it runs. Isolated compute
turns still run unattributed, because the compute-host frame has no author field.
2026-09-10 10:27:07 -07:00
Erosika 55b3ea0b11 fix(gateway): count each bot message once in the loop guard and consume the author variable
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it
again, and the busy path asks a third time. Each call counted one loop-guard event, so a
Telegram bot tripped the budget after a third of the configured messages. The verdict now only
refuses a chat that is cooling down. The ingress gate counts an admitted bot message once.

`parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag,
and returns None for an author with neither id nor name. Names keep format characters and
non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR
before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive
number. Issue numbers move out of code comments.
2026-09-10 10:27:07 -07:00
Erosika 16d15869ba feat(bot-mode): carry the sender through local and relay deliveries
a bot dm arrived as an ordinary user message. the only trace of the sender
was the "Message from" text prefix, which the model reads and nothing else
does. the recipient's memory provider saw its own configured user.

message_agent now passes the sender as {"id": "bot:<profile>", "name":
<handle>, "is_bot": true} to the delivery runner (--author <json>), which sets
HERMES_TURN_AUTHOR on the recipient one-shot only. the -Q turn reads it and
passes turn_author into run_conversation. the desktop relay forwards the
envelope's from_profile/from_handle to bot_relay.deliver, which sets the same
variable on its delivery turn. the runner drops any inherited author first so
a delivery without one stays unattributed. the text prefix is unchanged.
2026-09-10 10:27:07 -07:00
Teknium ac07e20407 fix: resolve subagent control authority from the live session slot
Subagent list/tail/steer/interrupt authorized against a per-record copy of
the owning session's transport (`owner_transport`). That copy had to be
re-synced at every reattach site; `_rebind_live_transport` did it for
session.resume/activate but prompt.submit and the queued-prompt drain still
attached bare, so a client that reconnected through a prompt (the common
path on a remote gateway / Bot Mode switch) streamed fine while
`subagent.list` returned [] and controls rejected.

Read `owner_session_record["transport"]` at check time instead: the slot is
already mutated by every attach/detach/viewer-failover path, so no site can
forget the sync. `owner_transport` stays as the capture-time "commissioned
by a gateway session" marker (None = no RPC authority ever); non-dict owners
keep the exact-object rule. Drops the registration-time re-read and the
attach-time registry loop.

Diagnosis credit: nftpoetrist (#106663) — their prompt.submit / drain
regression tests pass against this change with no call-site edits.
2026-09-10 10:23:12 -07:00
Erosika 37eb6e1b05 test(bot-mode): pin the waiter budget from Python constants, not relay.ts text
The waiter budget test read relay.ts with a regex, which AGENTS.md bans and which the Python CI lane would not rerun on an apps/-only PR. It now checks DESKTOP_DELIVER_TIMEOUT_SECONDS against the module's own constants and that REPLY_WAIT_SECONDS exceeds it. relay-deliver-budget.test.ts still pins the TS mirrors against those Python constants from the Desktop side.
2026-09-10 10:18:27 -07:00
Erosika b1915cb02d fix(bot-mode): keep the relay waiter watching past the Desktop deliver deadline
The sender-side waiter gave up at 900s while the Desktop held bot_relay.deliver open for 1500s, so a turn finishing between minute 15 and minute 25 wrote a reply nobody read. REPLY_WAIT_SECONDS now rebuilds the Desktop budget from the same numbers and waits 60s past it. The two turn constants move into tools/bot_relay.py so the gateway handler and the waiter share one definition.
2026-09-10 10:18:27 -07:00
Victor Kyriazakos 49ef015ca3 fix: join delegated work before finite chat exits 2026-09-10 05:02:41 -07:00
Siddharth Balyan b4d04eb8fd Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding

* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg

The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.

On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.

* docs(tool-search): connectors section — remote tools through the bridge

The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.

* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots

dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.

The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.

The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).

Live, 311 local tools + gateway, before -> after:
  "send gmail email":           5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
  "read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
  "linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.

* refactor(tool-search): connector leg into tools/connector_search.py

tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.

No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.

* fix(tool-search): at most 7 queries per call, the gateway's search limit

One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.

The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.

* fix(tool-search): the model is told that connectors__ names are manage_connections accounts

tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.

The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.

Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.

Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.

* fix(connectors): /stop halts a connector batch before the next remote call

dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.

The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.

Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.

* test(connections): schema assertions become dispatch contracts

test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.

Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.

Test count in the file goes from 26 to 25.

* docs(tool-search): connector batches are one gateway request per entry

The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.

Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.

Docs only, no test.

* fix(tools): the between-turns refresh never rewrites the bridge tools

The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.

The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.

Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.

* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation

The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.

* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote

bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.

The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.

Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.

* fix(connectors): search keeps the twin a colliding name reaches, and says so

format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.

Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
2026-09-10 02:21:16 +05:30
Siddharth Balyan cf4b78e91f fix(tool-search): a query no tool answers returns nothing, not five tools sharing one word (#106676)
search_catalog admitted every document with BM25 score > 0 and then padded
to `limit`. BM25 sums over the tokens a document shares with the query, so
on a 300-tool catalog "send gmail email" returned five incident tools that
shared only "email", and the discriminating word ("gmail", in no document)
had no say. The model read those as the answer.

Admission is now the query's rarest token: a document is a result only if it
contains the query token with the highest IDF, the one that names the intent.
Common verbs ("send", "read", "create") sit in dozens of documents and never
gate; vendor and object words ("gmail", "github", "incident") do. A token no
document carries admits nothing, and the existing empty-group hint tells the
model to retry without it. The name-substring fallback is deleted: it admitted
tools that matched no query token at all.

Result descriptions are clipped at 500 characters instead of 400. Over 353
vendor tool descriptions, 500 keeps 91% whole and every first sentence
(first-sentence max 329); 400 kept 82%.

Measured on the live 311-tool catalog with 25 hand-labelled queries:
precision@5 0.18 -> 0.43, wrong names returned 102 -> 66, false positives on
absent intents 17 -> 13. Live before/after: "send gmail email" went from five
betterstack tools to an empty group with the retry hint; "linear create issue"
and "betterstack incident" are unchanged.
2026-09-10 02:21:15 +05:30
Teknium b7bef04861 fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver
every surface funnels through, `kanban_db.create_task`: board-project
inheritance now runs only when the caller left `workspace_kind` open
(`None`), and `workspace_kind` defaults to scratch after that check. The
tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and
`board=` scoping but drops its handler-local sentinel logic, since the
resolver now owns the rule; the `self_task` project inheritance for
dispatcher-owned workers is unchanged.

Sibling surfaces had the same bug through the same line and are fixed by
the same change:
- CLI `hermes kanban create --workspace scratch` on a project-scoped board
  produced a project worktree; `--workspace` no longer defaults in
  argparse so the resolver can tell "omitted" from "scratch".
- Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same;
  `CreateTaskBody.workspace_kind` defaults to `None` for the same reason.
- `kanban_swarm.create_swarm` threads `None` through for consistency.

Tests: one resolver invariant in test_kanban_board_project.py (explicit
scratch stays scratch, omitted still inherits) and the salvaged tool test
folded into a single parametrized matrix over scoped/unscoped target boards.
2026-09-09 12:18:58 -07:00
KoNit-K dc5c87c6c1 fix(kanban): honor explicit scratch/project on MCP kanban_create
Pass the caller's board into create_task and treat explicit workspace_kind=scratch or empty project as no-project so ambient board project_id cannot override MCP args.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-09 12:18:58 -07:00
Teknium 85e423482a fix(tools): kill the browser_exec CLI tree on timeout on Windows too
The salvaged fix only ran the CLI in its own session on POSIX and kept
plain subprocess.run on Windows — the one platform where the wedge in
#106244 is actually reproducible: CPython's run() retries an unbounded
communicate() after kill() there, so a grandchild holding the capture
pipes blocks the worker forever. On POSIX run() wait()s the PID and
returns promptly; the grandchild merely leaks (live-reproduced on Linux).

One code path for both platforms:
- _group_popen_kwargs: start_new_session=True on POSIX,
  CREATE_NEW_PROCESS_GROUP + hide flags on Windows (replaces the
  hide-only _windows_popen_kwargs).
- _kill_cli_process_group: os.killpg SIGKILL on POSIX, taskkill /T /F
  on Windows (same kwarg set as the sibling taskkill sites).
- The Popen decodes with encoding="utf-8", errors="replace" like every
  other subprocess call in this file (windows footgun rule).
- Drain test patches the kill helper instead of os.killpg so it runs on
  every host.

Windows behaviour is not live-verifiable on this Linux host.
2026-09-09 12:11:13 -07:00
Teknium 6788c3224f test(tools): trim the browser_exec group-kill tests to two invariants
Drop test_gone_group_still_surfaces_timeout: the killpg-race branch is a
contextlib.suppress and the test only pinned the exact communicate()
timeout sequence (a change-detector). The real-grandchild test and the
bounded-drain test remain — the two behaviour contracts of the fix.
2026-09-09 12:11:13 -07:00
liuhao1024 60debff28d fix(tools): kill the whole browser-use CLI process group on browser_exec timeout
subprocess.run only kills the direct CLI child on TimeoutExpired; a
browser_harness daemon / Chrome helper grandchild inherits the stdout/
stderr pipes and keeps them open, so the internal communicate() blocks
on pipe EOF forever. The wedged tool call never returns, its activity
heartbeat keeps stamping last_activity_at every 30s, and the session is
pinned at "now" in the desktop sidebar indefinitely (#106244).

On POSIX, run the CLI in its own session (start_new_session=True) and
SIGKILL the whole process group on timeout, then drain the pipes under
a bounded deadline. Windows keeps subprocess.run.
2026-09-09 12:11:13 -07:00