Commit Graph

33712 Commits

Author SHA1 Message Date
srojk34 e70db09f51 fix(security): re-resolve checkpoint/sticker-cache paths per call
tools/checkpoint_manager.py's CHECKPOINT_BASE and gateway/sticker_cache.py's
CACHE_PATH are resolved once at import time via get_hermes_home(), which is
a context-local ContextVar under the multiplexed gateway (multiple profiles
sharing one process). Freezing the path at import time pins every later
checkpoint/cache read-write to whichever profile's HERMES_HOME was active
when the module was first imported -- the same bug class already fixed for
cache dirs, skills_hub, rich_sent_store, and (this session) the OAuth/auth.json/
sessions.json paths.

CheckpointManager is "owned by AIAgent" per-instance, but its methods read
the frozen module constant directly instead of taking the store root from
the instance, so a profile's CheckpointManager can read/write code-edit
checkpoints into a different profile's store.

Add a per-call resolver for each path, following the established "respect
an existing test monkeypatch of the constant, otherwise re-resolve through
get_hermes_home()" pattern so the extensive existing test seams in
tests/tools/test_checkpoint_manager.py and tests/gateway/test_sticker_cache.py
keep working unmodified.

(cherry picked from commit 03ae075d969094cb584e6ab38d2a773d15ff875c)
(cherry picked from commit b850c4b18e2ae2158a97c6cb87bd2057918b8170)
2026-09-11 15:44:00 -07:00
Teknium 819517fbac fix(gateway): routed profiles get their own max_turns, fallback chain, hooks, aux auth and media policy
One multiplexed gateway process serves every profile, but several per-turn
reads still went through state frozen from the LAUNCH profile:

- `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the
  module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS,
  so every secondary ran with the default profile's turn budget. A routed turn
  (HERMES_HOME override) now resolves `agent.max_turns` from its own config.
- `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide
  slot, so secondaries fell back through the default's provider/model with their
  own keys. It now reads the active gateway home and keeps a last-known-good
  chain per home.
- `_load_prefill_messages` resolved relative paths against the launch home.
- `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a
  secondary's compression/title/vision calls authenticated to Nous with the
  default profile's token when it had no pool entry. Resolved per call via
  `hermes_cli.auth._auth_file_path()` (patched constant still wins in tests).
- `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was
  loaded outside any profile scope, so secondaries' `hooks/` never ran and the
  default profile's handlers received every profile's messages, responses and
  user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the
  runner holds one registry per served home, picked from the active scope at
  emit time and front-loaded under each secondary's startup scope.
- Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME
  and the default profile's secrets). They now get the routed HERMES_HOME via
  `build_subprocess_env`, scrubbed under multiplexing, and the stdin payload
  carries `profile` so one script can tell which profile fired it.
- Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`,
  `trust_recent_files*`) was bridged once into env at startup and read from env
  per delivery; under a HERMES_HOME override the validator now reads the routed
  profile's config. Single-profile runs keep the env-bridge contract.

Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro
(temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7,
fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values.
2026-09-11 15:44:00 -07:00
srojk34 0c74353c86 security(gateway): re-resolve hooks directory per call to fix profile isolation
gateway/hooks.py::HOOKS_DIR is resolved once at import time via
get_hermes_home(), which is a context-local ContextVar under the
multiplexed gateway (multiple profiles sharing one process, each owning
its own Gateway/HookRegistry instance). Freezing the path at import time
pins every later HookRegistry.discover_and_load() call to whichever
profile's HERMES_HOME was active when this module was first imported --
so a later-starting profile silently discovers and executes the FIRST
profile's hook handlers (arbitrary Python code, not just data) against
its own live event context, including session_id/message/response text.
Same bug class already fixed for cache dirs, skills_hub, rich_sent_store,
and the OAuth/auth.json/sessions.json/checkpoint/sticker-cache paths.

Add a per-call resolver, following the established "respect an existing
test monkeypatch of the constant, otherwise re-resolve through
get_hermes_home()" pattern so the existing test seam in
tests/gateway/test_hooks.py keeps working unmodified.

(cherry picked from commit 1e4f96af681de90b942ccb77e3f96ee03913ddc2)
2026-09-11 15:44:00 -07:00
Teknium 53e32d0581 fix(env_passthrough): tolerate an unresolvable home when keying the allowlist cache
_make_run_env runs with a stripped environ on Windows children; hermes_home_key()
raises RuntimeError there (no HOME/USERPROFILE). Fall back to an unkeyed slot
instead of failing the sandbox env build.
2026-09-11 15:29:15 -07:00
Teknium d198082172 chore(contributors): map see-k's commit email for release credit 2026-09-11 15:29:15 -07:00
Teknium 72cc96578a test: trim salvaged multiplex tests to the invariant pair per fix 2026-09-11 15:29:15 -07:00
Teknium 651878e1e4 docs(multiplex): per-profile config.yaml behaviour, env_passthrough and write guards 2026-09-11 15:29:15 -07:00
Teknium 388b881b33 fix(gateway,tools): per-profile Yuanbao home, env_passthrough allowlist and Slack ignored-channel guard under multiplex
- gateway/platforms/yuanbao.py::AutoSetHomeMiddleware: the first authorized DM
  to a SECONDARY Yuanbao bot wrote YUANBAO_HOME_CHANNEL into os.environ, making
  that tenant's chat the default profile's cron/notification home. The write
  now only happens unscoped; reads go through the scoped reader + config.
- tools/env_passthrough.py::_config_passthrough: one module slot froze the
  first profile's terminal.env_passthrough for every profile's sandbox children;
  keyed by hermes_home_key().
- gateway/run.py::_slack_ignored_channels_from_gateway_config: the runner-level
  fail-safe only had the DEFAULT profile's GatewayConfig, so a secondary Slack
  bot's traffic was judged by the default's ignored list. It now takes the
  source's routed adapter (whose extra is the secondary's own config) and reads
  the env fallback through the scoped gate reader.
2026-09-11 15:29:15 -07:00
Teknium 545e74d0ea fix(gateway): a secondary profile's config.yaml no longer poisons the process env under multiplex
Under gateway.multiplex_profiles every secondary profile's config loads inside
_profile_runtime_scope, yet every apply_yaml_config_fn hook (feishu, matrix,
whatsapp, slack, dingtalk, discord non-gate keys, telegram non-gate keys) and
gateway/config_loader.py::bridge_core_env_settings still wrote os.environ there.
First-writer-wins: the first secondary with a require_mention / allowlist /
allow_bots / reactions block made that policy the DEFAULT profile's (live:
TELEGRAM_REQUIRE_MENTION written from a secondary load), and secondaries read
the default's env for the same keys.

- gateway/platforms/_shared.py::yaml_env_setter: the one env-write shape for
  YAML->env bridges — env wins, skipped under an active secondary scope.
- Every hook now seeds its values into the profile's PlatformConfig.extra and
  uses yaml_env_setter; bridge_core_env_settings seeds telegram/signal
  require_mention into extra and skips the env write under scope.
- Readers that bypassed extra/scope now consult extra first (matrix flags +
  session_scope, slack reactions, telegram reactions/mention_patterns/_extra_bool,
  discord reactions/auto_thread/history_backfill/approval_mentions/allow_mentions,
  feishu allow_bots, dingtalk mention_patterns, signal require_mention).

Salvages the shape of PR #100604 (whatsapp, earliest report #80099), #100435
(discord) and #100448 (telegram) by @nftpoetrist on top of current main.

Fixes #80099

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-11 15:29:15 -07:00
PRATHAMESH75 7af5006b24 fix(file-safety): bind the write-guard resolver fallback to the active profile
The per-call home/config getters fell back to `_expand_tilde("~/.hermes...")`
when the primary resolver raised. `_expand_tilde` follows the subprocess-HOME
contract, which under host `auto` mode can be the real/default user home rather
than the active multiplex `HERMES_HOME`. So on the exception path the guards
recreated the very cross-profile authority bug the happy path fixed: beta's
`config.yaml` was compared against the default/root config (hard-block fails
open), and the protected-instruction exemption resolved against the wrong home.

Re-derive both fallbacks from the same `get_hermes_home()` key the happy path
uses (`Path(home)/config.yaml`, `realpath(home)`), and substitute no unrelated
home if the active security path cannot be established — a `None` fails closed
at the protected-instruction consumer (exemption skipped, gate runs).

Adds opposite-side regressions: forcing the primary config resolver to raise
keeps beta's own config refused (and does not spuriously protect alpha's under
beta's scope); forcing the primary home resolver to raise keeps beta's
instruction-file exemption resolved against beta.

Addresses the fallback-authority review on #107335 (thanks @andrexibiza).

(cherry picked from commit 119d88b46f745ba081f12adf3f6ebba457d95d68)
2026-09-11 15:29:15 -07:00
PRATHAMESH75 1271622e4b fix(file-safety): resolve HERMES_HOME/config per call so multiplex profiles don't poison the write guards (#107327)
In a multiplexed gateway (`gateway.multiplex_profiles: true`) each profile turn
scopes `HERMES_HOME` through a per-turn contextvar. But
`tools/file_tools_write_guards.py` memoised the resolved home and config path in
process-global module state, filled once by whichever profile ran first. Both
the protected agent-instruction approval gate (`_get_real_hermes_home` →
exemption for a profile's own home) and the `config.yaml` hard-block
(`_get_hermes_config_resolved`) therefore became order-dependent: a later
profile's own `workspace/AGENTS.md` was gated against a *sibling* profile's home,
and — worse — its own `config.yaml` stopped matching the block, so a
prompt-injected agent could rewrite the very file the block exists to protect
(reproduced end-to-end in #107327).

Resolve both values per call instead. `get_hermes_home()` / `get_config_path()`
are contextvar-scoped, so the getters now track the active profile; the guard
already pays a `realpath` per call, so the extra cost is negligible. The two
module slots are kept purely as a test-override surface (set the slot + its
`_loaded` flag to pin a value); production leaves them unset and resolves live,
which also removes the cross-test poisoning the process memo could cause.

Adds regression coverage: both getters track the active profile after a prior
profile's scope, and the `config.yaml` hard-block fires for beta's own config
even after an alpha turn ran first.

(cherry picked from commit 36b257391da497ac31e6c440727dc57ecd584e71)
2026-09-11 15:29:15 -07:00
infinitycrew39 606903badc fix(tui): bind launch-profile terminal scope once multiplexing is active
After any secondary profile home is served, launch-profile turns used to stay
unscoped and fall back to ambient os.environ. Bind the launch home's own
terminal policy in that case so a poisoned ambient bridge can never become
the launch turn's authority (#107422 residual of #68559).

(cherry picked from commit f81147c1e5d283837e5e27f4da79710da7025235)
2026-09-11 15:29:15 -07:00
infinitycrew39 a5c801c8dc fix(tools): never ambient-bridge TERMINAL_* under a profile home override
A multiplexed dashboard can call _ensure_terminal_env_bridged while a
secondary profile's HERMES_HOME override is active. The one-shot latch then
wrote that profile's docker policy into process-global os.environ and poisoned
later unscoped launch-profile tool calls (#107422).

Skip the ambient bridge whenever a context-local home override is set —
ambient env is launch-profile authority only; routed profiles must use
terminal_scope (same rule as env_loader._reapply_terminal_config_bridge).

(cherry picked from commit 2050efb24fdc9d54a282b24d0042b90f47486c5a)
2026-09-11 15:29:15 -07:00
Chike Okonta 954113839b fix(discord): apply YAML allow_bots to ingress policy
config.yaml discord.allow_bots was accepted but ignored because
_get_allow_bots() only read DISCORD_ALLOW_BOTS. Seed the key through
_apply_yaml_config like the other Discord gates, then resolve via
_gate_raw so env still wins over YAML.

(cherry picked from commit ddeb1bd39253404a3c0b43bc65372e5ada61bbbf)
2026-09-11 15:29:15 -07:00
Teknium d807a34c87 docs(multiplex): cron ticker allowlist, named multiplexer, guild-scoped route delivery 2026-09-11 15:28:37 -07:00
Teknium 1c08edccf0 fix(gateway): a named-profile multiplexer ticks its own cron store and owns the shared adapters
profiles_to_serve(multiplex=True) yields default + allowlist, so a gateway run
as `hermes -p <name>` with multiplex on never ticked its own profile's jobs
unless allowlisted (which would start a second adapter on the same token). The
ticker's home list now unions the active profile. The shared-adapter owner
passed to the ticker is the runner's launch profile instead of the literal
"default", so that profile's jobs reuse its live adapters rather than the
fail-closed empty map.

Co-authored-by: Paul Pincente <101599379+pincente@users.noreply.github.com>
Co-authored-by: r3x443 <325334945+r3x443@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
Teknium a6c5b7ada8 fix(serve): SSH-isolated idle-exit keeps the backend alive while a cron job runs
turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.

Addresses #107485
2026-09-11 15:28:37 -07:00
Teknium 0e57543908 fix(desktop): cron ticker follows the multiplexer allowlist and stands down for served satellites
The Desktop/serve backend ticked every installed profile (ignoring
gateway.multiplex_profile_allowlist) and gated only on the profile's OWN
gateway.pid. A satellite served by the default multiplexer has no pid file, so
both tickers raced its fires and the Desktop one won nondeterministically —
adapter-less standalone delivery, and the environment behind #107485.

Homes now come from profiles_to_serve with the default profile's allowlist
(the multiplexer's served set); the per-tick gate also consults
named_profile_served_by_running_multiplexer.

Addresses #107485, #94590
Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-11 15:28:37 -07:00
Teknium 444fa8166a fix(cron): routed-profile cron delivers through the shared bot for guild-scoped routes and profiles without a platforms block
SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.

_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.

Fixes #89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
tachyon-r 32c538851a fix(desktop): yield single-profile cron to its running gateway
(cherry picked from commit db92ff258f0138a251b2c4e14b32f2fa0b7b8d5b)
2026-09-11 15:28:37 -07:00
fangliquanflq c17629a0a2 fix(cron): scope restart-safe worker environment
(cherry picked from commit e57f719f942736104ee7ff79999f41b8f2a23d63)
2026-09-11 15:28:37 -07:00
Teknium 3efbd79a30 chore: map salvaged contributor emails (remi-td, nmediaie) 2026-09-11 15:28:00 -07:00
Teknium cdedfe9770 test(gateway): invariants for multiplex routing/authz (#104933, #103717)
Five behaviour contracts, each red on origin/main: shared-bot route does not
hijack a dedicated secondary bot; secondary busy follow-up authorized against
its own allowlist; mid-turn authorization reads the admitting transport's
allowlist; shared-bot satellite resolves the primary transport for restored
sources while a downed secondary stays fail-closed; completion pre-flight runs
in the target profile's scope.
2026-09-11 15:28:00 -07:00
Teknium c632437c3b fix(gateway): deliver secondary-profile async completions under their own profile scope
The supervised `_async_delegation_watcher` and startup-recovered process
watchers run under the ROOT scope, so a secondary profile's completion was
classified against the DEFAULT profile's state.db (row absent → "terminal" →
"permanently-gone session" warning, delivery dropped) and every durable-ledger
op (`claim`/`complete`/`release`) hit the default's ledger, stranding the real
row `pending` forever.

`_deliver_completion_notification` and `_deliver_async_delegation_group` now
run their whole pre-flight + claim + inject + settle sequence inside
`_completion_event_scope(evt)` — the runtime scope of the profile the event's
session belongs to. At multiplex startup `_restore_secondary_completion_ledgers`
replays every secondary's pending rows, which the process registry (launch
home only) never saw.

Builds on the classify-only half of #107247.
2026-09-11 15:28:00 -07:00
Teknium 77180acc2c fix(gateway): mid-turn authorization reads the admitting bot's allowlist; shared-bot satellites keep a transport
Inside a routed satellite's turn the ambient scope is the satellite's, whose
.env has no token or allowlist. Five sites still called `_is_user_authorized()`
directly there — `/topic`, the sibling-thread `/stop` grant, plugin message
injection, Discord voice transcripts and startup auto-resume — so the shared
bot's owner was refused ("not authorized to use /topic") and a satellite that
DID copy an allowlist widened who may drive those commands. They now go through
`_is_user_authorized_for_source`, and `_under_authorization_profile` derives the
transport home from the delivering adapter's owner when ingress did not stamp
one (restored/cached sources).

`_adapters_for_profile` (the resolver behind `_authorization_adapter`,
`_adapter_for_source` and now `_resolve_injection_adapter`) returns the primary
map for a shared-bot satellite: a served profile with the `{}` startup
placeholder, no reconnect pending, targeted by a default-bot route. Kanban and
cron already applied that rule; the gateway's own resolvers returned None, so
heartbeats, process completions, goal notices and delegation results for such
profiles were undeliverable after a restart. A secondary that owns a credential
(connected on any platform, or queued for reconnect) still fails closed.
2026-09-11 15:28:00 -07:00
Teknium 7ecee9e78f refactor(gateway): share the primary admit/scope step between message and busy handlers
`_make_default_profile_busy_session_handler` (salvaged from #105357) duplicated
the stamp-transport-home / stamp-route / resolve-home block of
`_make_default_profile_message_handler`. Both now call `_admit_primary_source`,
so the two ingress paths cannot drift apart again (#103717 was exactly that
drift).
2026-09-11 15:28:00 -07:00
Teknium 27768a9fa5 fix(gateway): profile_routes no longer hijack DMs to a dedicated secondary bot (#104933)
Telegram DM chat_id == user_id for every bot, so a `chat_id` route meant to
pin a user's DM with the SHARED bot to profile `ops` also captured that
user's DM with team_b's dedicated bot: the turn, session key, state.db and
secrets were ops's while the reply left via team_b's transport.

`ProfileRoute` gains an optional `bot_profile` discriminator (None = the
default profile's bot) and `matches()` requires it to equal the receiving
adapter's owner profile. `build_source` passes `_owner_profile` and stamps it
as the fallback profile for secondary adapters; `_profile_name_for_source`
derives it from the transport ref for sources built elsewhere.
Secondary→secondary routes remain possible by naming the bot explicitly.

Salvages the discriminator idea from #105040 and #105049 in a smaller shape
(no bot-username sniffing across adapter internals; the owning profile is
already declared at `set_owner_profile`).

Co-authored-by: 0xAlyDev <agentai891@gmail.com>
Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
2026-09-11 15:28:00 -07:00
Hermes Agent 350b9c90fa fix(gateway): resolve authorization routing in multiplexed busy sessions
Route primary-adapter busy callbacks through the same transport authorization scope as normal multiplexed messages, while preserving the routed profile for session state.

Adds a regression test for a Signal primary transport routed to a secondary profile.

(cherry picked from commit 429171ae0dcb78fdd20500b93c415ee13d00091f)
2026-09-11 15:28:00 -07:00
chelsealong fc4f564568 fix(gateway): scope multiplex busy-session handler to its owning profile
_make_profile_busy_session_handler stamped source.profile but never
entered _async_profile_runtime_scope before delegating, so
_is_user_authorized fell through to os.environ and read the primary
profile's allowlist. Follow-up messages from a secondary profile's
owner while their session was busy were dropped as unauthorized. Wrap
the call the same way the cold-path _make_profile_message_handler
already does.

Fixes #103717

(cherry picked from commit 760bc35c60a75fbd2500501d80c8e00d20c951e0)
2026-09-11 15:28:00 -07:00
nmediaie 65579d8a46 fix(gateway): bind HERMES_SESSION_PROFILE per /p/<profile>/ api request
_bind_api_server_session never passed profile= to set_session_vars, so
every API-server turn bound HERMES_SESSION_PROFILE="" and the terminal
tool collapsed ALL api_server sessions (default AND org profiles) onto
the shared 'default' container key. In a multiplex gateway the first
caller after restart therefore won the cache slot and org-profile turns
could reuse the default profile's sandbox container (SSH key and secrets
exposure). Pass the request profile through so org turns key their own
profile-scoped container.

(cherry picked from commit b8b01b3d8cdafab292dda4388592044393222886)
2026-09-11 15:28:00 -07:00
Teknium 37dcc0a6e8 fix(gateway): a secondary API_SERVER_KEY no longer skips the profile; start/install/status honour the live multiplexer
Under gateway.multiplex_profiles the default gateway serves every profile, yet
four startup/status paths still reasoned from the wrong source:

* A secondary profile's API_SERVER_KEY (which the docs REQUIRE for /p/<profile>/
  auth) auto-enabled api_server in that profile's config, so
  _load_secondary_profile_config raised SecondaryPortBindingConfigError and the
  whole profile was skipped. gateway/config_env.py::_enable_from_env now leaves
  `enabled` alone for port-binding platforms while a multiplexer loads a
  NON-default profile (home override + multiplex flag, the same signal
  gateway.config uses for scoped reads); the credential still lands in extra so
  the shared listener can authenticate the prefix. Default profile unchanged.

* "Is this profile served?" was re-derived from the default config.yaml plus
  GATEWAY_MULTIPLEX_PROFILES as seen by the CLI process. `hermes -p coder ...`
  loads coder's .env, so an env-only opt-in on the default profile was invisible
  (guard never fired, status said stopped) and an allowlist edit flipped the
  answer before the restart. named_profile_served_by_running_multiplexer now
  reads the pid-verified default gateway_state.json served_profiles (written by
  _record_served_profiles) first and falls back to config derivation only when
  the key is absent. The record helpers live in hermes_cli/gateway_multiplex_served.py.

* The served-profile guard ran only inside `gateway run`. `hermes -p X gateway
  start|install|restart` reached the service manager, whose unit then exited 78
  forever (systemd parks it while the CLI prints "started"; launchd KeepAlive
  respawns every 30 s). The service verbs now run the same guard up front
  (exit 78, same message) and accept --force; the Desktop /api/gateway/start
  route returns 409 for a served profile instead of spawning a doomed child.

* Status surfaces disagreed: `hermes -p coder status` said stopped, `hermes -p
  coder cron status` said "cron jobs will NOT fire" while `cron list` said fine,
  and the default `hermes status` never listed served profiles. Both now route
  through the probe / the recorded served set. The -p/--profile matcher in
  _scan_gateway_pids and gateway.status._command_line_belongs_to_profile compares
  the flag token for equality (`-p ops` no longer claims -- or lets `gateway
  stop` SIGTERM -- an `-p ops-2` gateway).

Docs: multi-profile-gateways.md now describes the start/install refusal, the
--force flags, the API_SERVER_KEY behaviour and the single default-home
gateway_state.json (the per-profile runtime_status.json claim was wrong).

Fixes #100397
Addresses #89726 #97360 #71344

(cherry picked from commit d002c1864a7b6a22c53758b16b7b0cc79aea2edf)
2026-09-11 15:27:23 -07:00
Teknium ceaf622c6d fix(mcp): same-named MCP servers with different credentials connect per profile; owner /reload-mcp keeps adopters' tools
Under gateway.multiplex_profiles every connection ledger in tools/mcp_tool.py
(_servers, _server_scope_keys/_server_tool_scopes, connecting/error/cooldown
maps, the circuit breaker, lazy schema-cache configs, trust metadata) was keyed
by the bare server NAME. The common per-tenant layout — each profile names its
server `github`/`notion` with its own token — gave only the first profile a
connection: the second profile's register_mcp_servers saw the name as "already
connected", refused to adopt it (different credentials, 4ddbcbd35e), and left
the profile silently tool-less with a healthy-looking `configured` status
(#106005 Bug 1/2, #91654). Siblings of the same bug: profile A's failing `x`
put profile B's healthy `x` into A's 10-minute connect cooldown and A's open
circuit breaker short-circuited B's calls; toolsets._resolve_toolset_memo was
not scope-keyed, so B resolved A's `mcp-<server>` tool names.

Keys are now the connection key from the new tools/mcp_tool_scope.py: the bare
name outside a multiplexer (single-profile processes are unchanged) and
(owner_scope, name) under one. Call-time lookups (_resolve_server_key) prefer
the calling scope's own connection, then a shared connection it adopted, so
identical-route profiles still share one subprocess. _select_new_servers,
the cooldown/breaker/trust maps, lazy registration and get_mcp_status all read
and write through the composite key; teardown resolves a task's key by
identity (the MCP loop has no profile context). The toolset memo key includes
registry.current_scope_key().

An owner's scoped /reload-mcp tore down its connection and, with it, every
adopting profile's tool overlay; nothing re-ran the adopters' discovery until
they reloaded. shutdown_mcp_servers(scope=) now records the orphaned adopters
and register_mcp_servers re-registers them under their own home + secret scope
once the owner's rediscovery pass completes.

Docs: multi-profile-gateways.md states the per-profile connection rule.

Fixes #106005
Fixes #91654
Co-authored-by: Bergmann89 <info@bergmann89.de>
Co-authored-by: Izzy-Gottz <srulynj@gmail.com>
2026-09-11 15:27:23 -07:00
Teknium 5e04961072 test(hindsight): trim salvaged #103957 tests to the three invariants
Keep: scopeless worker resolves the on-disk key (durability core), a keyed
build still rewrites (rotation), and the API-prefixed vault name resolves.
Drop the mock-patched gate predicate test, the inlined copy of the same
predicate, and the empty-build/empty-disk arm — they re-assert the helper's
shape rather than a behaviour contract.
2026-09-11 15:26:46 -07:00
Teknium a9838c2100 fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.

Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.

Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.

Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.

Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.

Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.

Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.

session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.

Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.

Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
2026-09-11 15:26:46 -07:00
이민재 a1b689f12a fix(mem0): skip platform key lookup in OSS mode
Conflict resolution on top of main also routes MEM0_MODE / MEM0_HOST / MEM0_AGENT_ID /
MEM0_USER_ID through get_secret: identity and host are .env values like the key, and a
raw environ read hands a secondary profile the default profile's mem0 account.

(cherry picked from commit 7a5863ebafdb193d8ec2f89a7aca309774d6d69a)
2026-09-11 15:26:46 -07:00
red4711 778ad61017 fix(hindsight): accept HINDSIGHT_API_LLM_API_KEY vault name in key resolution
The vault item is named HINDSIGHT_API_LLM_API_KEY (matching the daemon env var), while the setup wizard uses HINDSIGHT_LLM_API_KEY. Vault-fed scopes carrying only the API-prefixed name resolved empty, so the scopeless path — and the whole fail-closed machinery — never engaged on real deployments. Accept both names.

Proven end-to-end: hydrated scope now resolves the 73-char key and the rewrite gate passes.

(cherry picked from commit 2ce8c2980188ae8e4a3ed398e14fa8ed89a40644)
2026-09-11 15:26:46 -07:00
red4711 d0d5c7745a test(hindsight): run pure gate tests on Windows too
test_gate_refuses_keyless_build_against_keyed_disk and test_rewrite_allowed_when_build_carries_key touch no filesystem modes, so the nt skip guard was needlessly excluding them. Addresses review feedback on PR #103957.

(cherry picked from commit 196cf723571c87829be61e17df27869948db00a0)
2026-09-11 15:26:46 -07:00
red4711 2c24ed4b6e fix(hindsight): keep scopeless daemon worker from clobbering profile env key
The embedded daemon-start worker usually runs with no secret scope, so _embedded_llm_api_key resolved empty and both the profile-env compare and the upstream manager merge overwrote the file's good key with emptiness.

- _embedded_llm_api_key: explicit config, then secret scope, then on-disk profile env fallback; UnscopedSecretError treated as empty without touching os.environ (multiplex safety).
- Add _may_rewrite_profile_env gate: refuse rewrite on keyless-build plus keyed-disk, and skip the daemon stop in that case.
- Document key resolution order in plugins/memory/hindsight/README.md.
- Add 5 tests pinning disk fallback, the gate's refuse branch, rotation, keyless providers, and file survival.

(cherry picked from commit 48a4bcd536fa799b666a8d5eaae73c1816ca9e33)
2026-09-11 15:26:46 -07:00
Teknium 915c0f806d test(goals): patch the registry construction seam, not hermes_state.SessionDB
goals.py now acquires its store through hermes_state_registry (one shared
writer per home), so the off-loop bootstrap tests' fake must be installed at
hermes_state_registry._open_session_db — the seam the registry documents for
tests — and accept the db_path the registry passes.
2026-09-11 15:26:09 -07:00
Teknium 956793aeac docs(multi-profile): profile sessions keep their own keys and store on every gateway path 2026-09-11 15:26:09 -07:00
Teknium 6806a38084 fix(goals): GoalManager borrows the registry's shared state.db handle
hermes_cli/goals.py cached a bare SessionDB() per HERMES_HOME. The path was
right per profile, but it was a SECOND writer connection beside the
gateway's registry handle for the same file — its own token-writer thread
and close-time checkpoint the registry knows nothing about (the #90837
corruption shape), doubled per profile under multiplexing. This is the
same topology #100682 removed from the API server. Acquire through
hermes_state_registry and drop lost-race references with release_or_close.
2026-09-11 15:26:09 -07:00
Teknium e7136f1694 fix(tui_gateway): profile sessions never read or write through the launch store
Three wrong-store paths in the TUI/Desktop backend under multiplexing:

- insights.get was scoped=True and then called _get_db(): before the launch
  handle was pinned (#102526, #108074) a scoped first touch bound the
  process-wide handle to the requested profile's state.db; even pinned, the
  RPC answered for the launch profile whatever `profile` said. Route it
  through _profile_db so it counts the requested profile's sessions.

- prompt.background side agents were handed the launch handle, so a
  named-profile Bot Chat's bg_* rows appeared in the default profile's
  session list and were missing from the profile's own history. Inherit the
  parent agent's dedicated _session_db.

- session_lifecycle's gateway-owned-source guard and the notification
  poller's compression-tip resolver looked the session up in the launch
  store, where a named-profile row does not exist: the guard was dead and
  a compression-rotated profile session lost every post-compression
  delegation/background completion (the fail-closed owner gate never
  matched the compressed parent's key). Both now use the session's own
  store via _session_db(session).

Addresses #102157, #102526.
2026-09-11 15:26:09 -07:00
Teknium 75ae2859b9 fix(gateway): default-profile rows and /undo eviction stay on the profile that owns them
Two multiplex key/store mismatches in the gateway session layer:

- SessionStore._db_for_key resolved legacy agent:main keys to the AMBIENT
  store. Named-profile keys were already pinned to profiles/<x>/state.db
  (#97309), but a default-profile session touched inside a secondary's
  runtime scope (a coalesced async-delegation drain, a cron mirror into a
  default chat, a scoped watcher) was written into the secondary's state.db
  — the profile_name='<A>' row inside B's store reported in #102157, after
  which the routing index and row disagree and the #54878 self-heal drops a
  live conversation. Default keys now resolve to _routing_home/state.db, the
  launch home already pinned for the routing index. Single-profile gateways
  (multiplex off) keep the ambient store byte-for-byte.

- /undo evicted the cached agent under build_session_key(source) —
  agent:main:… — while the cache is keyed by the profile namespace, so on a
  secondary profile the eviction missed and the next turn reused an agent
  whose in-memory history still held the undone turns. Use
  _session_key_for_source like /reset and /retry.

Addresses #102157.
2026-09-11 15:26:09 -07:00
Teknium 5fc784a8cf fix(gateway): shutdown notices parse named-profile keys directly
With _parse_session_key accepting the agent:<profile>: namespace (salvaged
from #105942), the rewrite-to-agent:main workaround in
_shutdown_notification_target is dead weight: take the profile from the
parsed result like _build_process_event_source now does.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-11 15:26:09 -07:00
LovePlayCode adc4fad919 fix(qqbot): authorize approval buttons for named-profile session keys
`_parse_gateway_session_key()` required the literal `main` in the
namespace slot of `agent:<ns>:<platform>:<chat_type>:<chat_id>`, but
`_session_key_namespace()` (gateway/session.py) puts the multiplex
profile name there for named profiles. Every approval button click in a
named-profile QQ session therefore failed authorization ("Rejected
unauthorized approval click") and the pending approval timed out
fail-closed.

The namespace slot takes no part in any authorization decision, so the
parser now accepts any non-empty namespace; the platform and operator
checks are unchanged and pinned by new negative tests.

Fixes #98292

(cherry picked from commit 8e0f02ee770bcf931dd0adc8464f0edc185d2546)
2026-09-11 15:26:09 -07:00
Konstantin Khlopkov bdb38e3fc6 fix(gateway): parse named-profile namespaces in session keys
_parse_session_key() only accepted the agent:main namespace, so under
multiplex_profiles named-profile keys (agent:<profile>:...) parsed to None and
notification/shutdown routing fell back to the LRU source cache; and
_sibling_thread_run_keys() hardcoded agent:main, so a per-user thread /stop
never matched a sibling run on a named profile. Accept any valid profile-id
namespace slot and report it as profile (main keys keep their exact historical
shape), build the sibling prefix from _session_key_namespace(source.profile),
and drop the manual agent:main re-wrap in _build_process_event_source().

(cherry picked from commit 562792015a60da3c675498edc2552a6234015457)
2026-09-11 15:26:09 -07:00
ethernet 31d0a2428e fix: add MagicMock to gitignore 2026-09-11 16:11:50 -04:00
ethernet 35a9d6b39d fix: delete magicmock crud 2026-09-11 16:11:50 -04:00
Teknium 939e45c91d chore: release v0.21.2 (2026.9.11) 2026-09-11 12:20:08 -07:00
Teknium 759024bdff fix(deepseek): honor Flash 1M window leftovers and native vision
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.

Two catalog gaps still produced the reported symptoms:

- A leftover context_length_cache.yaml entry of 128K (the old
  ``deepseek`` catch-all) outlived the 1M catalog keys because
  deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
  image routing falls through to lossy text. Vendor docs (2026-09-10)
  mark deepseek-flash as vision-capable and deepseek-v4-pro as not.

Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
2026-09-11 12:09:41 -07:00