Commit Graph

28271 Commits

Author SHA1 Message Date
Ayush Nangia 9189e42842 fix(desktop): reconnect an active profile whose gateway is closed 2026-09-02 06:47:55 -07:00
Teknium dfa74b5815 fix(desktop): boot session pop-out/watch windows against the session's owning profile
`openSessionInNewWindow` → IPC `hermes:window:openSession` →
`buildSessionWindowUrl` emitted no `profile`, so a secondary window (⇧⌘-click
pop-out, subagent watch) was a full renderer that adopted the PRIMARY
backend's profile and resolved the session id against the wrong store —
blank/wrong session for any non-primary profile (#82768, #61286).

The owning profile now rides the URL as `&profile=`, exactly the carry the
HUD already does (buildHudWindowUrl / windowProfileOverride in
use-gateway-boot); the renderer picks it with the same ladder openHud uses:
the session's stamped owner wins, an unstamped/uncached id (a brand-new
subagent child) inherits the profile the user is looking at.

Diagnosis credit: @DomGrieco (#82794).

Co-authored-by: DomGrieco <6556434+DomGrieco@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
liuhao1024 74b00d7a97 fix(dashboard): route every per-row session request at the row's owning profile (salvage #99387)
The Sessions page listed rows stamped with their owning profile but sent
delete/bulk-delete to the global management profile, which stays "" while
the sticky active profile equals the dashboard process's own — so the
request opened the process store, missed, and returned a false
`already_absent` success while the row survived in profiles/<p>/state.db.

Same class at three sibling sites the PR didn't touch: renameSession,
exportSessionUrl and the expanded-row getSessionMessages read. One
`rowProfile(id)` owner now feeds all four (+ bulk delete); unstamped rows
(search results) fall back to the management profile as before.

Test trimmed to one jsdom scenario driving all four row actions.

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-02 06:47:55 -07:00
Teknium 2ce6f5c538 fix(desktop): refresh routed-profile Bot Chat transcripts (salvage #99333)
Two sites dropped profile ownership on the tile-transcript refresh path:

- tui_gateway/server.py _sessions_sig statted only the launch home's
  state.db, so a turn landing in a served sibling profile's store never
  produced sessions.changed. The watcher now also probes every profile
  home _profile_home() has resolved for this backend (empty set on
  single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
  transcript unscoped; it now passes the tile's ownerRoute scope, the
  same way reconcileActiveTranscript already does for the main pane, and
  keys the change signature by owner.

Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).

Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
SZWzz ff8c1aca7a test(desktop): cover combined remote route precedence 2026-09-02 06:47:55 -07:00
SZWzz ae81786580 fix(desktop): let a per-profile remote override win over the forced-local route (#90477)
resolveRegistryLocalRoute collapsed globalRemote and profileRemoteOverride
into one forced-local branch. The two cases are different:

- globalRemote: forcing "This device" to spawn genuinely-local children is
  the intended migration behavior — unchanged.
- profileRemoteOverride: the per-profile SSH/remote override is an explicit,
  authoritative routing decision for that profile. Forcing local made the
  roster enumerate the profile via its override but open the thread in a
  forced-local child, which dies with 'Profile "x" no longer exists' when
  the profile only exists on the remote — reproduced on a macOS Desktop in
  global SSH mode where mythony-agent/q-agent exist only on the NAS.

The registry 'local' entry now delegates to the legacy profile route when a
per-profile override is present, so the override stays authoritative.

Tests: the override case now pins delegation, and a new witness pins that
globalRemote alone still forces local; both contracts are asserted together.
88 connection-registry + 91 remote-lifecycle + 73 routing tests pass;
tsc --build clean.
2026-09-02 06:47:55 -07:00
Teknium 3fcfa647ed docs(google-chat): document per-profile scoping and ADC fail-closed under multiplex 2026-09-02 06:47:45 -07:00
pierrenode 3245264668 fix(relay): carry routed profile through the passthrough-plane forward
The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.

Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.

Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
2026-09-02 06:47:45 -07:00
Kong dca9a54427 fix(gateway): match WhatsApp profile_routes across JID/LID/number forms
`ProfileRoute.matches()` compared `chat_id` as an exact string, so a
WhatsApp route written as a phone number never matched the JID
(`…@s.whatsapp.net`) or LID (`…@lid`) the bridge actually delivers, and
the inbound fell through to the default profile. Allowlists and session
keys already canonicalize these via `gateway.whatsapp_identity`.

Exact compare still wins first; only whatsapp / whatsapp_cloud *user*
chats get the alias intersection fallback (applied to both `chat_id` and
`parent_chat_id`). Groups, broadcasts and every other platform stay exact.

Salvaged from #85081 (Kong); parent_chat_id fallback added on top.
2026-09-02 06:47:45 -07:00
Hudson db639e1023 fix(gateway): bind every adapter to the runner at the _create_adapter boundary
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.

Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.

Salvaged from #70831 (Hudson). First reported in #68332.
2026-09-02 06:47:45 -07:00
Jony b1bc9bb650 fix(google-chat): scope multiplex profile config and fail ADC closed
Route every GOOGLE_CHAT_* / GOOGLE_APPLICATION_CREDENTIALS read through a
module-local `_get_scoped_secret` (scope-authoritative under multiplex,
os.environ fallback only for the unscoped default-profile constructor, so
startup/reconnect never hits UnscopedSecretError — #70652 class). Snapshot
Pub/Sub callback knobs on the instance while the scope is still installed,
and seed them into `extra` from `_env_enablement`.

When a scoped profile has no service-account setting, do NOT fall through
to google.auth.default(): ADC reads the process env directly and would
authenticate the profile as another profile's SA. Fail closed with an
explicit error (adapter and standalone send).

Also resolve the bot-id cache path at call time via get_hermes_home() so
profiles don't share one identity cache.

Fixes #73439.
Salvaged from #73445 (Jony) with the ADC guard from #57674 (Ray, first submitter).

Co-authored-by: Ray <rayjun0412@gmail.com>
2026-09-02 06:47:45 -07:00
SHT 14f20d142e fix(feishu): isolate lark_oapi WS globals per profile and supervise the client thread (#73779)
lark_oapi.ws.client keeps the asyncio loop used by Client.start() in a
module-level global, and Hermes monkey-patched websockets.connect on the
shared module. Under multiplex every profile runs its own WS client on a
dedicated thread, so N threads overwrote each other's globals
(last-write-wins): clients scheduled tasks on a sibling's loop ('Future
attached to a different loop') or bound to the wrong loop and went deaf.

Install process-wide shims once: the module loop becomes a proxy that
forwards to the calling thread's registered loop, and websockets.connect
a dispatcher merging the calling thread's ping overrides. Add a per-adapter
supervisor: the executor future was only awaited by disconnect(), so a
dead WS thread left the profile silently deaf; now it is rebuilt with
capped exponential backoff while the adapter is meant to be connected.

Salvaged from PR #84165, trimmed: the legacy fallback path when the shim
cannot install was dropped (the shim only touches two attributes the
adapter already depended on); tests reduced to three. The loop-isolation
approach was first proposed by @zmlgit in #64247/#69904.

Co-authored-by: zmlgit <6995990+zmlgit@users.noreply.github.com>
2026-09-02 06:47:30 -07:00
SHT 9f40d3bdd0 fix(feishu): resolve DM admission config per-profile under multiplex (#86905)
Route FEISHU_ALLOW_BOTS, FEISHU_GROUP_POLICY, FEISHU_ALLOWED_USERS,
FEISHU_BOT_*, FEISHU_APP_ID and FEISHU_REQUIRE_MENTION through
_get_scoped_secret, and snapshot FEISHU_ALLOW_ALL_USERS /
GATEWAY_ALLOW_ALL_USERS into settings.allow_all_dm at construct time so
_admit (running on the lark_oapi WS thread, no secret scope) no longer
reads the default profile's os.environ. Feishu open_ids are app-scoped,
so a secondary bot's DMs were always dm_policy_rejected against the
default profile's allow-list.

Salvaged from PR #86908 trimmed to the adapter call-site migration; the
fail-closed _get_scoped_secret / _auth_env rewrites are already on main
(2912c36aa4, agent/secret_scope.py get_secret policy).
2026-09-02 06:47:30 -07:00
Teknium 5b4cff976d fix(gateway): fingerprint Teams client_id and WeCom bot_id for the multiplex credential guard
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
2026-09-02 06:47:30 -07:00
webtecnica 97b8a97861 fix(feishu): include app credentials in multiplex fingerprint (#76793) 2026-09-02 06:47:30 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium d45bc39667 fix(gateway): resolve the ephemeral personality prompt per turn from the scoped profile
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.

Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.

Fixes #89161

Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
StanleyStetson 94a2d7f8af fix(gateway): persist slash-command config writes into the routed profile
Slash dispatch already runs inside _profile_runtime_scope under multiplex,
but _save_gateway_config_key (/reasoning --global, /fast, show/hide),
/memory approval, /skills approval, /verbose and /footer built their write
path from the module constant gateway.run._hermes_home — the launch home —
so a routed profile's toggles landed in the default profile's config.yaml
while the reads (via _gateway_config_home()) saw the routed one.

Resolve the write path through _gateway_config_home() at all five sites so
reads and writes agree. Single-profile gateways never install the override
and keep resolving the launch home.

Fixes #87939
Fixes #75684

Co-authored-by: Bao <nnqbao@gmail.com>
2026-09-02 06:36:16 -07:00
Teknium 257da5ca07 fix(gateway): route /review and /reload-skills executor hops through the scoped helper
Sibling sites of the bare loop.run_in_executor(None, …) class fixed for
/insights, /debug and /goal draft: the worker started with an empty
context, so get_hermes_home()-relative reads (skills.external_dirs,
disabled skills, the reviewer subagent's home and secret scope) resolved
the launch home instead of the routed profile under multiplex.
2026-09-02 06:36:16 -07:00
Drexuxux fbd9730e30 fix(gateway): run /insights, /debug and /goal draft inside the routed profile
The multiplexed inbound handler wraps every message in _profile_runtime_scope,
which installs the routed profile's HERMES_HOME override and its secret scope
as contextvars. A bare loop.run_in_executor(None, fn) starts the worker with an
EMPTY context, so neither reaches the blocking work.

GatewaySlashCommandsMixin already knows this -- /compress goes through
_run_in_executor_with_context and the call site says why. Three siblings in the
same file still used the bare hop:

  /insights   SessionDB() with no explicit path resolves get_hermes_home() at
              call time (_default_db_path), so the worker opened the DEFAULT
              profile's state.db. Under multiplexing the command reported
              another profile's conversations, session counts and sources to
              this profile's user.

  /debug      collects that home's logs/config and uploads them to a public
              paste, so it published the default profile's diagnostics from
              another profile's chat.

  /goal draft calls the auxiliary LLM, whose provider/credential resolution
              reads the profile secret scope -- unscoped it falls back to
              process-global os.environ, which under multiplexing may hold a
              different profile's keys.

Route all three through _run_in_executor_with_context.

/reload-skills is deliberately left alone: tools.skills_tool binds SKILLS_DIR
at import time, so it does not follow the contextvar either way. Fixing that
needs the module-global retarget web_server._profile_scope performs under a
lock, which is a different change from context propagation.

Single-profile gateways never enter the scope, so their behaviour is unchanged.
2026-09-02 06:36:16 -07:00
Teknium eed481bd34 docs(profiles): cron delivery for routed profiles rides the shared bot only for exact routed targets (#101113) 2026-09-02 06:27:24 -07:00
Teknium a6351a71e5 fix(cron): route a credentialless satellite's cron delivery through the primary adapter for exact profile_routes targets (#101113)
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.

- cron/scheduler.py: factor the preflight's primary-config route loader into
  `_primary_profile_routes_for_current_home()` (one owner for both halves,
  so route semantics cannot drift) and add `SharedRouteAdapters`, a
  read-only view over the primary adapter map that resolves an adapter for
  a (platform, target) ONLY when an enabled primary route with a
  chat_id/thread_id maps that exact target to the current profile —
  using the same `ProfileRoute.matches` predicate as inbound routing.
  `_deliver_result` resolves the transport per target from it; everything
  else (unmatched chat, disabled route, route for another profile, no
  primary adapter, guild-only route) is a miss and never uses the primary
  bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
  gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
  fallback: with no matching route the view is falsy and delivers nothing.

Fixes #101113
2026-09-02 06:27:24 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Cyber-Yichen a2fea79de6 fix(cron): isolate multiplex profile failures per profile (#74878)
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:

- startup recovery loop: a per-profile exception (e.g. an unreadable
  executions.db raising sqlite3.DatabaseError) was uncaught and killed the
  ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
  exception escaped to the cycle-wide handler, skipping every remaining
  profile that cycle and marking all of them failed.

Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).

Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).

Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
이민재 e48bb828d4 test(cron): preserve explicit suggestions path override 2026-09-02 06:27:24 -07:00
wanliqin 9da8842585 fix(cron): resolve profile store paths per call
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.

Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 0fd9218e5a docs: ${VAR} config refs resolve per-profile under a multiplexed gateway 2026-09-02 06:19:03 -07:00
Teknium 638c1f3204 chore(contributors): map patryk.kopycinski@elastic.co -> patrykkopycinski 2026-09-02 06:19:03 -07:00
Teknium b6f0106602 docs(design): note loader-boundary dotenv guard, scoped ${VAR} expansion and scoped .env publish in multiplexing design doc
Accuracy pass on the #89950 doc for the fixes landing alongside it (#77562, #84079, #88441).
2026-09-02 06:19:03 -07:00
Eva 011a60b7cc docs(design): add the multiplexing-gateway design doc referenced by secret_scope
agent/secret_scope.py has pointed at docs/design/multiplexing-gateway.md
("Workstream A") since the fail-closed secret scope landed, but the file was
never added. This writes the missing doc from the code as it stands today:
the mode flag, scope composition (_profile_runtime_scope seams), the
context-local secret scope and HERMES_HOME override, routing/serving/
persistence/session-lane isolation, the control-plane RPCs, failure modes,
and an honest table of what is still process-global (per-profile MCP
registries are tracked in #67605).

Doc-only change; no code touched. Style follows docs/profile-routing.md.
2026-09-02 06:19:03 -07:00
Teknium 2475335443 fix(config): keep .env publishes inside the routed profile scope under multiplex (#88441)
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.

Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
2026-09-02 06:19:03 -07:00
Teknium b53c50cf8a fix(config): resolve ${VAR} config refs through the profile secret scope (#84079)
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).

Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
2026-09-02 06:19:03 -07:00
webtecnica d7dc75ccff fix(gateway): multiplex must not apply default-profile creds to unconfigured profiles (#84079)
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.

Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.

Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Patryk Kopycinski 1eb71756af toolchain-self-improve: route API_SERVER_KEY to profile env 2026-09-02 06:19:03 -07:00
Teknium 0a6aa7cce1 fix(env_loader): log routed-scope dotenv skip once per home; port single-profile control test
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.

Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Lester Liang 1aa62ceb45 fix(security): isolate multiplex dotenv reloads 2026-09-02 06:19:03 -07:00
Teknium 0058bde251 fix(bot-relay): a DM into a live Bot Chat queues as the next turn, never interrupts the one in flight 2026-09-02 06:18:55 -07:00
Teknium d29a7936e4 fix(bot-mode): DMs to a Desktop-owned Bot Chat land in the live session instead of being dropped (#100523)
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.

bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.

On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.

Closes #100523
Supersedes #100544, #100542

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 06:18:55 -07:00
Teknium 16370ae539 fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.

`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.

The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.

Supersedes #94147 (python half)

Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
aeonsong 8076c78c87 fix(desktop): scope handoff config to session profile 2026-09-02 06:17:47 -07:00
Teknium 7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Teknium af86ad0479 fix(doctor): reuse resolved memory config at Memory Provider section
Follow-up to salvaged #100677: the file checks read the memory section via the
run's hermes_home while the Memory Provider section re-read config with no
argument (module-global HERMES_HOME). Resolve once and reuse so both sections
report against the same config.
2026-09-02 06:17:25 -07:00
fangliquanflq 21cebbfd68 fix(doctor): honor disabled built-in memory stores
Salvaged from #100677. Fixes #100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
2026-09-02 06:17:25 -07:00
Teknium 70dc1606c6 test(cli): pin OSC 9 / Warp OSC 777 bell emitters; docs + contributor mappings
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
  only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
  of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
2026-09-02 06:17:10 -07:00
Teknium 632078bca7 feat(cli): OSC 9 + Warp OSC 777 notifications ride on the bell flags
Extend _ring_bell() so display.bell_on_prompt / bell_on_complete also
emit terminal-native desktop notifications from the same six call sites
(clarify, clarify batch, approval incl. computer_use, sudo password,
secret capture, turn complete). No new config keys.

- OSC 9 (ESC ] 9 ; body BEL): Ghostty / iTerm2 / Kitty / WezTerm raise an
  OS notification; unknown terminals drop it. Body is "Hermes: <context>"
  with C0 controls and DEL stripped. Written to /dev/tty (prompt_toolkit's
  stdout wrapper can buffer/strip raw escapes) with a sys.stdout fallback.
- Warp OSC 777 warp://cli-agent (agent "hermes", event permission_request
  / stop, compact JSON mirroring build-payload.sh). Gated on
  TERM_PROGRAM=WarpTerminal + WARP_CLI_AGENT_PROTOCOL_VERSION + the
  should-use-structured.sh broken-build floor (stable/preview builds at or
  before v0.2026.03.25.08.24.*_05 rejected). Never raises.

Salvages #58957 and #100805.

Co-authored-by: glitchbunny0 <glitchbunny0@proton.me>
Co-authored-by: harsha-usethread <harsha@usethread.io>
2026-09-02 06:17:10 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
Teknium 11f932c935 fix(slack): interactive-caller and pre-fetch auth prefer the injected profile check; gate reads never fall through to os.environ
`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.

- Prefer the wired `set_authorization_check` callback (profile-bound
  `_make_adapter_auth_check`) at both sites; keep `__self__` introspection
  only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
  (scoped miss under multiplex → "", never os.environ); drop the raw
  `os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.

Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.

Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Teknium bbb087f3d1 fix(gateway): egress adapter and channel directory no longer follow the per-turn active profile
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.

`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.

Extracted from #87240 (topic-table half handled separately via #76487).

Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
2026-09-02 06:08:09 -07:00