Commit Graph

27289 Commits

Author SHA1 Message Date
Teknium ca42d7a034 fix(gateway): isolate /voice state and voice-channel input per multiplexed profile (#84872)
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.

- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
  `<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
  (`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
  now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
  transcript handler (functools.partial); used at primary connect,
  primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
  `_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
  `_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
  primary path and both reconnect paths already did.

Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium e8231f01da fix(gateway): hydrate secondary-profile secrets off-loop at startup too (#99519 class)
_run_secondary_profile_reconnect now pre-hydrates external secret sources
in a worker thread (PR #99519); _start_one_profile_adapters entered
_profile_runtime_scope on the event loop three times per profile, each
running the same synchronous network-bound hydration under
_SECRET_SOURCE_CACHE_LOCK. Hydrate once via asyncio.to_thread and enter
every scope in that method with hydrate_secrets=False.

The reconnect test is parametrized over both entry points; the startup
reconnect handoff waits are deadline-based since the runner now hops to a
worker thread before publishing the replacement adapter.

Co-authored-by: GoBeromsu <37897508+GoBeromsu@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
pierrenode 956f493cf1 fix(plugins): route platform_actions through profile-aware adapter resolution
PlatformActions._resolve_adapter() unconditionally read
runner.adapters — the default profile's adapter registry — with no
awareness of multiplex/Team-Gateway secondary profiles. Every other
adapter-resolution path in this codebase (GatewayAuthorizationMixin.
_authorization_adapter, and the plugin message-injection path that
uses it) is careful about this split: a secondary profile's adapters
live in runner._profile_adapters[profile], and a stamped secondary
profile with no registry entry must fail closed rather than fall back
to the default profile's adapter of the same platform — falling back
would send replies out the wrong bot.

ctx.platform_actions had none of that: a plugin scoped to one profile
calling add_reaction/set_thread_title in a Team-Gateway deployment
would silently act through the DEFAULT profile's adapter/bot identity
instead of its own profile's, whenever the default profile also ran
that platform.

Route _resolve_adapter() through the same _authorization_adapter
lookup, resolving the calling profile via
hermes_cli.profiles.get_active_profile_name() (which reflects the
per-task HERMES_HOME override multiplex profiles already propagate).
Falls back to the old bare runner.adapters lookup only when the
runner predates _authorization_adapter (defensive, not expected in
practice).
2026-09-02 06:59:11 -07:00
Fangliquan df50b5287e fix(gateway): coerce profile route IDs to str for YAML ints 2026-09-02 06:59:11 -07:00
gobeumsu 04836886fa fix(gateway): run secondary-profile adapter auth setup off the event loop 2026-09-02 06:59:11 -07:00
Teknium b1193b27e2 test(desktop): corrupt-exe integrity test follows stage-and-swap — corrupt pack lands in staging, live app untouched
The Windows-only test still modelled the old in-place pack (corrupt exe already
at release/win-unpacked + a .bak to roll back to). Under stage-and-swap the pack
writes into -c.directories.output=<staging>; the integrity gate runs there and a
failure discards staging while the live tree is never touched. Verified on Linux
with sys.platform patched to win32 inside the test; sabotage (skip the discard)
fails it.
2026-09-02 06:53:20 -07:00
Teknium 8ecf19b328 chore: map contributor deathxdefeat@users.noreply.github.com -> @deathxdefeat 2026-09-02 06:53:20 -07:00
Teknium 1398c0f5ca fix(update): stage-and-swap the Desktop rebuild so a failed pack never removes the working app
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).

Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.

- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
  `_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
  retry purge and the integrity self-heal only ever clear the staging tree,
  never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.

Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.

Closes #86443

Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
2026-09-02 06:53:20 -07:00
Teknium 87ad7aa0b0 fix(gateway): route multiplexed profile logs to their own logs/ dir
setup_logging(mode="gateway") binds agent.log / errors.log / gateway.log
to the launch home, so under multiplex_profiles every secondary profile's
records (emitted inside _profile_runtime_scope) landed in the DEFAULT
profile's files. Main already has the routing primitives from #99440
(record.hermes_home factory + _ProfileRoutingFileHandler +
enable_profile_log_routing) but only the Desktop cron ticker used them.
Enable them at gateway startup for the served profile set; single-profile
gateways are untouched (routing is a no-op below two homes).

Supersedes #84954, which introduced a parallel routing mechanism; the
cross-profile isolation test is translated from its suite.

Co-authored-by: Michał Dziwisz <michal@dziwisz.net>
2026-09-02 06:48:31 -07:00
fangliquanflq 04c640a183 fix(gateway): honor secondary adapters under profile runtime scope
Multiplex turns enter _profile_runtime_scope, so get_active_profile_name()
equals the secondary profile and the old authz lookup consulted empty
self.adapters. Prefer _profile_adapters before the active-profile shortcut.
2026-09-02 06:48:31 -07:00
fangliquanflq 8167cfa4b6 fix(a2a): scope multiplexed peer authorization
ThreadingHTTPServer request threads do not inherit the gateway's profile
ContextVars, so security.authenticate()/is_trusted_peer() read the
process-global A2A_* env on every request — every secondary profile's
listener authenticated against the default profile's tokens. Capture an
immutable A2ASecurityContext at adapter construction (which runs inside
_profile_runtime_scope for secondary profiles) and have the request
handler consult it instead of re-reading env per request.

Salvage note: the original `get_secret() / except UnscopedSecretError:
os.getenv()` fallthrough in _startup_env was replaced by _profile_scoped()
gating (Buzz/Raft pattern) — inside a secondary profile's scope the scope
is authoritative and a miss never falls through to os.environ.

Salvaged from #80956.
2026-09-02 06:48:31 -07:00
Teknium 64bbbaa896 fix(gateway): honor explicit msgraph_webhook disable under env override
Sibling of the webhook branch fixed in the previous commit (#85637): the
MSGRAPH_WEBHOOK env-override branch force-set enabled=True whenever
MSGRAPH_WEBHOOK_ENABLED was truthy, ignoring an explicit
`platforms.msgraph_webhook.enabled: false` in config.yaml. Apply the same
_enabled_explicit guard; port/client_state extras still wire through.
2026-09-02 06:48:31 -07:00
webtecnica dd666d2418 fix(gateway): honor explicit webhook disable in _apply_env_overrides
The webhook-specific env block set config.platforms[Platform.WEBHOOK].enabled
= True directly whenever WEBHOOK_ENABLED was truthy, without honoring the
_enabled_explicit marker that the YAML merge sets for profiles that pin
platforms.webhook.enabled: false. In multiplex mode a secondary profile that
explicitly disables webhook still inherits the process-level WEBHOOK_ENABLED
(get_secret() falls back to the default profile's .env), so the env var
force-enabled the listener and tripped the MultiplexConfigError check.

Apply the same guard used by the api_server env block (and the generic
_enable_from_env path): pop the _enabled_explicit marker and only set
enabled = True when the platform was not explicitly disabled.

Closes #85637
2026-09-02 06:48:31 -07:00
liuhao1024 7d509657a8 fix(tts): resolve default output dir from the active profile
DEFAULT_OUTPUT_DIR was resolved once at import time, so long-lived
multi-profile runtimes (dashboard console, TUI/Desktop backend, cron,
kanban workers) kept writing synthesized audio into the launch
profile's cache/audio instead of the requesting profile's (#98749).

Same bug class and fix as skills_tool (f8723c478) and skills_sync
(#65828): keep the legacy module attribute for tests and external
patchers, but re-resolve from the live profile-scoped HERMES_HOME on
every synthesis call.
2026-09-02 06:48:31 -07:00
liuhao1024 39fca697ae fix(tools): log boot-time UnscopedSecretError probes at debug, not warning
With multiplexing on, check_fns evaluated before any profile secret scope
exists fail closed by design: get_secret raises UnscopedSecretError and the
tool re-probes on the first scoped turn. _run_check_fn_uncached logged that
expected signal like a crashed check_fn (WARNING + exc_info), so every
multiplexed gateway start printed three full tracebacks that drowned real
check_fn failures.

Split the handler: an unscoped read reported while the profile cache scope
was unresolved logs one debug line without a traceback; the same error with
the scope resolved is a genuinely lost scope and keeps the loud
warning + traceback.

Fixes #100697
2026-09-02 06:48:31 -07:00
nftpoetrist 79732c6452 fix(a2a): scope multiplex secondary-profile construction, not shared env
A2AAdapter.__init__ / _default_agent_name / _load_served_agents read
A2A_PORT, A2A_AGENT_NAME and A2A_AGENT_DESCRIPTION from raw os.environ,
so a secondary multiplex profile borrowed the default profile's port and
Agent Card identity. Skip the env read when constructed inside a
secondary profile's scope (_profile_scoped(), #98738 pattern) and fall
to config.extra / module defaults instead. The default profile keeps its
unscoped env precedence.

Salvaged from #100382 (tests trimmed to two).
2026-09-02 06:48:31 -07:00
nftpoetrist 504dd1f24f fix(raft): scope RAFT_PROFILE resolution to the active multiplex profile
_spawn_bridge, _env_enablement and register()'s platform_hint read
RAFT_PROFILE via raw os.environ. Under a multiplexed secondary profile
os.environ holds the DEFAULT profile's bridged value, so the bridge
subprocess / CLI hint pointed at another profile's Raft identity.
Resolve through get_secret() only when running inside a secondary
profile's scope (Buzz/SimpleX _profile_scoped() pattern, #98738); the
default profile keeps its unscoped os.environ read. No fallthrough to
os.environ after a scoped miss.

Salvaged from #100392 (tests trimmed to two).
2026-09-02 06:48:31 -07:00
Teknium d9b051a9d6 fix(desktop): pin the project '+' new chat to the profile its tree is shown under
workspace-session-target.ts never set $newChatProfile, so a new session
started from a project's "+" reached desktopSessionCreateParams with no
intent and fell back to $activeGatewayProfile — which an in-flight profile
swap can move between the click and Send, landing session.create on the
wrong backend (#79005 flaw 3, second path). Pin the tree's profile
(projectProfile()) via the same intent write newSessionInProfile uses.

Fixes #79005
2026-09-02 06:47:55 -07:00
Ayush Nangia 9189e42842 fix(desktop): reconnect an active profile whose gateway is closed 2026-09-02 06:47:55 -07:00
Teknium dfa74b5815 fix(desktop): boot session pop-out/watch windows against the session's owning profile
`openSessionInNewWindow` → IPC `hermes:window:openSession` →
`buildSessionWindowUrl` emitted no `profile`, so a secondary window (⇧⌘-click
pop-out, subagent watch) was a full renderer that adopted the PRIMARY
backend's profile and resolved the session id against the wrong store —
blank/wrong session for any non-primary profile (#82768, #61286).

The owning profile now rides the URL as `&profile=`, exactly the carry the
HUD already does (buildHudWindowUrl / windowProfileOverride in
use-gateway-boot); the renderer picks it with the same ladder openHud uses:
the session's stamped owner wins, an unstamped/uncached id (a brand-new
subagent child) inherits the profile the user is looking at.

Diagnosis credit: @DomGrieco (#82794).

Co-authored-by: DomGrieco <6556434+DomGrieco@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
liuhao1024 74b00d7a97 fix(dashboard): route every per-row session request at the row's owning profile (salvage #99387)
The Sessions page listed rows stamped with their owning profile but sent
delete/bulk-delete to the global management profile, which stays "" while
the sticky active profile equals the dashboard process's own — so the
request opened the process store, missed, and returned a false
`already_absent` success while the row survived in profiles/<p>/state.db.

Same class at three sibling sites the PR didn't touch: renameSession,
exportSessionUrl and the expanded-row getSessionMessages read. One
`rowProfile(id)` owner now feeds all four (+ bulk delete); unstamped rows
(search results) fall back to the management profile as before.

Test trimmed to one jsdom scenario driving all four row actions.

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-02 06:47:55 -07:00
Teknium 2ce6f5c538 fix(desktop): refresh routed-profile Bot Chat transcripts (salvage #99333)
Two sites dropped profile ownership on the tile-transcript refresh path:

- tui_gateway/server.py _sessions_sig statted only the launch home's
  state.db, so a turn landing in a served sibling profile's store never
  produced sessions.changed. The watcher now also probes every profile
  home _profile_home() has resolved for this backend (empty set on
  single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
  transcript unscoped; it now passes the tile's ownerRoute scope, the
  same way reconcileActiveTranscript already does for the main pane, and
  keys the change signature by owner.

Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).

Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
SZWzz ff8c1aca7a test(desktop): cover combined remote route precedence 2026-09-02 06:47:55 -07:00
SZWzz ae81786580 fix(desktop): let a per-profile remote override win over the forced-local route (#90477)
resolveRegistryLocalRoute collapsed globalRemote and profileRemoteOverride
into one forced-local branch. The two cases are different:

- globalRemote: forcing "This device" to spawn genuinely-local children is
  the intended migration behavior — unchanged.
- profileRemoteOverride: the per-profile SSH/remote override is an explicit,
  authoritative routing decision for that profile. Forcing local made the
  roster enumerate the profile via its override but open the thread in a
  forced-local child, which dies with 'Profile "x" no longer exists' when
  the profile only exists on the remote — reproduced on a macOS Desktop in
  global SSH mode where mythony-agent/q-agent exist only on the NAS.

The registry 'local' entry now delegates to the legacy profile route when a
per-profile override is present, so the override stays authoritative.

Tests: the override case now pins delegation, and a new witness pins that
globalRemote alone still forces local; both contracts are asserted together.
88 connection-registry + 91 remote-lifecycle + 73 routing tests pass;
tsc --build clean.
2026-09-02 06:47:55 -07:00
Teknium 3fcfa647ed docs(google-chat): document per-profile scoping and ADC fail-closed under multiplex 2026-09-02 06:47:45 -07:00
pierrenode 3245264668 fix(relay): carry routed profile through the passthrough-plane forward
The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.

Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.

Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
2026-09-02 06:47:45 -07:00
Kong dca9a54427 fix(gateway): match WhatsApp profile_routes across JID/LID/number forms
`ProfileRoute.matches()` compared `chat_id` as an exact string, so a
WhatsApp route written as a phone number never matched the JID
(`…@s.whatsapp.net`) or LID (`…@lid`) the bridge actually delivers, and
the inbound fell through to the default profile. Allowlists and session
keys already canonicalize these via `gateway.whatsapp_identity`.

Exact compare still wins first; only whatsapp / whatsapp_cloud *user*
chats get the alias intersection fallback (applied to both `chat_id` and
`parent_chat_id`). Groups, broadcasts and every other platform stay exact.

Salvaged from #85081 (Kong); parent_chat_id fallback added on top.
2026-09-02 06:47:45 -07:00
Hudson db639e1023 fix(gateway): bind every adapter to the runner at the _create_adapter boundary
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.

Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.

Salvaged from #70831 (Hudson). First reported in #68332.
2026-09-02 06:47:45 -07:00
Jony b1bc9bb650 fix(google-chat): scope multiplex profile config and fail ADC closed
Route every GOOGLE_CHAT_* / GOOGLE_APPLICATION_CREDENTIALS read through a
module-local `_get_scoped_secret` (scope-authoritative under multiplex,
os.environ fallback only for the unscoped default-profile constructor, so
startup/reconnect never hits UnscopedSecretError — #70652 class). Snapshot
Pub/Sub callback knobs on the instance while the scope is still installed,
and seed them into `extra` from `_env_enablement`.

When a scoped profile has no service-account setting, do NOT fall through
to google.auth.default(): ADC reads the process env directly and would
authenticate the profile as another profile's SA. Fail closed with an
explicit error (adapter and standalone send).

Also resolve the bot-id cache path at call time via get_hermes_home() so
profiles don't share one identity cache.

Fixes #73439.
Salvaged from #73445 (Jony) with the ADC guard from #57674 (Ray, first submitter).

Co-authored-by: Ray <rayjun0412@gmail.com>
2026-09-02 06:47:45 -07:00
SHT 14f20d142e fix(feishu): isolate lark_oapi WS globals per profile and supervise the client thread (#73779)
lark_oapi.ws.client keeps the asyncio loop used by Client.start() in a
module-level global, and Hermes monkey-patched websockets.connect on the
shared module. Under multiplex every profile runs its own WS client on a
dedicated thread, so N threads overwrote each other's globals
(last-write-wins): clients scheduled tasks on a sibling's loop ('Future
attached to a different loop') or bound to the wrong loop and went deaf.

Install process-wide shims once: the module loop becomes a proxy that
forwards to the calling thread's registered loop, and websockets.connect
a dispatcher merging the calling thread's ping overrides. Add a per-adapter
supervisor: the executor future was only awaited by disconnect(), so a
dead WS thread left the profile silently deaf; now it is rebuilt with
capped exponential backoff while the adapter is meant to be connected.

Salvaged from PR #84165, trimmed: the legacy fallback path when the shim
cannot install was dropped (the shim only touches two attributes the
adapter already depended on); tests reduced to three. The loop-isolation
approach was first proposed by @zmlgit in #64247/#69904.

Co-authored-by: zmlgit <6995990+zmlgit@users.noreply.github.com>
2026-09-02 06:47:30 -07:00
SHT 9f40d3bdd0 fix(feishu): resolve DM admission config per-profile under multiplex (#86905)
Route FEISHU_ALLOW_BOTS, FEISHU_GROUP_POLICY, FEISHU_ALLOWED_USERS,
FEISHU_BOT_*, FEISHU_APP_ID and FEISHU_REQUIRE_MENTION through
_get_scoped_secret, and snapshot FEISHU_ALLOW_ALL_USERS /
GATEWAY_ALLOW_ALL_USERS into settings.allow_all_dm at construct time so
_admit (running on the lark_oapi WS thread, no secret scope) no longer
reads the default profile's os.environ. Feishu open_ids are app-scoped,
so a secondary bot's DMs were always dm_policy_rejected against the
default profile's allow-list.

Salvaged from PR #86908 trimmed to the adapter call-site migration; the
fail-closed _get_scoped_secret / _auth_env rewrites are already on main
(2912c36aa4, agent/secret_scope.py get_secret policy).
2026-09-02 06:47:30 -07:00
Teknium 5b4cff976d fix(gateway): fingerprint Teams client_id and WeCom bot_id for the multiplex credential guard
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
2026-09-02 06:47:30 -07:00
webtecnica 97b8a97861 fix(feishu): include app credentials in multiplex fingerprint (#76793) 2026-09-02 06:47:30 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium d45bc39667 fix(gateway): resolve the ephemeral personality prompt per turn from the scoped profile
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.

Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.

Fixes #89161

Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
StanleyStetson 94a2d7f8af fix(gateway): persist slash-command config writes into the routed profile
Slash dispatch already runs inside _profile_runtime_scope under multiplex,
but _save_gateway_config_key (/reasoning --global, /fast, show/hide),
/memory approval, /skills approval, /verbose and /footer built their write
path from the module constant gateway.run._hermes_home — the launch home —
so a routed profile's toggles landed in the default profile's config.yaml
while the reads (via _gateway_config_home()) saw the routed one.

Resolve the write path through _gateway_config_home() at all five sites so
reads and writes agree. Single-profile gateways never install the override
and keep resolving the launch home.

Fixes #87939
Fixes #75684

Co-authored-by: Bao <nnqbao@gmail.com>
2026-09-02 06:36:16 -07:00
Teknium 257da5ca07 fix(gateway): route /review and /reload-skills executor hops through the scoped helper
Sibling sites of the bare loop.run_in_executor(None, …) class fixed for
/insights, /debug and /goal draft: the worker started with an empty
context, so get_hermes_home()-relative reads (skills.external_dirs,
disabled skills, the reviewer subagent's home and secret scope) resolved
the launch home instead of the routed profile under multiplex.
2026-09-02 06:36:16 -07:00
Drexuxux fbd9730e30 fix(gateway): run /insights, /debug and /goal draft inside the routed profile
The multiplexed inbound handler wraps every message in _profile_runtime_scope,
which installs the routed profile's HERMES_HOME override and its secret scope
as contextvars. A bare loop.run_in_executor(None, fn) starts the worker with an
EMPTY context, so neither reaches the blocking work.

GatewaySlashCommandsMixin already knows this -- /compress goes through
_run_in_executor_with_context and the call site says why. Three siblings in the
same file still used the bare hop:

  /insights   SessionDB() with no explicit path resolves get_hermes_home() at
              call time (_default_db_path), so the worker opened the DEFAULT
              profile's state.db. Under multiplexing the command reported
              another profile's conversations, session counts and sources to
              this profile's user.

  /debug      collects that home's logs/config and uploads them to a public
              paste, so it published the default profile's diagnostics from
              another profile's chat.

  /goal draft calls the auxiliary LLM, whose provider/credential resolution
              reads the profile secret scope -- unscoped it falls back to
              process-global os.environ, which under multiplexing may hold a
              different profile's keys.

Route all three through _run_in_executor_with_context.

/reload-skills is deliberately left alone: tools.skills_tool binds SKILLS_DIR
at import time, so it does not follow the contextvar either way. Fixing that
needs the module-global retarget web_server._profile_scope performs under a
lock, which is a different change from context propagation.

Single-profile gateways never enter the scope, so their behaviour is unchanged.
2026-09-02 06:36:16 -07:00
Teknium eed481bd34 docs(profiles): cron delivery for routed profiles rides the shared bot only for exact routed targets (#101113) 2026-09-02 06:27:24 -07:00
Teknium a6351a71e5 fix(cron): route a credentialless satellite's cron delivery through the primary adapter for exact profile_routes targets (#101113)
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.

- cron/scheduler.py: factor the preflight's primary-config route loader into
  `_primary_profile_routes_for_current_home()` (one owner for both halves,
  so route semantics cannot drift) and add `SharedRouteAdapters`, a
  read-only view over the primary adapter map that resolves an adapter for
  a (platform, target) ONLY when an enabled primary route with a
  chat_id/thread_id maps that exact target to the current profile —
  using the same `ProfileRoute.matches` predicate as inbound routing.
  `_deliver_result` resolves the transport per target from it; everything
  else (unmatched chat, disabled route, route for another profile, no
  primary adapter, guild-only route) is a miss and never uses the primary
  bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
  gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
  fallback: with no matching route the view is falsy and delivers nothing.

Fixes #101113
2026-09-02 06:27:24 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Cyber-Yichen a2fea79de6 fix(cron): isolate multiplex profile failures per profile (#74878)
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:

- startup recovery loop: a per-profile exception (e.g. an unreadable
  executions.db raising sqlite3.DatabaseError) was uncaught and killed the
  ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
  exception escaped to the cycle-wide handler, skipping every remaining
  profile that cycle and marking all of them failed.

Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).

Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).

Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
이민재 e48bb828d4 test(cron): preserve explicit suggestions path override 2026-09-02 06:27:24 -07:00
wanliqin 9da8842585 fix(cron): resolve profile store paths per call
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.

Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 0fd9218e5a docs: ${VAR} config refs resolve per-profile under a multiplexed gateway 2026-09-02 06:19:03 -07:00
Teknium 638c1f3204 chore(contributors): map patryk.kopycinski@elastic.co -> patrykkopycinski 2026-09-02 06:19:03 -07:00
Teknium b6f0106602 docs(design): note loader-boundary dotenv guard, scoped ${VAR} expansion and scoped .env publish in multiplexing design doc
Accuracy pass on the #89950 doc for the fixes landing alongside it (#77562, #84079, #88441).
2026-09-02 06:19:03 -07:00
Eva 011a60b7cc docs(design): add the multiplexing-gateway design doc referenced by secret_scope
agent/secret_scope.py has pointed at docs/design/multiplexing-gateway.md
("Workstream A") since the fail-closed secret scope landed, but the file was
never added. This writes the missing doc from the code as it stands today:
the mode flag, scope composition (_profile_runtime_scope seams), the
context-local secret scope and HERMES_HOME override, routing/serving/
persistence/session-lane isolation, the control-plane RPCs, failure modes,
and an honest table of what is still process-global (per-profile MCP
registries are tracked in #67605).

Doc-only change; no code touched. Style follows docs/profile-routing.md.
2026-09-02 06:19:03 -07:00