Commit Graph

14240 Commits

Author SHA1 Message Date
nftpoetrist 504dd1f24f fix(raft): scope RAFT_PROFILE resolution to the active multiplex profile
_spawn_bridge, _env_enablement and register()'s platform_hint read
RAFT_PROFILE via raw os.environ. Under a multiplexed secondary profile
os.environ holds the DEFAULT profile's bridged value, so the bridge
subprocess / CLI hint pointed at another profile's Raft identity.
Resolve through get_secret() only when running inside a secondary
profile's scope (Buzz/SimpleX _profile_scoped() pattern, #98738); the
default profile keeps its unscoped os.environ read. No fallthrough to
os.environ after a scoped miss.

Salvaged from #100392 (tests trimmed to two).
2026-09-02 06:48:31 -07:00
Teknium 2ce6f5c538 fix(desktop): refresh routed-profile Bot Chat transcripts (salvage #99333)
Two sites dropped profile ownership on the tile-transcript refresh path:

- tui_gateway/server.py _sessions_sig statted only the launch home's
  state.db, so a turn landing in a served sibling profile's store never
  produced sessions.changed. The watcher now also probes every profile
  home _profile_home() has resolved for this backend (empty set on
  single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
  transcript unscoped; it now passes the tile's ownerRoute scope, the
  same way reconcileActiveTranscript already does for the main pane, and
  keys the change signature by owner.

Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).

Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
pierrenode 3245264668 fix(relay): carry routed profile through the passthrough-plane forward
The relay text lane stamps `SessionSource.profile` from the wire frame
(#60586), but `PassthroughForward` had no profile field, so a relayed
Discord slash-command/button/modal always landed in the default profile's
agent:main namespace even when the connector resolved a specific profile.

Add an optional `profile` to `PassthroughForward` (read off the wire in
`_passthrough_from_wire`) and stamp it on the interaction's SessionSource.
Absent on the wire → None → legacy routing, byte-identical for
single-profile gateways. Contract doc updated.

Salvaged from #61012 (pierrenode); two context conflicts resolved
(delivered_via_upstream_relay / _platform_by_chat landed on main).
2026-09-02 06:47:45 -07:00
Kong dca9a54427 fix(gateway): match WhatsApp profile_routes across JID/LID/number forms
`ProfileRoute.matches()` compared `chat_id` as an exact string, so a
WhatsApp route written as a phone number never matched the JID
(`…@s.whatsapp.net`) or LID (`…@lid`) the bridge actually delivers, and
the inbound fell through to the default profile. Allowlists and session
keys already canonicalize these via `gateway.whatsapp_identity`.

Exact compare still wins first; only whatsapp / whatsapp_cloud *user*
chats get the alias intersection fallback (applied to both `chat_id` and
`parent_chat_id`). Groups, broadcasts and every other platform stay exact.

Salvaged from #85081 (Kong); parent_chat_id fallback added on top.
2026-09-02 06:47:45 -07:00
Hudson db639e1023 fix(gateway): bind every adapter to the runner at the _create_adapter boundary
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.

Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.

Salvaged from #70831 (Hudson). First reported in #68332.
2026-09-02 06:47:45 -07:00
Jony b1bc9bb650 fix(google-chat): scope multiplex profile config and fail ADC closed
Route every GOOGLE_CHAT_* / GOOGLE_APPLICATION_CREDENTIALS read through a
module-local `_get_scoped_secret` (scope-authoritative under multiplex,
os.environ fallback only for the unscoped default-profile constructor, so
startup/reconnect never hits UnscopedSecretError — #70652 class). Snapshot
Pub/Sub callback knobs on the instance while the scope is still installed,
and seed them into `extra` from `_env_enablement`.

When a scoped profile has no service-account setting, do NOT fall through
to google.auth.default(): ADC reads the process env directly and would
authenticate the profile as another profile's SA. Fail closed with an
explicit error (adapter and standalone send).

Also resolve the bot-id cache path at call time via get_hermes_home() so
profiles don't share one identity cache.

Fixes #73439.
Salvaged from #73445 (Jony) with the ADC guard from #57674 (Ray, first submitter).

Co-authored-by: Ray <rayjun0412@gmail.com>
2026-09-02 06:47:45 -07:00
SHT 14f20d142e fix(feishu): isolate lark_oapi WS globals per profile and supervise the client thread (#73779)
lark_oapi.ws.client keeps the asyncio loop used by Client.start() in a
module-level global, and Hermes monkey-patched websockets.connect on the
shared module. Under multiplex every profile runs its own WS client on a
dedicated thread, so N threads overwrote each other's globals
(last-write-wins): clients scheduled tasks on a sibling's loop ('Future
attached to a different loop') or bound to the wrong loop and went deaf.

Install process-wide shims once: the module loop becomes a proxy that
forwards to the calling thread's registered loop, and websockets.connect
a dispatcher merging the calling thread's ping overrides. Add a per-adapter
supervisor: the executor future was only awaited by disconnect(), so a
dead WS thread left the profile silently deaf; now it is rebuilt with
capped exponential backoff while the adapter is meant to be connected.

Salvaged from PR #84165, trimmed: the legacy fallback path when the shim
cannot install was dropped (the shim only touches two attributes the
adapter already depended on); tests reduced to three. The loop-isolation
approach was first proposed by @zmlgit in #64247/#69904.

Co-authored-by: zmlgit <6995990+zmlgit@users.noreply.github.com>
2026-09-02 06:47:30 -07:00
SHT 9f40d3bdd0 fix(feishu): resolve DM admission config per-profile under multiplex (#86905)
Route FEISHU_ALLOW_BOTS, FEISHU_GROUP_POLICY, FEISHU_ALLOWED_USERS,
FEISHU_BOT_*, FEISHU_APP_ID and FEISHU_REQUIRE_MENTION through
_get_scoped_secret, and snapshot FEISHU_ALLOW_ALL_USERS /
GATEWAY_ALLOW_ALL_USERS into settings.allow_all_dm at construct time so
_admit (running on the lark_oapi WS thread, no secret scope) no longer
reads the default profile's os.environ. Feishu open_ids are app-scoped,
so a secondary bot's DMs were always dm_policy_rejected against the
default profile's allow-list.

Salvaged from PR #86908 trimmed to the adapter call-site migration; the
fail-closed _get_scoped_secret / _auth_env rewrites are already on main
(2912c36aa4, agent/secret_scope.py get_secret policy).
2026-09-02 06:47:30 -07:00
Teknium 5b4cff976d fix(gateway): fingerprint Teams client_id and WeCom bot_id for the multiplex credential guard
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
2026-09-02 06:47:30 -07:00
webtecnica 97b8a97861 fix(feishu): include app credentials in multiplex fingerprint (#76793) 2026-09-02 06:47:30 -07:00
Teknium 527da60844 fix(cli): report a named profile as running when the default multiplexer serves it
`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.

Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.

Salvage of #69118 rebased onto the shared helper (which post-dates it).

Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
Teknium d45bc39667 fix(gateway): resolve the ephemeral personality prompt per turn from the scoped profile
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.

Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.

Fixes #89161

Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
2026-09-02 06:36:16 -07:00
StanleyStetson 94a2d7f8af fix(gateway): persist slash-command config writes into the routed profile
Slash dispatch already runs inside _profile_runtime_scope under multiplex,
but _save_gateway_config_key (/reasoning --global, /fast, show/hide),
/memory approval, /skills approval, /verbose and /footer built their write
path from the module constant gateway.run._hermes_home — the launch home —
so a routed profile's toggles landed in the default profile's config.yaml
while the reads (via _gateway_config_home()) saw the routed one.

Resolve the write path through _gateway_config_home() at all five sites so
reads and writes agree. Single-profile gateways never install the override
and keep resolving the launch home.

Fixes #87939
Fixes #75684

Co-authored-by: Bao <nnqbao@gmail.com>
2026-09-02 06:36:16 -07:00
Drexuxux fbd9730e30 fix(gateway): run /insights, /debug and /goal draft inside the routed profile
The multiplexed inbound handler wraps every message in _profile_runtime_scope,
which installs the routed profile's HERMES_HOME override and its secret scope
as contextvars. A bare loop.run_in_executor(None, fn) starts the worker with an
EMPTY context, so neither reaches the blocking work.

GatewaySlashCommandsMixin already knows this -- /compress goes through
_run_in_executor_with_context and the call site says why. Three siblings in the
same file still used the bare hop:

  /insights   SessionDB() with no explicit path resolves get_hermes_home() at
              call time (_default_db_path), so the worker opened the DEFAULT
              profile's state.db. Under multiplexing the command reported
              another profile's conversations, session counts and sources to
              this profile's user.

  /debug      collects that home's logs/config and uploads them to a public
              paste, so it published the default profile's diagnostics from
              another profile's chat.

  /goal draft calls the auxiliary LLM, whose provider/credential resolution
              reads the profile secret scope -- unscoped it falls back to
              process-global os.environ, which under multiplexing may hold a
              different profile's keys.

Route all three through _run_in_executor_with_context.

/reload-skills is deliberately left alone: tools.skills_tool binds SKILLS_DIR
at import time, so it does not follow the contextvar either way. Fixing that
needs the module-global retarget web_server._profile_scope performs under a
lock, which is a different change from context propagation.

Single-profile gateways never enter the scope, so their behaviour is unchanged.
2026-09-02 06:36:16 -07:00
Teknium a6351a71e5 fix(cron): route a credentialless satellite's cron delivery through the primary adapter for exact profile_routes targets (#101113)
Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.

- cron/scheduler.py: factor the preflight's primary-config route loader into
  `_primary_profile_routes_for_current_home()` (one owner for both halves,
  so route semantics cannot drift) and add `SharedRouteAdapters`, a
  read-only view over the primary adapter map that resolves an adapter for
  a (platform, target) ONLY when an enabled primary route with a
  chat_id/thread_id maps that exact target to the current profile —
  using the same `ProfileRoute.matches` predicate as inbound routing.
  `_deliver_result` resolves the transport per target from it; everything
  else (unmatched chat, disabled route, route for another profile, no
  primary adapter, guild-only route) is a miss and never uses the primary
  bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
  gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
  fallback: with no matching route the view is falsy and delivers nothing.

Fixes #101113
2026-09-02 06:27:24 -07:00
Teknium 62a4599f89 fix(cron): keep profile scope on the standalone fallback pool; desktop ticker stands down for profiles with their own gateway (#100489)
Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
2026-09-02 06:27:24 -07:00
berg 357062f39e fix(dashboard): multiplex cron fire URL uses the default profile's listener port
Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
2026-09-02 06:27:24 -07:00
Cyber-Yichen a2fea79de6 fix(cron): isolate multiplex profile failures per profile (#74878)
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:

- startup recovery loop: a per-profile exception (e.g. an unreadable
  executions.db raising sqlite3.DatabaseError) was uncaught and killed the
  ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
  exception escaped to the cycle-wide handler, skipping every remaining
  profile that cycle and marking all of them failed.

Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).

Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).

Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
이민재 e48bb828d4 test(cron): preserve explicit suggestions path override 2026-09-02 06:27:24 -07:00
wanliqin 9da8842585 fix(cron): resolve profile store paths per call
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.

Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
2026-09-02 06:27:24 -07:00
Teknium 4155ea97e8 perf(serve): Desktop backend announces its socket before MCP discovery imports the SDK
`cmd_dashboard` started the background MCP discovery thread before importing
`hermes_cli.web_server`. The thread's first act is the ~350ms `mcp` SDK
import, which holds the GIL against the main thread's own web_server import,
so the HERMES_BACKEND_READY sentinel — and every renderer paint behind it —
moved ~300ms later on every Desktop cold start with any MCP server configured.

Desktop `serve` (headless + HERMES_DESKTOP=1) now arms discovery one second
after the sentinel instead. Starting it AT the bind was measured to give back
most of the gain (the renderer's WebSocket connect + first hydration reads
contend on the same loop). An agent build inside that window pulls the
deferred start forward itself via `wait_for_mcp_discovery`, so the bounded
join and the late-binding tool refresh behave exactly as before. Dashboard
and non-Desktop `serve` keep the eager pre-import ordering.

Minimal reimplementation of the MCP-deferral slice of #96751 by @helix4u;
the plugin-route deferral / 503 middleware / cron-after-bind slices were
measured at ~0-10ms each and are not taken.

Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
2026-09-02 06:19:08 -07:00
Teknium 2475335443 fix(config): keep .env publishes inside the routed profile scope under multiplex (#88441)
`save_env_value` / `remove_env_value` already write the right FILE
(`get_env_path()` honors the profile-home override, so a routed turn lands
in `profiles/<p>/.env`, not the root -- #77490's premise), but the
in-process mirror went to `os.environ` unconditionally. Under a
multiplexed gateway a `/pair` grant mirrored into `DISCORD_ALLOWED_USERS`
from profile B therefore published B's allowlist into the SHARED process
env, and B's own installed scope never saw the new value.

Add `_publish_env_value`: when multiplex is active and a secret scope is
installed, update the installed scope mapping (so same-turn scope reads see
the grant) and leave `os.environ` untouched; every other caller keeps the
legacy `os.environ` publish. Replace the stale TODO in gateway/pairing.py.
2026-09-02 06:19:03 -07:00
Teknium b53c50cf8a fix(config): resolve ${VAR} config refs through the profile secret scope (#84079)
`_env_expand_match` read `os.environ` directly, so under a multiplexed
gateway every secondary profile whose config.yaml carried
`${MATRIX_ACCESS_TOKEN}` (or `${env:...}`) expanded to the DEFAULT
profile's token loaded at startup -- each profile "had" the credential and
one inbound message fanned out across all of them. This is the residual
half of #84079 the secondary credential gate cannot see (the expanded
token is non-empty).

Add `_env_ref_lookup`: outside a secret scope it is the same
`os.environ.get`; inside a scope it goes through `get_secret`, which is
authoritative under multiplexing and an environ overlay otherwise -- the
same policy `gateway.config._getenv` and `get_env_value` already follow.
The cache env-snapshot (#58514) uses the same lookup so a scoped load is
not served another scope's cached expansion.
2026-09-02 06:19:03 -07:00
webtecnica d7dc75ccff fix(gateway): multiplex must not apply default-profile creds to unconfigured profiles (#84079)
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.

Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.

Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Patryk Kopycinski 1eb71756af toolchain-self-improve: route API_SERVER_KEY to profile env 2026-09-02 06:19:03 -07:00
Teknium 0a6aa7cce1 fix(env_loader): log routed-scope dotenv skip once per home; port single-profile control test
Follow-up to the #77592 salvage: emit a once-per-home debug line where the
multiplex guard skips the process-global dotenv load (requested on #77562),
and port the single-profile control test from #77970 so the guard is pinned
to the multiplex flag rather than the home override alone.

Co-authored-by: DonShelly <25538402+DonShelly@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Lester Liang 1aa62ceb45 fix(security): isolate multiplex dotenv reloads 2026-09-02 06:19:03 -07:00
Teknium 0058bde251 fix(bot-relay): a DM into a live Bot Chat queues as the next turn, never interrupts the one in flight 2026-09-02 06:18:55 -07:00
Teknium d29a7936e4 fix(bot-mode): DMs to a Desktop-owned Bot Chat land in the live session instead of being dropped (#100523)
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.

bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.

On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.

Closes #100523
Supersedes #100544, #100542

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 06:18:55 -07:00
Teknium 16370ae539 fix(tui_gateway): setup.status / setup.runtime_check answer for the requested profile
Port the Python half of PR #94147: both readiness RPCs accept an optional
`profile` and bind that profile's HERMES_HOME + .env secret scope for the
duration of the check via `_session_profile_runtime_scope` (ContextVars,
so concurrent checks stay isolated). Unknown profile → ok=False with an
explicit error instead of quietly reporting the launch profile's readiness.

`_has_any_provider_configured(strict_profile_scope=True)` reads provider
env only from the bound secret scope (never os.environ) and skips the
host-wide fallbacks (gh auth, Claude Code credentials, api-key
active_provider in auth.json) that describe the launch host, not the
target profile. Unscoped callers are byte-identical to before.

The desktop TS half of #94147 targets plugin.js, which was deleted on main;
it needs a recut on create-dialog.tsx.

Supersedes #94147 (python half)

Co-authored-by: Zeus-Deus <100132710+Zeus-Deus@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
aeonsong 8076c78c87 fix(desktop): scope handoff config to session profile 2026-09-02 06:17:47 -07:00
Teknium 7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
fangliquanflq 21cebbfd68 fix(doctor): honor disabled built-in memory stores
Salvaged from #100677. Fixes #100668: hermes doctor reported MEMORY.md/USER.md
char counts even when memory.memory_enabled / memory.user_profile_enabled were
false. Resolve the flags via get_builtin_memory_store_flags (same resolver the
agent uses), only inspect enabled targets, and point at the Memory Provider
section when both are disabled.
2026-09-02 06:17:25 -07:00
Teknium 70dc1606c6 test(cli): pin OSC 9 / Warp OSC 777 bell emitters; docs + contributor mappings
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
  only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
  of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
2026-09-02 06:17:10 -07:00
Teknium c2954c8934 feat(model-catalog): picker catalogs refresh every 20 minutes, gateway keeps them warm
The /model picker's remote catalogs (curated manifest, OpenRouter live
filter, Nous Portal recommendations) only refreshed when someone opened
the picker on a stale cache, with a 1h TTL. A delisted model (tencent/hy3:free
after the free promo ended) or a newly published one could sit stale for
an hour after the manifest deploy, and indefinitely in a gateway nobody
opened /model in.

- model_catalog.ttl_minutes: 20 replaces ttl_hours: 1 as the default;
  an explicitly set legacy ttl_hours is still honoured.
- model_catalog.refresh_catalogs() force-refreshes all three sources to
  disk; refresh_interval_seconds() exposes the cadence.
- Gateway spawns a supervised _model_catalog_refresh_watcher that calls
  it off-thread every TTL window, so every surface on the machine reads
  a cache no older than 20 minutes.
- Config migration v39→v40 drops the old ttl_hours: 1 default only.
- Docs: reference/model-catalog.md updated.
2026-09-02 06:16:54 -07:00
Teknium 11f932c935 fix(slack): interactive-caller and pre-fetch auth prefer the injected profile check; gate reads never fall through to os.environ
`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.

- Prefer the wired `set_authorization_check` callback (profile-bound
  `_make_adapter_auth_check`) at both sites; keep `__self__` introspection
  only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
  (scoped miss under multiplex → "", never os.environ); drop the raw
  `os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.

Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.

Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Teknium bbb087f3d1 fix(gateway): egress adapter and channel directory no longer follow the per-turn active profile
`_authorization_adapter` compared a stamped profile against
`_active_profile_name()`, which reads the per-turn HERMES_HOME override.
Inside a secondary profile's `_profile_runtime_scope` (cron, restored or
hand-built sources without transport provenance) that reported the
secondary itself, so it was handed the DEFAULT bot for egress instead of
the fail-closed None. Capture the launch identity once in `__init__`
(`_primary_profile_name`) and compare against that; the
`_active_profile_name()` fallback remains for partial fixtures.

`gateway/channel_directory.py` resolved `DIRECTORY_PATH` /
`CHANNEL_ALIASES_PATH` at import time, pinning every multiplexed profile's
directory to whichever home imported the module first. Resolve lazily
from the current home; the module attributes stay as explicit overrides
(tests patch them) and default to None.

Extracted from #87240 (topic-table half handled separately via #76487).

Co-authored-by: cherryb16 <166878179+cherryb16@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
Teknium 74775df53f fix(gateway): route-stamp primary callback auth and carry is_bot through the adapter auth check
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.

- `_make_adapter_auth_check`: for the shared primary adapter under
  multiplex, mirror the inbound message path exactly — stamp the
  `profile_routes` match so the routed profile's pairing store is
  consulted, and authorize under the TRANSPORT home via
  `_is_user_authorized_for_source` (same split as
  `_make_default_profile_message_handler`, 2afed50863). A rejected route
  fails closed like the ingress gate. Retain the receiving adapter as
  `_transport_adapter_ref` so config.yaml policy reads stay on it.
  Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
  `thread_id` as keywords only when set, so legacy 3-positional callbacks
  keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
  the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
  honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
  introspection class — fall back to the injected `gateway_runner` and
  the adapter's owner profile.

Fixes #86296
Fixes #92840

Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
2026-09-02 06:08:09 -07:00
elphamale 4346721117 fix(telegram): resolve button-caller authorization via the injected auth check, not handler introspection
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).

Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
2026-09-02 06:08:09 -07:00
Teknium 4afbecb429 test(cli): trim fast-serve coverage to the parity + dispatch invariants 2026-09-02 06:06:47 -07:00
Gille 5180601a6a perf(cli): dispatch serve without the full parser tree 2026-09-02 06:06:47 -07:00
Teknium 458e2ef1d2 refactor(state): collapse telegram topic v3 migration into one table-driven rebuild
Same behavior as the salvaged #76487 migration (fresh installs get the v3
shape; v1/v2 tables rebuild with profile_name leading the PK, legacy rows
into 'default' only, CASCADE FK supplied on the way), with the per-table
DDL written once instead of three times and the now-redundant v1->v2
CASCADE-only rebuild dropped (the v3 rebuild subsumes it).

Co-authored-by: Celio Monteiro <crdesign8@hotmail.com>
2026-09-02 05:59:24 -07:00
Celio Monteiro d55d9d128a fix(gateway): route profile into topic prune, cooldowns, and docs
Address hermes-sweeper review on #76487:

- Prefer hermes_profile from send metadata when pruning stale topic
  bindings so profile_routes cannot delete the transport adapter's
  namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
  (profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
2026-09-02 05:59:24 -07:00
Celio Monteiro 62be7043ff fix(gateway): pass routed source.profile into telegram topic state
Issue #76423 follow-up: wire SessionDB profile_name through gateway paths.

- Resolve profile from source.profile (never process-global active profile)
- Stamp adapter._hermes_profile_name for prune under multiplex
- /topic enable/status and binding record/recover/disable/restore paths
2026-09-02 05:59:24 -07:00
Celio Monteiro 351e4c0067 fix(state): namespace telegram topic tables by profile_name
Issue #76423: under multiplex_profiles a shared state.db keyed topic mode
and bindings only by Telegram chat_id/thread_id, so private-chat ids
collided across bots/profiles.

- Add profile_name to telegram_dm_topic_mode and telegram_dm_topic_bindings
- Schema v2→v3 rebuild; legacy rows migrate into the "default" namespace
- Keyword-only profile_name="default" on SessionDB topic APIs (compat)
2026-09-02 05:59:24 -07:00
Teknium 0437fe66f7 fix(discord): native slash commands honor guild/channel profile_routes
_build_slash_event and _dispatch_thread_session built their SessionSource
without guild_id/parent_chat_id, while on_message passes both. Route
matching in build_source keys off exactly those fields, so under
gateway.multiplex_profiles a guild- or channel-routed profile never matched
a native slash command: /new, /reset, /model, /profile, /status ... all ran
against the default profile and reset the wrong session (#69178, #91633).

Pass guild_id (interaction.guild_id, falling back to channel.guild like the
message path) and the thread's parent channel id into build_source at both
sites. One test pins channel + thread routing parity with messages.

Fixes #69178
Fixes #91633
Co-authored-by: Sora-bluesky <179361977+Sora-bluesky@users.noreply.github.com>
Co-authored-by: jondgilbert <42873618+jondgilbert@users.noreply.github.com>
Co-authored-by: tensorbit89-netizen <257030052+tensorbit89-netizen@users.noreply.github.com>
2026-09-02 05:55:36 -07:00
fangliquanflq f069ffd471 fix(gateway): offload handoff secret scope loading
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.

Fixes #100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-02 05:55:36 -07:00
Adolanium 55e1698979 fix(gateway): a secondary profile handoff fails closed when its config cannot load
_process_handoff caught a config load failure for a secondary profile,
logged a warning, and kept going with self.config, which is the primary
profile's config. The handoff then went out through the right bot to the
primary's home channel and the row was reported completed. That is the
exact wrong delivery the multi-profile handoff work exists to prevent,
and the same fail closed posture the no-live-adapters branch already
takes.

A load failure now logs an error and raises, which marks the row failed
so the CLI can report and retry it. The default profile path is
untouched, it never reloaded config.
2026-09-02 05:55:36 -07:00
kshitijk4poor e9fa7bc05e fix(state): keep the #94736 teardown self-heal alive under the deleted-WAL guard
Follow-ups on the salvaged #101081 guard:

- A clean close() lets SQLite unlink the WAL sidecars legitimately; the
  guard treated that as a lost generation and permanently halted the
  handle, so the #94736 late-write self-heal reopen dropped transcript
  tails (4 existing tests failed). close() now clears the recorded
  sidecar generation, and _wal_generation_was_lost() re-adopts the
  current sidecars after a clean /proc/self probe instead of relying on
  a stale snapshot.
- Healthy writes no longer walk /proc/self/fd: once a sidecar
  generation is recorded, the stat-based inode check alone detects an
  unlink/replace. The fd probe only runs in the empty-identity state
  (fresh DB, post-close reopen).
- DeletedWalGenerationError now subclasses StateDbReplacedError, so the
  gateway retry queue and run_agent flush divert transcripts to the
  JSONL fallback exactly as they do for a replaced store, instead of
  retrying forever against a halted handle.
- __init__ refuses once (under the startup lock) instead of twice per
  open, halving the system-wide /proc scan; dropped the dead
  include_self parameter and the dead _IS_WINDOWS clause.
- Test fixes: rstrip(' (deleted)') char-set bug -> removesuffix; the
  non-linux test now patches sys.platform (the real gate) instead of
  _IS_WINDOWS.
2026-09-02 18:24:54 +05:30
Cursor Agent 7f7df1ce44 fix(state): refuse SessionDB open and writes on a deleted WAL generation
A live writer can keep a deleted state.db-wal inode while a second opener
mints a fresh WAL at the same path. Fail closed on writable open (before
connect) and on the write-path sidecar identity check so the second
generation is never created.

Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
2026-09-02 18:24:54 +05:30