Commit Graph

2 Commits

Author SHA1 Message Date
fangliquanflq f069ffd471 fix(gateway): offload handoff secret scope loading
The handoff watcher entered _profile_runtime_scope synchronously on the
event loop each tick; hydrate_profile_secret_sources + build_profile_secret_scope
do blocking file/secret-source IO, so a slow profile secret read stalled every
adapter (#100014). Load the secret scope via asyncio.to_thread, then enter
the existing sync scope with prepared_secret_scope=.

Fixes #100014
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-02 05:55:36 -07:00
69k4xmdfm2-blip fc5fdb8c2a fix(gateway): handoff is broken on multi-profile installs (wrong DB, wrong key, wrong bot)
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.

1. The watcher polls only the ROOT store.
   `_handoff_watcher` resolves `self._session_db` with no profile scope, which
   always yields the root `state.db`. But `/handoff` run under
   `hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
   store. Nothing ever reads it, so the CLI times out after 60s while the
   gateway is alive and connected. The watcher now iterates
   `[(None, None), *secondary_profiles]` and polls each inside
   `_profile_runtime_scope`.

2. The destination session key is built without the profile namespace.
   `_process_handoff` called `build_session_key()` with no `profile=`,
   producing `agent:main:...` while that profile's own adapter routes organic
   inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
   reads.

3. Delivery uses the PRIMARY profile's adapter and config.
   `self.adapters` holds only the default profile's adapters (secondaries live
   in `self._profile_adapters[name]`) and `self.config` only the default's home
   channel. A secondary profile's handoff was therefore sent by the wrong bot,
   to the wrong chat, while persisting the right session key and reporting
   `handoff_state='completed'` — a false positive that looks correct in the
   database and is wrong on the wire.

Two robustness fixes in the same path:

4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
   turn plus delivery; awaiting it inline meant one slow handoff stopped the
   watcher from even polling the other profiles. With the CLI's 60s deadline, a
   valid handoff could time out purely because another profile's was ahead of
   it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
   never claimed twice, and a bounded drain on shutdown.

5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
   one in-process dispatch, so a row still in that state at startup belongs to
   a gateway that died mid-dispatch. It can never reach a terminal state, and
   `request_handoff` refuses new requests unless the state is
   NULL/completed/failed — that session could never hand off again, silently.
   `reclaim_stale_running_handoffs()` now fails those rows once per store at
   watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
   may already have switched the session key and dispatched, so a blind retry
   risks double delivery.

Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.

Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
2026-08-28 11:45:19 -07:00