Commit Graph

31 Commits

Author SHA1 Message Date
Teknium febd2af391 fix(gateway): skip credential-less WhatsApp on secondary multiplex profiles
_start_one_profile_adapters skipped only Platform.RELAY as shared
process-level ingress. WhatsApp is the same shape: the bridge is one
authenticated session tied to a single phone number, so a secondary
profile has no credential of its own to bring; constructing an adapter
for it only produced a connect/retry loop that stalled startup for every
profile queued behind it. Treat WhatsApp like Relay -- the active profile
owns the connection and route-stamped source.profile fans inbound turns
out to secondary profiles.

Salvage of #69042 (narrowed by its author to this one behavioral line);
test re-expressed on the current secondary-startup fixtures.

Co-authored-by: sshawn <28279366+lsshawn@users.noreply.github.com>
2026-09-02 07:01:23 -07:00
Joel Taylor 9d5c58be89 fix(gateway): guard Teams multiplex listener ownership 2026-09-02 07:01:23 -07:00
chelsealong 29640e3b5a fix(gateway): register secondary profiles' shell hooks and outbound webhooks
Multiplex gateway startup only ever calls agent.shell_hooks/
outbound_webhooks register_from_config() once, against the
root/default profile's config, before any profile scope exists.
_start_one_profile_adapters() discovers Python plugins per profile
but never registered that profile's own declarative `hooks:` block,
so a secondary profile's shell hooks (e.g. a deny-writes gate) and
outbound webhooks silently never fire.

Load and register each profile's own config inside its
_profile_runtime_scope, and key the module-level idempotence sets in
shell_hooks.py/outbound_webhooks.py by resolved Hermes home so two
profiles configuring an identical hook/webhook both register on
their own plugin manager instead of the second being dropped as a
duplicate of the first.

Fixes #92672
2026-09-02 07:00:13 -07:00
Teknium ca42d7a034 fix(gateway): isolate /voice state and voice-channel input per multiplexed profile (#84872)
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.

- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
  `<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
  (`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
  now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
  transcript handler (functools.partial); used at primary connect,
  primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
  `_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
  `_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
  primary path and both reconnect paths already did.

Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
Teknium e8231f01da fix(gateway): hydrate secondary-profile secrets off-loop at startup too (#99519 class)
_run_secondary_profile_reconnect now pre-hydrates external secret sources
in a worker thread (PR #99519); _start_one_profile_adapters entered
_profile_runtime_scope on the event loop three times per profile, each
running the same synchronous network-bound hydration under
_SECRET_SOURCE_CACHE_LOCK. Hydrate once via asyncio.to_thread and enter
every scope in that method with hydrate_secrets=False.

The reconnect test is parametrized over both entry points; the startup
reconnect handoff waits are deadline-based since the runner now hops to a
worker thread before publishing the replacement adapter.

Co-authored-by: GoBeromsu <37897508+GoBeromsu@users.noreply.github.com>
2026-09-02 06:59:11 -07:00
gobeumsu 04836886fa fix(gateway): run secondary-profile adapter auth setup off the event loop 2026-09-02 06:59:11 -07:00
Teknium 5b4cff976d fix(gateway): fingerprint Teams client_id and WeCom bot_id for the multiplex credential guard
Same class as the Feishu _app_id gap (#76793): Teams and WeCom authenticate
with an id/secret pair and store no token attribute, so
_adapter_credential_fingerprint returned None and cloned profiles started
competing adapters against one app. Add both ids to the attr tuple.
2026-09-02 06:47:30 -07:00
webtecnica 97b8a97861 fix(feishu): include app credentials in multiplex fingerprint (#76793) 2026-09-02 06:47:30 -07:00
webtecnica d7dc75ccff fix(gateway): multiplex must not apply default-profile creds to unconfigured profiles (#84079)
Secondary profile startup and reconnect now call the existing
`_platform_has_bot_credential` gate (the same one the primary loop and
primary reconnect use since #64674), so an enabled-in-YAML platform whose
credential is absent from that profile's secret scope is skipped instead
of built with an empty token and fanned out.

Independently reported and fixed in #72313 (@manny3), which added a
duplicate helper; the shared main helper is used here instead.

Co-authored-by: manny3 <16465310+manny3@users.noreply.github.com>
2026-09-02 06:19:03 -07:00
Teknium 6879a621b1 fix(gateway): live foreign token lock at startup exits 78 instead of retry-queueing forever
BasePlatformAdapter._acquire_platform_lock emits `{scope}_lock` with
retryable=True on purpose (#54167): a MID-RUN reconnect must be able to
recover once the live holder exits or a stale record is cleared. The
startup router keyed solely off that flag, so a live foreign holder of the
bot token at zero-connected startup landed in `_failed_platforms` with
gateway_state=running — alive, deaf, and retry-storming the token every
backoff — instead of the exit-78 (EX_CONFIG / startup_failed) contract
that #51228 established for single-writer conflicts.

Minimal class fix, salvaged from #83183 (@alexgunsberg) against current
main:

- gateway/restart.py: `is_global_startup_conflict(error_code)` — matches
  the `*_lock` / `lock_conflict` code families every adapter emits for
  scoped-lock and identity conflicts. Code only, never message text.
- gateway/run.py primary startup routing: a lock-conflict failure is
  routed as non-retryable (parked `fatal`, not queued). Nothing else
  connected → exit 78; alongside a transient peer → NS-609 mixed mode,
  gateway stays alive and only the peer retries.
- gateway/run.py `_schedule_secondary_profile_startup_reconnect`: the same
  contract for multiplex secondaries — park `<profile>:<platform>` fatal
  like `duplicate_credential` instead of scheduling a reconnect storm.
- Mid-run behavior is untouched: `_handle_adapter_fatal_error_impl` and
  the reconnect watcher still treat `*_lock` as retryable (#54167).

Not carried over from #83183 (superseded on main or out of scope): the
`degraded` lifecycle write only fires on the all-retryable path and the
runner immediately overwrites it with `running` (so busy/drain already
see `running`); the secondary retry bridge landed separately in
96489f3c1b (#92064); Buzz/IRC/LINE lock-tuple unpack and the reconnect
ownership registry are separate class fixes.

Live repro (real GatewayRunner.start(), isolated HERMES_HOME + lock dir,
live holder subprocess owning the lock via production
acquire_scoped_lock): before — exit_code=None, gateway_state=running,
telegram `retrying`, queued in _failed_platforms; after — exit_code=78,
gateway_state=startup_failed, telegram `fatal`, _failed_platforms={}.

Co-authored-by: alexgunsberg <alex@gunsberg.fi>
2026-09-02 00:17:54 -07:00
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
ruangraung dce4abe917 fix(gateway): log secondary startup-reconnect handoff failures instead of dropping them
Review follow-up to the AI code-review pass on PR #92074: the bridge task's
handoff into _schedule_secondary_profile_reconnect was unguarded at both call
sites inside the parked coroutine. The scheduler touches live registries
(_profile_failed_platforms slot creation, background-task registration), so an
unexpected raise there would kill the parked task as an unretrieved-task
exception — logged only at GC time via "Task exception was never retrieved",
where no operator ever looks. A fix whose entire purpose is to stop a platform
dying silently should not contain its own silent-death path; both handoff sites
now wrap the scheduler call with logger.exception so the failure lands in
gateway.log with profile and platform context.

The early-exit branch (gateway already _running when the bridge starts) had the
identical exposure and is guarded the same way — same bug class, fixed together.
Regression test drives a handoff raise end-to-end through the real bridge task:
the await completes cleanly, the error is captured in gateway.run's logger, and
no adapter or failed-platform slot leaks behind the failed handoff.
2026-08-25 22:55:07 -07:00
ruangraung 96489f3c1b fix(gateway): schedule secondary-profile reconnect when initial adapter connect fails
When gateway.multiplex_profiles is active, a secondary profile whose platform
adapter fails its initial connect at startup was silently given up on: the
failure branches in _start_one_profile_adapters() logged and disconnected, but
never scheduled recovery. One unlucky connect window during a Telegram API
outage left the profile permanently silent until manual restart (~80 min in
the observed incident), while the mid-run fatal path already recovers via
_handle_profile_adapter_fatal_error() -> _schedule_secondary_profile_reconnect().
The same gap hit both failure shapes: a clean False return from
_connect_initial_adapter_with_timeout() and an exception escaping it.

Fix: call _schedule_secondary_profile_startup_reconnect() from both startup
failure branches after _safe_adapter_disconnect(). Because secondary adapters
are started mid-start(), before self._running flips True, the regular
scheduler's not-self._running guard would silently drop the request — so the
new bridge parks a background task until startup completes (or shutdown
begins) and then hands off to _schedule_secondary_profile_reconnect()
verbatim: backoff, fresh-adapter rebuild under the profile runtime scope,
slot dedupe, and shutdown cancellation all come from the existing path.
Non-retryable failures are dropped at scheduling time exactly as the regular
scheduler drops them, keeping duplicate-credential/auth-failed startups dead
instead of looping.

Fixes #92064
2026-08-25 22:55:07 -07:00
Shannon Sands f755ed5e90 fix(gateway): surface multiplex profile failures (OOF-3) 2026-08-16 20:32:09 -07:00
Teknium f8d75db026 test(gateway): accept keyword args in _connect_adapter_with_timeout mocks 2026-08-14 21:57:41 -07:00
Teknium 9994bc9ec9 fix(gateway): enforce post-auth normalized reaction observer
Builds on Paolo Antinori's #68431 salvage for #64176. Move plugin dispatch behind the profile-scoped runner authorization boundary, fail closed on malformed reaction identities, preserve observer registration across Telegram app rebuilds, and document the deliberately observer-only contract.

Co-authored-by: Paolo Antinori <pantinor@redhat.com>
2026-08-12 16:42:28 -07:00
Trevin Chow c8f235a106 feat(gateway): allow selective multiplex profile serving 2026-08-10 22:48:24 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
Nick Edson 201be3e546 fix(gateway): reserve retrying Photon listener ownership 2026-07-28 21:45:52 -07:00
Nick Edson c608a6937a fix(gateway): release failed Photon listener claims 2026-07-28 21:45:52 -07:00
Nick Edson a908c62d28 fix(gateway): guard Photon sidecar listener collisions 2026-07-28 21:45:52 -07:00
Nick Edson cf19ac8ff3 fix(gateway): prevent duplicate Photon sidecar storms 2026-07-28 21:45:52 -07:00
StellarisW f57157a128 fix(gateway): recover Discord websocket and event-loop stalls
Replace REST-based Discord liveness probe with local WebSocket/heartbeat
state detection. REST success doesn't prove Gateway event delivery — a
half-closed WebSocket can leave Bot.start() alive while REST returns 200.
Now samples ready/open/ACK state and heartbeat latency; consecutive
unhealthy samples emit one retryable fatal code so GatewayRunner rebuilds
the adapter through the existing reconnect path.

Also fixes three lifecycle gaps in the recovery path:
1. asyncio.wait_for() can remain blocked if adapter cleanup swallows
   cancellation — now uses bounded asyncio.wait() with task detachment.
2. Multiplexed secondary-profile adapters had no profile-scoped reconnect
   owner — now uses one runner-owned reconnect slot per profile.
3. An in-flight turn could send its final text through the disconnected
   adapter after a replacement was registered — now resolves the live
   same-profile replacement for unsent final responses only (message IDs
   never migrate, edits/deletes stay on the old transport).

Adds an opt-in Linux/systemd event-loop watchdog (gateway.systemd_watchdog_seconds,
default 0) for the failure mode where the whole asyncio loop stops making
progress and no in-process liveness task can run. stdlib-only sd_notify,
Type=notify/WatchdogSec generation, READY/STOPPING lifecycle.

Co-authored-by: 王鑫 <wx.xw@bytedance.com>
2026-07-18 20:01:55 +05:30
Teknium c3b2af95e3 test: accept profile_name kwarg in auth-check stubs 2026-07-16 07:17:55 -07:00
Teknium 0cc9426c6d test: feishu port-binding report expects webhook mode after #52563 integration
With the mode-conditional check centralized, default (websocket) Feishu
no longer counts as port-binding in the secondary batch report — pin the
fixture to connection_mode=webhook so the test still exercises the
multi-platform report path.
2026-07-16 07:17:55 -07:00
liuhao1024 9cb3569e97 fix(gateway): allow Feishu websocket mode in multiplex profiles
Feishu was unconditionally listed in _PORT_BINDING_PLATFORM_VALUES,
causing the multiplexer to reject ALL Feishu secondary profiles. But
Feishu in websocket mode (the default) uses an outbound WebSocket
connection and does NOT bind an HTTP port — only webhook/callback mode
needs a listener.

Add _platform_binds_port() helper that checks connection-dependent
platforms (currently only Feishu) against their actual config before
raising MultiplexConfigError. Feishu websocket profiles are now allowed;
Feishu webhook profiles still raise as before.

Fixes #52563
2026-07-16 07:17:55 -07:00
Christopher dd9e75335c fix(gateway): skip port-conflicting multiplex profiles 2026-07-16 07:17:55 -07:00
Ben Barclay 2ea39daeb1 fix(gateway): share relay adapter in multiplex mode (#65366) 2026-07-16 14:30:42 +10:00
Burgunthy a1d6654264 fix(gateway): read adapter token from config for fingerprint check
_adapter_credential_fingerprint only looked at adapter.token directly,
but Discord (and similar) adapters store the bot token on their config
sub-object, not on self. Every Discord adapter in a multiplexed
gateway therefore returned None, the same-token conflict check was
silently skipped, and N adapters all polled the same bot token —
producing a per-message race where whichever adapter won the GIL
answered the user.

Adds a config-token fallback (token, then bot_token) so the check
actually fires for config-backed adapters. Direct adapter.token
still takes precedence when both exist.

Tests cover: config-backed token produces a fingerprint, distinct
tokens produce distinct fingerprints, direct token wins over config,
config without token attributes returns None.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-15 09:50:05 -07:00
Teknium d297568299 fix(gateway): detect config token credential collisions (#59321)
Co-authored-by: markoub <2418548+markoub@users.noreply.github.com>
2026-07-07 18:25:33 +10:00
Ben Barclay d5d02eabb0 feat(gateway): multiplex phase 3 — secondary-profile adapter registry + conflict detection
Bring up adapters for every profile the gateway serves, not just the active
one. Keeps self.adapters as the default/active profile's map (the ~93 existing
self.adapters[...] sites are untouched) and adds secondary profiles under
self._profile_adapters[profile][platform].

- _start_secondary_profile_adapters loops profiles_to_serve(multiplex=True),
  skips the active profile (handled by the primary startup loop), and for each
  other profile loads its gateway config and creates+connects its enabled
  adapters under that profile's _profile_runtime_scope (home + secret scope).
- Each secondary adapter gets _make_profile_message_handler(profile): stamps
  source.profile (when unset) before delegating to the shared _handle_message,
  so the agent turn and session key resolve to that profile.
- Same-platform credential-conflict detection: _adapter_credential_fingerprint
  hashes the adapter's bot token (salted, truncated — never logs the token);
  two profiles claiming the same (platform, token) refuse the duplicate with a
  clear error naming both, since one token can't be polled twice.
- Port-binding hard-error: a SECONDARY profile that enables a port-binding
  platform (webhook, api_server, msgraph_webhook, feishu, wecom_callback,
  bluebubbles, sms) is a config error and aborts startup via MultiplexConfigError
  — the default profile owns the single shared HTTP listener and serves every
  profile through the /p/<profile>/ prefix, so a second bind can only collide.
  Distinct from a transient connect failure (which logs + stays alive to retry):
  a config error writes gateway_state=startup_failed and exits cleanly with an
  actionable message (names the profile, the platform, and the fix). There is no
  valid reason to bind a second port once you've opted into a multiplexer.
- Shutdown tears down secondary adapters alongside the primary ones.
- Defensive getattr guards keep partial-construction unit tests (stop(),
  _run_agent on bare instances) working.

No-op when multiplex_profiles is off (self._profile_adapters stays empty).

Tests: fingerprint stability/log-safety/distinctness, profile message-handler
stamping (and not overriding an already-stamped source), port-binding hard-error
raises + names the profile/platform, non-binding platform is not rejected, and
the guard set covers every TCP-binding adapter.
2026-06-19 07:34:15 -07:00