Commit Graph

83 Commits

Author SHA1 Message Date
teknium1 6a66a5d481 fix(desktop,dashboard): served profile's api_server/webhook read connected with their /p/<profile>/ URL; shared-gateway restart asks first
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.

- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
  webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
  (`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
  adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
  the real one, not a config guess. `hermes status` lists those URLs beside the other
  shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
  when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
  (statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
  shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
  Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
  start/stop on a served profile renders as an inline notice instead of a raw error toast.
2026-09-12 12:52:19 -07:00
Teknium c0d8d6b5d5 feat(multiplex): serve secondary profiles' inbound-port platforms on the default's shared listener
Under gateway.multiplex_profiles a secondary profile with Twilio / LINE / Teams /
BlueBubbles / Microsoft Graph / WhatsApp Cloud / WeCom-callback / Feishu-webhook
credentials was refused WHOLE at config load (SecondaryPortBindingConfigError): every
one of its platforms was skipped because these adapters bind their own port and the
default profile owns the one listener.

The refusal is gone. gateway/platforms/shared_ingress.py gives port-binding adapters
a shared-listener mode: the runner stamps `_shared_listener_profile` on a secondary's
port-binder, `bind_listener()` publishes the adapter's fully wired aiohttp app instead
of starting a TCPSite, and the default listener (api_server, or the webhook adapter
when there is no api_server) forwards `/p/<profile>/<tail>` to the served profile's
adapter whose router matches `/<tail>`, under that profile's runtime scope. The request
is therefore verified by the NAMED profile's adapter with its own secret and replies
leave through that adapter; the un-prefixed path keeps serving the default byte for
byte; an unknown profile or a profile without an adapter for the path is a 404, never
another profile's bot. api_server and webhook stay MIRRORS (the default's own adapter
answers /p/<profile>/ for them) and are the only port-binders a secondary must not
enable; `SHARED_LISTENER_MIRROR_PLATFORMS` is that set in gateway/config.py.

The adapter records its public callback URL (`ingress_url`) in gateway_state.json
under `<profile>:<platform>` and logs it once at connect, so the operator knows what
to paste into the vendor console.
2026-09-12 01:53:15 -07:00
Teknium bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00
Teknium 9848e22ed6 feat(multiplex)!: drop gateway.multiplex_profile_allowlist — serve every profile
The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.

Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.

BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
2026-09-11 19:50:46 -07:00
Teknium c3e01c753d fix(gateway): /p/<profile>/ webhook delivery and api_server callbacks stay on the routed profile
Under gateway.multiplex_profiles a /p/<profile>/ webhook route (or one with
`profile: <name>`) executed under that profile but delivered its reply as the
FIRST profile owning the target platform: `_find_adapter` took
`runner.adapters` then iterated `_profile_adapters` in dict order, and the
home-channel fallback read `runner.config` (the default profile's). The gh leg
for `github_comment` inherited the process environ, i.e. the default profile's
GH_TOKEN. Both directions leaked: a secondary route posted through the default
bot, and a default route borrowed a platform parked only on a secondary.

The api_server `/p/<profile>/api/platforms/<platform>/events` callback had the
same shape — `_get_platform_callback_adapter` read `runner.adapters` regardless
of `_api_request_profile`, so a secondary's Google Chat / Teams events were
verified and dispatched by the default adapter.

Now:
- webhook: `_delivery_info` / the deliver_only dict carry the resolved
  profile; `_find_adapter(platform, profile)` resolves through the runner's
  shared fail-closed `_authorization_adapter`; the delivery leg runs inside
  `_profile_scope(profile)` and takes the home channel from that profile's
  `load_gateway_config()`; `gh` gets GH_TOKEN/GITHUB_TOKEN from the profile
  secret scope (default's values dropped from the child env when the profile
  has none).
- api_server: the callback adapter resolves via
  `_authorization_adapter(platform, _api_request_profile.get())`; a named
  profile without the adapter is a 503, never the primary's adapter.

A profile without the target platform fails closed ("not connected" → 502 /
503) instead of a silent cross-profile send.

Fixes #65939
Fixes #84266

Co-authored-by: 604maestro <604maestro@protonmail.com>
Co-authored-by: mjshorty <mjshorty@users.noreply.github.com>
Co-authored-by: StellarisW <stellarisw@users.noreply.github.com>
2026-09-10 18:16:43 -07:00
kshitijk4poor ab2f4602de refactor: MessageEvent to gateway/platforms/event.py; ElicitationHandler takes a call_context thunk
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.

gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).

tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.

ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)

Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
2026-09-07 22:47:33 +05:30
Teknium e40192118e fix(webhook): only webhook-signature commits a delivery to the Standard Webhooks path
GitLab (lib/gitlab/web_hooks.rb, app/services/web_hook_service.rb) sends
webhook-id and webhook-timestamp on EVERY delivery and adds webhook-signature
only when a signing token is configured. Selecting the HMAC path on any
webhook-* header would 401 every legacy X-Gitlab-Token install the moment it
upgrades to GitLab 19, so only the signature header selects the path; id and
timestamp then travel with it and the validator still fails closed when either
is missing. svix-* keeps its existing any-header selection.

Tests trimmed to the invariant bar: one parameterized contract (whsec_ and raw
secrets accept; wrong secret, tampered body and stale timestamp reject) plus the
GitLab legacy-token coexistence contract. evals/webhook_auth/standard_webhooks_ab.py
drives a real aiohttp WebhookAdapter on a dedicated loopback port with real
signed requests for before/after evidence.

Related: #47849 (HwangJohn, cherry-picked here), #92024 (earlier salvage of
#47849), #102080 and #103167 (same alias fix, same target).
2026-09-06 05:35:48 -07:00
HwangJohn cb3b5bb64c feat(webhook): accept standard webhook signatures 2026-09-06 05:35:48 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 4009eeb59f refactor(gateway/platforms): webhook/msgraph/keyboards/chunked_upload — walrus guards, fold small helpers 2026-09-02 23:13:26 -07:00
Teknium 0452566961 refactor(gateway/platforms): signal.py/webhook.py — fold multi-line messages and calls, walrus guards 2026-09-02 22:47:05 -07:00
Teknium 1005056619 refactor(gateway/platforms): bluebubbles.py second pass — walrus guards, fold layouts, compact docs 2026-09-02 22:24:40 -07:00
Teknium af1e13f13e refactor(gateway/platforms): webhook.py second pass — unify replay-window check, compact comments/layout 2026-09-02 22:08:30 -07:00
Teknium 20ec9ef9d3 refactor(gateway/platforms): webhook.py — table-driven simple-provider signature checks (constant-time compare kept) 2026-09-02 20:00:09 -07:00
Teknium c02feadce3 refactor(gateway/platforms): webhook.py — _is_known_platform helper, prune loop collapse 2026-09-02 19:55:40 -07:00
Teknium f46a9d4341 refactor(gateway/platforms): webhook.py — _peek_session_id helper, merged session-close guards 2026-09-02 19:45:18 -07:00
Teknium eb4222f415 refactor(gateway/platforms): webhook.py — _dispatch_agent_run phase helper, compact docstrings 2026-09-02 19:43:16 -07:00
Teknium c4a6970d87 refactor(gateway/platforms): group D minor layout packs 2026-09-02 19:37:41 -07:00
Teknium 223593621c refactor(gateway/platforms): webhook.py — suppress() for swallowed errors, one-line section banners, extra local 2026-09-02 19:19:49 -07:00
Teknium ba57157cf0 refactor(gateway/platforms): group D docstring closer folds 2026-09-02 19:10:08 -07:00
Teknium 7b831c7145 refactor(gateway/platforms): group D layout hug + string-literal joins (AST-neutral) 2026-09-02 19:01:23 -07:00
Teknium 50556a4d86 refactor(gateway/platforms): webhook.py — route/skill/validation phase helpers, module-level Svix validator, shared timestamp-age check 2026-09-02 18:33:02 -07:00
Teknium c4f96565d0 refactor(adapters/whatsapp_webhook): 6449->4465; whatsapp_cloud/common/plugin send + media unification, webhook route/filter dispatch tables 2026-09-02 14:06:38 -07:00
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
kshitijk4poor c0ff25a1f8 fix(gateway): stop blocking the event loop — off-loop hot sites + ASYNC lint ratchet
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).

Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
  now runs via asyncio.to_thread — a webhook delivery no longer freezes
  every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
  5s worst case) become await asyncio.sleep() (re-lands #36163 at current
  line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
  the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
  move off the loop.

Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
  (blocking HTTP / Popen / subprocess.run / time.sleep in async def).
  These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
  sites (detached restart watchers; router sweep in flight via #84376;
  two platform adapters), each documented for burn-down. New files or
  new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).

Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
  subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
  liveness DURING delivery (ticker coroutine): 1 tick on the old
  blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.

Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
2026-08-28 07:50:51 -07:00
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
Teknium 6cb1085d3d fix(gateway): /p/<profile>/ on a non-multiplex gateway fails closed instead of serving the owner profile
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.

Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.

Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.

Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.

Fixes #91583 (defect 2). Repro and live validation by @kubaboski.
2026-08-23 03:57:47 -07:00
liuhao1024 f13f3401a1 fix(webhook): authenticate Linear deliveries via linear-signature HMAC (#87348) 2026-08-16 01:45:37 -07:00
Teknium e4aeb65599 feat(webhook): per-route toolset overrides for webhook agent runs
Webhook agent runs default to the constrained hermes-webhook toolset
(web/vision/clarify) because payloads can carry untrusted third-party
content. That default is right for public webhooks but wrong for trusted
local pushes (e.g. an OOM monitor daemon that needs the agent to run
ps/free/py-spy): the only workaround was widening platform_toolsets.webhook,
which elevates EVERY webhook route at once.

This adds a 'toolsets' key on individual webhook route configs (static
routes in config.yaml and dynamic subscriptions in
webhook_subscriptions.json) that replaces the platform-level resolution
for that route only:

- BasePlatformAdapter.toolsets_for_source(): per-source override hook,
  default None (no behavior change for any other platform).
- WebhookAdapter.toolsets_for_source(): maps the session chat_id
  (webhook:{route}:{delivery_id}) back to its route config and returns
  the route's toolsets list.
- GatewayRunner._resolve_enabled_toolsets_for_source(): shared resolver
  used by both agent-run call sites; validates the override through the
  SAME _get_platform_tools path as platform config, so unknown names and
  platform-restricted toolsets (e.g. discord_admin) are dropped rather
  than trusted.

Deliberately NOT exposed via 'hermes webhook subscribe': granting elevated
tools is a manual config edit only, so an agent-created subscription
cannot self-grant terminal at runtime.

Cache-safe: the toolset list is resolved before agent construction and is
constant for a route, so the per-session agent signature and frozen system
prompt are unaffected mid-conversation.
2026-08-13 01:51:19 -07:00
Trevin Chow c8f235a106 feat(gateway): allow selective multiplex profile serving 2026-08-10 22:48:24 -07:00
teknium1 5b751dc0ad chore: remove unused imports and dead locals (ruff F401/F841 sweep)
Cleans F401 unused imports and F841 dead local assignments across
root *.py, agent/, hermes_cli/, tools/, gateway/, cron/, tui_gateway/
(tests/, plugins/, skills/ excluded).

Intentionally KEPT (false positives / test-patch surfaces):
- agent/transports/__init__.py package re-exports
- cli.py browser_connect re-exports (DEFAULT_BROWSER_CDP_URL area,
  used by tests/cli/test_cli_browser_connect.py)
- hermes_cli/main.py _prompt_auth_credentials_choice /
  _model_flow_bedrock_api_key (accessed via main_mod attr in tests)
- gateway/run.py aliased replay_cleanup + whatsapp_identity re-exports
  and _PORT_BINDING_PLATFORM_VALUES (test-referenced)
- hermes_cli/web_server.py get_running_pid (tests monkeypatch it) and
  _OAUTH_TOKEN_URL availability probe
- hermes_cli/config.py get_process_hermes_home re-export (noqa'd F811
  chain) and yaml availability-probe import
- hermes_cli/nous_subscription.py managed_nous_tools_enabled
  (tests patch hermes_cli.nous_subscription.managed_nous_tools_enabled)
- try/except ImportError availability probes (env_loader, tts_tool,
  mcp_tool, web_server anthropic OAuth block)
- tools/web_tools.py noqa F401 re-exports
- hermes_cli/setup_whatsapp_cloud.py:263 'proceed' skipped: possible
  missing-guard bug, flagged for separate review
- unused function parameters (signature changes out of scope)

Side-effect RHS calls preserved where only the binding was dead
(e.g. web_server proc = _spawn_hermes_action -> bare call).
2026-07-29 11:53:39 -07:00
Atakan 8c5e846536 fix(gateway): bind HTTP auth to routed profiles 2026-07-28 14:22:18 -07:00
teknium1 136f8dab67 refactor(gateway): promote autonomous silence matcher to shared response_filters helper
Follow-up to the salvaged #71756: instead of webhook importing cron's
private _is_cron_silence_response, the loose autonomous-lane matcher now
lives in gateway/response_filters.py as is_autonomous_silence_response,
sharing LIVE_GATEWAY_SILENT_MARKERS with the interactive exact-marker
rule so the marker sets can never drift. Cron and webhook both delegate
to it. Interactive gateway behavior unchanged.
2026-07-26 16:55:42 -07:00
Gumclaw 55d3272286 fix(webhook): honor [SILENT] when the agent explains its own silence
A webhook route that answered `[SILENT]` still delivered, whenever the model
added a sentence saying why it was staying quiet:

    [SILENT]

    The new inbound was the same email quoted back a second time, on a ticket
    we already answered. Nothing new to reply to, so I closed it.

Webhook subscription prompts tell the agent to answer `[SILENT]` on a tick that
produced no story — a duplicate inbound, a stand-down because a sibling lane
already replied, a routine close. Nobody is waiting on the other end of a
webhook, so a "nothing happened" message has no reader.

Delivery went through the live gateway's `is_intentional_silence_response`,
which requires the response to be EXACTLY a marker. That rule is right for an
interactive chat: swallowing a real answer because it opens with a marker is
much worse than showing a stray marker. It is the wrong trade for an autonomous
lane, where a leaked non-story is a pointless notification on every tick and
models reliably append the explanation that flips the check back to "deliver".
Cron already resolved this the other way — `cron/scheduler.py` treats a marker
on its own first or last line as silence — so the two autonomous lanes
disagreed while the interactive path was fine.

Suppress in `WebhookAdapter.send`, before the deliver-type switch, so every
route (log, github_comment, cross-platform) behaves the same. Reuses cron's
`_is_cron_silence_response` rather than restating the rule, so the two lanes
cannot drift; prose that merely mentions a marker mid-sentence still delivers.
The interactive gateway path is untouched.

Tests: six cases in tests/gateway/test_webhook_adapter.py — bare marker,
marker + trailing prose (the reported shape), marker on the last line, a real
report, a report quoting a marker mid-sentence, and a `log` route. Verified
red-first: with the suppression removed the three silence cases fail
("Expected send to not have been awaited") while the three delivery cases still
pass, so the tests assert the fix rather than the framework.
2026-07-26 16:55:42 -07:00
Stoltemberg c89481db5e fix: add explicit UTF-8 encoding to all subprocess text=True calls (#53428)
On Windows with Chinese locale (GBK), subprocess.run(text=True) without
explicit encoding causes UnicodeDecodeError crashes. This fix adds
encoding='utf-8', errors='replace' to all subprocess.run() and
subprocess.Popen() calls that use text=True across 76 non-test Python files.

Fixes #53428 (master tracker for Windows GBK locale crash).

Note: credential_pool.py and electron changes excluded per reviewer request —
those will be submitted as separate focused PRs.
2026-07-24 11:45:57 -07:00
Teknium 9fc0074bac fix(gateway): unify reset boundaries vs recovery — promote accidental ends, honor mode=none, adapter-aware resume guidance
Unifies the two gateway subsystems that were fighting each other: the
'never lose a session' recovery machinery (#54878 stale-route self-heal,
find_latest_gateway_session_for_peer reopening agent_close/ws_orphan_reap
rows) and the session reset/expiry machinery (expiry watcher, /new,
/resume, resume_pending freshness gate).

The unified contract:
- INTENTIONAL boundaries (expiry finalization, auto-reset, /new,
  /resume switch) are recorded durably via promote_to_session_reset(),
  which upgrades accidental recoverable end_reasons (agent_close,
  ws_orphan_reap) to the explicit boundary while preserving other
  explicit reasons (compression, etc.). Recovery then correctly refuses
  to resurrect them.
- ACCIDENTAL ends (crash, cleanup bug, mistaken reaper) stay
  recoverable — genuine crash recovery is untouched.

On top of the cherry-picked contributor commits:
- promote_to_session_reset widened to ws_orphan_reap + parameterized
  reason so auto-reset paths stay auditable (idle/daily/suspended/
  resume_pending_expired) (#61220, #61993, #63539)
- get_or_create_session auto-reset, reset_session (/new), and
  switch_session (/resume) all write through the promote path — the
  first-reason-wins end_session no-op could previously leave a reset
  session resurrectable behind a stale agent_close row (#61993)
- resume_pending freshness gate now honors session_reset.mode=none:
  explicit opt-out of automatic resets also opts out of the zombie
  gate (#61052)
- resume recovery note extracted to build_resume_recovery_note() and
  made adapter-aware via a new interactive_resume capability flag:
  webhook/api_server auto-resume turns now CONTINUE the interrupted
  task instead of emitting an unanswerable 'session restored'
  acknowledgement that abandoned the work (#57056)
- tests updated to call the real note builder instead of mirroring it

E2E-validated against a real SessionDB + SessionStore in a temp
HERMES_HOME: expiry->agent_close->no-resurrection, /new promote,
crash recovery preserved, mode=none opt-out, routing-table flag sync.
2026-07-17 04:51:39 -07:00
Drexuxux 1b69c47e97 fix(webhook): reject a non-ASCII signature header instead of crashing the endpoint
_validate_signature backs the public webhook receiver. It compared each
attacker-supplied signature/token header (GitHub X-Hub-Signature-256,
GitLab X-Gitlab-Token, generic X-Webhook-Signature / -V2, and the Svix v1
header) against a computed hex/base64 digest with hmac.compare_digest on
two str values. compare_digest raises TypeError on a str containing
non-ASCII characters, and the header is raw client input on an
unauthenticated endpoint — so any internet client could POST a single
non-ASCII byte in the signature header and raise out of the handler,
returning a 500 instead of a clean 401. Fail-closed, but an on-demand
crash of the request path.

Route all five comparisons through a small _hmac_str_equal() helper that
encodes both sides to UTF-8 bytes before the constant-time compare
(compare_digest has no ASCII restriction on bytes). Semantics are
unchanged for valid signatures; a hostile non-ASCII header now fails
closed with a rejection instead of raising.

Adds regression tests: non-ASCII GitHub/GitLab/generic/V2 signature
headers return False (no raise), and a non-ASCII configured secret still
matches its exact token value.

Also maps drexux0@gmail.com in scripts/release.py AUTHOR_MAP.
2026-07-16 07:22:24 -07:00
Agung Subastian 8fc989b416 fix(gateway): multiplex secret_scope for authz, Slack, webhooks
Secondary profiles under gateway multiplex keep tokens/allowlists in
profile secret_scope, not process os.environ. Auth and Slack were still
reading os.getenv, so Slack on a secondary profile failed allowlist and
socket mode. Webhook deliver also only looked at default adapters.

- Prefer get_secret for allowlists / allow-all flags (authz_mixin)
- Slack app token + allowlist via secret_scope with getenv fallback
- Wrap secondary profile message handlers in _profile_runtime_scope
  before auth runs
- Resolve home-channel env from secret_scope / PlatformConfig
- Webhook deliver falls back to _profile_adapters for target platform
- Template key event_type for webhook prompts
2026-07-16 05:39:58 -07:00
Teknium 9420ad946a fix(webhook): scope reuse_address=False to macOS only (#65482)
On Linux, SO_REUSEADDR only allows rebinding past TIME_WAIT (a second
live listener would need SO_REUSEPORT, which we never set), so
disabling it bought no protection there while making a quick gateway
restart fail to rebind for up to ~60s. Keep the BSD silent-split guard
on darwin, default semantics elsewhere.

E2E verified: dual-stack v4+v6 bind on one port, immediate rebind
after disconnect, and live-listener conflict still rejected.
2026-07-16 01:28:18 -07:00
kshitijk4poor 92876effe2 fix(webhook): make dual-stack bind exclusive
Disable address reuse so an existing family-specific listener cannot silently split traffic with the webhook server. Normalize wildcard bind hosts for local CLI URLs and align setup documentation with the dual-stack default.
2026-07-16 12:36:51 +05:30
Ben d542894adf fix(webhook): default to dual-stack bind so 6PN (IPv6) can reach the adapter
The webhook adapter defaulted to host='0.0.0.0' — IPv4 only. On Fly.io
hosted agents the edge router (hermes-agent-router) reverse-proxies public
webhook traffic to <app>.internal:8644 over 6PN, Fly's private network,
which is IPv6-only (.internal resolves to an fdaa:… address). An IPv4-only
listener is unreachable there, so public webhook POSTs to
https://<agent>.agents.nousresearch.com/webhooks/<route> never landed on
the adapter — the router's dial was refused.

Fix: DEFAULT_HOST = None, which makes aiohttp/asyncio create_server bind
BOTH address families. '::' is NOT a valid substitute: on hosts where the
kernel sets bindv6only=1 (verified on Fly machines) it yields an IPv6-only
socket, breaking the IPv4 loopback /health check and the AF_INET
port-conflict probe in connect(). None binds per-family regardless of the
sysctl. An explicit empty-string/null host in config now also normalises to
None (dual-stack) rather than an invalid host=''. Users can still pin a
specific host via platforms.webhook.extra.host.

Validated live on a Fly staging agent: with this default and no config
override, the adapter binds both v4 and v6 (127.0.0.1:8644 and [::1]:8644
both answer), and a public signed webhook POST through the router returns
202 (valid sig) / 401 (bad sig) instead of the router's 502.

Tests: new TestDualStackBind asserts the None default, config resolution
(missing/empty→None, pinned preserved), and a real dual-stack bind opens
both AF_INET and AF_INET6 listeners. Red-proof: these fail on the old
'0.0.0.0' default.
2026-07-16 12:36:51 +05:30
teknium1 ae5e39005b fix(gateway): run webhook route scripts off the event loop + AUTHOR_MAP entry
- run_route_script shells out with subprocess.run (up to 30s timeout); wrap
  the call in asyncio.to_thread so a slow script can't stall every other
  webhook and gateway task on the loop.
- scripts/release.py: map grace@weeb.onl -> evelynburger for the salvaged
  contributor commit.
2026-07-08 08:10:55 -07:00
Grace 0cf2e39c41 feat(gateway): add webhook payload filters 2026-07-08 08:10:55 -07:00
MorAlekss d577408f3f fix(webhook): reject generic V2 signature missing timestamp instead of falling back to V1 2026-07-05 02:25:38 -07:00
teknium1 708b57e009 fix(webhook): rate-limit V1 deprecation warning + document V2 signature
- warn once per route instead of on every request (busy senders would
  spam the log)
- document X-Webhook-Signature-V2 / X-Webhook-Timestamp in the webhooks
  user guide

Follow-ups for salvaged #58461.
2026-07-05 01:36:53 -07:00
MorAlekss 70449a4939 fix(security): add timestamp-bound V2 signature for generic webhook replay protection 2026-07-05 01:36:53 -07:00
Gutslabs ec29590a0f fix(webhook): enforce body-size limit on chunked requests
The webhook adapter enforced max_body_bytes only via the Content-Length
header; a Transfer-Encoding: chunked request (content_length=None) or a
spoofed small Content-Length bypassed the cap entirely and read the full
body (bounded only by aiohttp's implicit 1 MiB default, above any
operator-configured smaller limit).

- web.Application(client_max_size=max_body_bytes): aiohttp enforces the
  cap on every read path, chunked included
- catch HTTPRequestEntityTooLarge -> 413 (was swallowed into generic 400)
- post-read length re-check as defense in depth
- chunked-upload regression test

Manual port of PR #3955 by @Gutslabs onto current main (handler had
been restructured since); authorship preserved.
2026-07-04 15:33:31 -07:00
Teknium 64ed99a6e6 fix(webhook): close per-delivery session at the true end of the run (#57423)
The merged webhook session-close fix (#57370, salvaging #57322) wrapped
handle_message in a try/finally — but BasePlatformAdapter.handle_message
is fire-and-forget: it spawns _process_message_background and returns
before the agent run starts. The finally-close therefore ran BEFORE
get_or_create_session created the session row, found no session_id, and
silently no-op'd — the ghost-session leak persisted on the real path.
(The shipped test masked this by stubbing handle_message with a fake
that created the row synchronously.)

Move the close to an on_processing_complete override — the lifecycle
hook the base class fires at the TRUE end of the run, on the success,
failure, and cancellation paths alike. Empirically verified through the
real fire-and-forget pipeline: before, ended_at stayed NULL; after,
ended_at is set with end_reason=webhook_complete and the row is
prunable.

Tests now stub only the runner-side _message_handler (the seam the live
gateway injects) so handle_message / _process_message_background /
on_processing_complete all run for real; adds an AsyncSessionDB-facade
coverage test for the coroutine-await branch.
2026-07-02 17:39:09 -07:00
kshitijk4poor 65cb70b8d0 refactor(gateway): add SessionStore.peek_session_id public accessor for webhook close
Replace the webhook delivery-close path's direct reach into private
SessionStore._entries (which also bypassed the store lock) with a public,
lock-held peek_session_id(session_key) accessor. Mirrors the existing
lookup_by_session_id inverse helper. Keeps a getattr fallback for older
stores / test doubles. Adds a unit test for the accessor.
2026-07-03 03:26:53 +05:30
Gumclaw 14882bab7e fix(gateway): close webhook sessions on delivery completion so prune can reap them
Webhook deliveries created a unique one-shot session (delivery_id baked into
the session key at gateway/platforms/webhook.py:668) but the adapter fired
handle_message via asyncio.create_task WITHOUT ever ending the session
(webhook.py:713, pre-fix). Nothing else closes it: the gateway caches/expires
the agent per session_key but never calls end_session for the webhook path,
and _end_session_on_close teardown doesn't run for these fire-and-forget tasks.

SessionDB.prune_sessions (hermes_state.py:4965) only deletes rows WHERE
ended_at IS NOT NULL. So every webhook session stayed with ended_at NULL ->
unprunable -> unbounded state.db growth. This was the primary driver of the
SQLite lock-contention gateway outage.

Fix: wrap the delivery in _run_delivery_and_close, which awaits
handle_message and then (in finally, so failures still reap) calls
_end_webhook_session -> SessionDB.end_session(session_id, 'webhook_complete').
This mirrors how cron closes its session with 'cron_complete'
(cron/scheduler.py:3065). end_session is first-reason-wins and no-ops on an
already-ended row, so it never clobbers a compression/agent_close reason.

Adds tests/gateway/test_webhook_session_close.py asserting the invariant
(a completed webhook session has ended_at set + is prunable), including the
error-path case, against a real SessionStore + SessionDB.
2026-07-03 03:26:53 +05:30