Commit Graph

3388 Commits

Author SHA1 Message Date
teknium1 735831f776 fix(mcp): reconcile chore lives in run_profile_reconcile; prune lazy + mid-connect servers
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.

reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.

test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
2026-09-13 06:29:12 -07:00
Teknium 4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
kshitijk4poor dbd027287e refactor(telegram): tighten the dispatch-stall check shape from review
- Correct the per-generation reset comment: an in-place updater restart keeps
  PTB's update_queue, so old-generation dispatches can briefly exceed
  received; the check already treats that as no backlog.
- Make the once-per-stall gate explicit (cap the heartbeat count) instead of
  relying on `!=`.
- Drop the dead getattr in _record_updates_received (only reachable after
  _record_polling_progress dereferenced the same instance state); keep the
  fallbacks in the heartbeat check and group-99 handler, which sibling
  watchdogs share because object.__new__ adapter doubles exist in tests.
- Split the single invariant test so a failure names the broken guard:
  stall-once, re-arm-on-progress, generation-reset; drop the unused _app mock.
2026-09-13 18:55:15 +05:30
kshitijk4poor 78b98032c5 fix(telegram): key the deaf-ingress report on dispatcher progress, not update age
Follow-up to the cherry-picked #102383 commit. The check as written was neither
sensitive nor specific:

- It aged the newest received update, so a wedged PTB dispatcher was never
  reported while new updates kept arriving more often than every 300s (probe:
  1 update/250s for an hour -> 0 reports).
- `delivered` counted only MessageEvents reaching the gateway handler, while
  `received`/`dispatched` counted every Update; a single handled callback_query,
  reaction, unauthorized user or unmentioned group message produced a false
  ERROR after 300s of quiet.
- `_record_updates_received` skipped the generation/teardown guard
  `_record_polling_progress` applies, and the counters never reset across
  polling generations, so a late response from a fenced poll or a reconnect
  inflated the backlog.

Now `received` and `dispatched` count the same population (every fetched
update reaches the group-99 catch-all) and the report fires when a backlog
persists with no dispatch progress across two 90s heartbeats, once per stall,
re-armed on progress, reset per generation, at WARNING (diagnostic only;
#71240 owns recovery). The delivered counter and the `note_inbound_delivered`
facade method are dropped; the once-per-adapter "no message handler" error on
BasePlatformAdapter.handle_message stays. `_record_polling_progress` returns
whether the round-trip was accepted so the received stamp reuses its gate.
Tests trimmed to two invariants; every guard proven red by mutation.

Refs #102260
2026-09-13 18:55:15 +05:30
joaomarcos db407dd078 fix(telegram): report a healthy-but-deaf ingress instead of nothing (#102260)
Every Telegram health probe measures the transport. A getUpdates round-trip
that returns 200 proves bytes are moving and nothing else: the stall watchdog
(#92991), the pending-update probe (#42909/#55769), the get_me() heartbeat
(#66377) and the polling-progress instrumentation all stay green while updates
arrive and then die downstream. The adapter then publishes "connected", logs
nothing at all, and is indistinguishable from a bot nobody has messaged.

That is #102260: three weeks of telegram.state "connected" plus "polling
confirmed healthy: getUpdates progressing (generation 1)" with zero inbound
reaching the agent, surviving every restart. Two of the issue's three
hypotheses do not hold on this code — _record_polling_progress fires on every
round-trip (not only at start_polling), and _send_path_degraded is cleared on
the first confirmed round-trip — and the reporter's own observation that fresh
messages are received but not processed places the failure downstream of the
transport, in the one stretch with no instrumentation at all.

Add the missing delivered side of the accounting:

- received: updates Telegram handed the process, read from the getUpdates
  envelope the adapter already parses (an empty result proves the transport,
  not arrival, so only non-empty results count).
- dispatched: updates PTB's dispatcher carried through the whole handler
  chain, stamped in the existing group-99 catch-all before its early returns.
- delivered: inbound events that reached the gateway's message handler,
  stamped in BasePlatformAdapter.handle_message for every platform.

_check_ingress_delivery_gap runs on the existing heartbeat and, when updates
arrived but nothing was delivered for 300s, names the broken hop: received >
dispatched means the dispatcher is not draining, dispatched > delivered means
Hermes is dropping what arrives. Diagnostic only — a received update
legitimately reaches no gateway turn, and reconnecting a healthy transport
cannot repair a dropped update, so this never drives recovery.

Also make the silent discard on the shared funnel speak: handle_message
returned with no log when no message handler was installed, so a mis-wired
adapter discarded 100% of inbound while connected and able to send.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018pT5hFJBRfLj8KMqFhm3qz
2026-09-13 18:55:15 +05:30
teknium1 0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
Will Lynas 9939e3375e fix(slack): render compact tool previews as inline code 2026-09-13 05:34:48 -07:00
teknium1 008caa88a2 fix(platforms): extra_or_secret keeps blank-string YAML values where the old readers did
dingtalk _extra_get, mattermost _extra_or_env and slack _extra_or_env_flag/_channel_set fell
through to env only on None, so `allowed_channels: ""` / `free_response_channels: ""` meant
"no whitelist" rather than "use the env CSV". The shared reader treated blank as unset and
silently widened those to the env value. New `blank_is_unset=False` knob restores the old
semantics at those seven call sites; the default (blank = unset) stays for the readers whose
old body was `extra.get(k) or env`.
2026-09-13 05:32:38 -07:00
teknium1 c0d7b05faa fix(whatsapp_cloud): allowlist "*" admits DM intake again
The mixin routes DM intake through _is_dm_allowed, and the Cloud override was a bare
wa_id set lookup with no "*" handling, so allow_from={"*"} (documented for
WHATSAPP_ALLOWED_USERS, inherited by Cloud) went from admitting intake on main to
denying it. Cloud now overrides _entry_matches instead: bare-wa_id membership, then
the shared WhatsApp matcher ("*" + aliases) — one predicate for strict DM auth, DM
intake and groups. whatsapp_cloud joins the parity matrix; a new wildcard row covers
every wildcard host on all three paths (sabotage: revert -> [whatsapp_cloud] fails).
2026-09-13 05:32:38 -07:00
teknium1 7370c82c2d fix(wecom): WECOM_ALLOW_ALL_USERS opens DMs again; mixin refuses hosts with no env prefix
WeComAdapter mixed in OwnAccessPolicyMixin without ALLOW_ALL_ENV_PREFIX, so
_allow_all_env_names() read "_ALLOW_ALL_USERS" and every open-DM WeCom deployment
(setup still writes WECOM_ALLOW_ALL_USERS) silently denied all DMs.

- WeComAdapter.ALLOW_ALL_ENV_PREFIX = "WECOM"
- OwnAccessPolicyMixin.__init_subclass__ raises TypeError on an empty prefix so the
  omission cannot ship again (every other host already sets one).
- Parity test now runs a third matrix row that sets each host's own
  <PREFIX>_ALLOW_ALL_USERS; the GATEWAY-only row is why this slipped through.
  Sabotage: prefix removed -> [platform] row fails even with the guard reverted.
2026-09-13 05:32:38 -07:00
teknium1 de114b3af1 refactor(platforms): one scoped-secret reader and spec-driven enablement/YAML-bridge boilerplate across all adapters
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:

- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
  docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
  `external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
  Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
  keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
  (SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
  bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
  slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
  `_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.

Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:

- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
  irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
  matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
  nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
  `yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
  mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.

Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
  neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
  seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
  env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
  does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
  all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
  "Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
  reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
  as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
  value (was list/tuple/set only) — same env text for every real YAML shape.

Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.

Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
2026-09-13 05:32:38 -07:00
teknium1 9b1990583d refactor(gateway): adapters share helpers.cancel_task / MessageDeduplicator / bounded_put
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).

- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
  task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
  definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
  whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
  wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
  max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
  LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
  refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
  `~/.hermes`, so profiles with HERMES_HOME set find them.

Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.

Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
2026-09-13 05:32:38 -07:00
teknium1 3a179fe524 refactor(gateway): every Ogg/Opus voice transcode routes through base.transcode_to_ogg_opus
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).

`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.

Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
2026-09-13 05:32:38 -07:00
teknium1 8185834bc9 refactor(gateway): one OwnAccessPolicyMixin replaces five copied DM/group intake predicates
Weixin, WeCom, QQBot, WhatsApp (common + cloud) and Yuanbao's AccessPolicy each carried
their own `_open_dm_opted_in` / `_is_dm_allowed` / `_is_dm_intake_allowed` /
`_is_group_allowed`, differing only in the platform prefix of the allow-all env var and in
small drifts. The same "allow-all must be scoped / fail closed" fix has landed on this trio
at least four times; a shared rule means it lands once.

gateway/platforms/access_policy_mixin.py::OwnAccessPolicyMixin owns the predicates,
reads every env name through the scoped `get_scoped_secret` and exposes two hooks:
`_entry_matches` (platform allowlist matching) and `_live_dm_allow_from` (env-seeded
lists re-read live). `ALLOW_ALL_ENV_PREFIX` is the only per-adapter datum.

Retained overrides (behavior genuinely differs):
- whatsapp_cloud `_allow_all_env_names` adds WHATSAPP_CLOUD_ALLOW_ALL_USERS;
  `_is_dm_allowed` keeps its bare-wa_id normalisation.
- whatsapp_common `_entry_matches` -> phone/LID alias matching; `_live_dm_allow_from`.
- wecom `_is_group_allowed(chat_id, sender_id)` adds the per-group sender allowlist on
  top of the shared chat-level rule; `_entry_matches` strips `wecom:user:` prefixes.
- qqbot `_entry_matches` (case-insensitive, `*`).
- yuanbao `AccessPolicy.is_group_allowed`: `open` groups still require the allow-all
  opt-in (no runner-side mention gate), so it wraps the shared rule.

Behavior change: Weixin `_is_dm_intake_allowed` now denies a blank/whitespace principal
(the other four already did; safe variant chosen). Weixin's inline group gate now goes
through `_is_group_allowed` with identical verdicts.
2026-09-13 05:32:38 -07:00
teknium1 81d77c280a refactor(platforms): photon rides the base send-retry loop; slack/discord parse Retry-After with the shared parser
Photon forked BasePlatformAdapter._send_with_retry before three fixes landed there: the server's
retry_after is honoured over exponential backoff, a long server penalty (> 60s) returns a typed
failure instead of sleeping inline (#91969), and an exhausted rate-limited send no longer posts the
failure notice inside the flood penalty. The fork got none of them. Photon's two genuine differences
are now hooks on the base — `_send_retry_is_final(result)` (structured auth/target refusals are
returned as-is, no retry, no plain-text resend) and `_send_plain_fallback(...)` (no Markdown banner,
richlink() bypassed) — and the 42-line fork is deleted.

Slack's `_retry_after_from_exc` and Discord's `_extract_discord_retry_after` parsed the header by
hand and only understood the numeric form; both now call agent.retry_utils.parse_retry_after_seconds
(numeric or HTTP-date, either header casing). Discord keeps its `retry_after` attribute path, the
`X-RateLimit-Reset-After` fallback and the 1s floor.

Behavior change: Photon retries now add up to 1s of jitter to the backoff and honour a server
retry_after; Slack/Discord recognise an HTTP-date Retry-After they previously ignored.
2026-09-13 05:21:39 -07:00
teknium1 73eadd54f5 fix(platforms): standalone senders return redacted error envelopes
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.

`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.

Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
2026-09-13 05:21:39 -07:00
teknium1 ad305bead5 refactor(gateway): send_exec_approval is a base template method; 9 adapters only render buttons
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.

BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.

The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.

Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
  fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
2026-09-13 05:21:39 -07:00
teknium1 aac36fd96f refactor(gateway): text-batch flush lives on BasePlatformAdapter; shield + cancel-race fixes reach all 8 adapters
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.

The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.

`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
2026-09-13 05:21:39 -07:00
teknium1 74a315bd32 fix(platforms): irc/line/buzz identity-lock conflicts actually fire
gateway.status.acquire_scoped_lock returns (acquired, existing_record). The irc, line and
buzz adapters tested `if not acquire_scoped_lock(...)`, and a non-empty tuple is always
truthy, so two profiles could drive one IRC nick / LINE channel / Buzz identity in
parallel. Route the three through BasePlatformAdapter._acquire_platform_lock (the seam
the other 8 adapters use), which unpacks the tuple, names the owning profile + PID in the
fatal error and honours the `--replace` takeover. Release goes through
_release_platform_lock; the private _lock_key bookkeeping is gone.

Error code changes from `lock_conflict` to `{scope}_lock` — both families are already
matched by gateway.restart.is_global_startup_conflict.

The buzz test mocked acquire_scoped_lock as a bare False, which masked the bug; it now
returns the real (False, record) contract, and irc/line gain the same conflict test.
2026-09-13 05:21:39 -07:00
teknium1 11576390fe refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.

Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.

Sites -> canonical:
  hermes_cli/cli_model_switch_mixin.py::_persist_global_switch          -> deleted; _commit_model_switch calls persist_model_selection
  hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
  gateway/slash_commands_model.py::_persist_model_switch_to_config       -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
  tui_gateway/model_switch.py::_persist_model_switch                     -> deleted; _apply_model_switch calls persist_model_selection
  hermes_cli/web_server_config.py::_apply_main_model_assignment          -> apply_model_selection(result) (+ explicit custom api_key)
  hermes_cli/web_server_config.py::_validated_main_model_selection       -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
  hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
  acp_adapter/server.py::_resolve_model_selection                        -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError

Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.

Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.

Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
2026-09-13 05:21:02 -07:00
teknium1 c1e58f4cb2 fix(process-identity): desktop reap, Windows updater and profile liveness use the canonical matchers
Two kill/relaunch predicates decided identity by argv substring, the bug class root AGENTS.md
forbids: hermes_cli/dashboard_procs.py::_is_desktop_local_serve_cmdline (`"serve" not in cmd`,
on the orphan-reap KILL path) and hermes_cli/update_cmd_windows.py::_is_backend_argv
(`" serve" in argv_low`, in the very file that defines _hermes_holder_subcommand). Both now ask
the canonical token classifier; host/port are read as flag values, not substrings.

hermes_cli/profiles.py::_check_gateway_running open-coded rungs 1/3 of
gateway.status.resolve_gateway_liveness and skipped the multiplexer rung; it is now that ladder
scoped to the profile dir (pid probe keeps cleanup_stale=False so a probe for another profile
never unlinks its PID file). The gateway/status.py ladder itself is untouched.

Behavior change: `hermes kanban --preserve-cache --host 127.0.0.1 --port 0` and
`-m dashboard serve`-style argv are no longer classified as serve backends (never killed /
relaunched as one); a named profile served by the live default multiplexer now reads as
running from _check_gateway_running (previously only via the separate
_served_by_running_multiplexer OR at some call sites).
2026-09-13 05:21:02 -07:00
teknium1 e3281fe249 fix(state): gateway /retry re-sends a multi-part carrier's stored bytes, not a "\n"-joined view
`rewind_user_turn` called `retryable_user_text` for validation only and returned
`flatten_message_text(...)` (`"\n".join` of the parts) as `live_text`, which the
gateway used as the re-sent prompt: `[{text:"a"},{text:"b"}]` went back to the
model as `"a\nb"` where origin/main sent `"ab"`. When `require_retryable` the
outcome now carries the lossless `"".join` (wire bytes == stored bytes); the
display flattening stays for /undo prefill.

Review follow-up on #109610.
2026-09-13 05:20:26 -07:00
teknium1 c0be0e0826 refactor(display): one think-tag list; ACP tool titles derive from agent/display previews
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.

acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
2026-09-13 05:20:26 -07:00
teknium1 5aa1a50c57 fix(config): good-config backup only for the active home; fixture-based effective-config contract
- load_user_config_effective wrote backups/config/*.good.* into ANY home it
  read, so doctor and the TUI cwd lookup created backup dirs inside other
  profiles. The copy is now taken only when the path is the active home's
  config (the only home load_config ever backed up).
- The effective-config test re-composed the implementation's own primitives
  (a mirror); it now pins a literal expected dict for a fixture of user file +
  managed overlay + env, and fail_closed asserts yaml.YAMLError.
- Four repointed gateway tests carried duplicate _load_gateway_config setattr
  lines (one silently overriding the other); deduped to the intended dict.
2026-09-13 05:09:06 -07:00
teknium1 b91088d768 refactor(config): one effective-user-config loader replaces 9 hand-rolled raw→overlay→expand pipelines
Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.

hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.

Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
2026-09-13 05:09:06 -07:00
teknium1 6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
teknium1 9401cc1643 test(redact): a2a superset test asserts every credential class; drop dead mask branch
The a2a invariant test skipped any synthesized token the canonical
redactor itself let through, so a boundary/shape regression would have
passed silently. Every class scrubs today, so the escape hatch goes and
the soft ">= 50" count becomes the exact registry size.

The gateway body test still asserted a "[REDACTED]" fallback marker that
no longer exists; redact_for_egress masks via _mask_token ("***"), so
assert that alone. honcho oauth.py's `re` import became unused when its
private pattern list moved to the registry.
2026-09-13 05:07:50 -07:00
teknium1 dd1baee0e4 refactor(secrets): drop scope-aware env shims; runtime_provider and the voice/xai tools read the canonical getters
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.

tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.

Behavior change: none.
2026-09-13 05:07:50 -07:00
teknium1 3a75078c3d test(channel_directory): fault the canonical JSON serializer, not json.dump
atomic_json_write now serializes with json.dumps before touching the file (the
surrogate-escape fix), so a json.dump stub never fired and the disk-full test
silently passed the write. The invariant is unchanged (a failed write keeps the
previous cache); the fault is injected at utils._dump_json.
2026-09-13 05:07:11 -07:00
teknium1 2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Franci Penov f361971eed feat(gateway): fire agent_loop_stopped plugin hook on interrupt
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.

_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.

Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.

Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
2026-09-12 22:19:44 -07:00
Teknium a5522f69c0 feat(slack): route status/title through the Agent Sessions API (slack-sdk 3.44.0)
Slack deprecates the Assistant messaging experience (assistant_view) in
February 2027: assistant.threads.setStatus/setTitle are replaced by
agents.sessions.setStatus/rename. slack-sdk 3.44.0 (Aug 27 2026) ships
the typed methods with drop-in-compatible signatures.

- adapter: capability probe on the AsyncWebClient CLASS (never instance —
  mock auto-attributes lie), cached; status set/clear + thread title route
  through agents.sessions.* when available, legacy otherwise
- pins: slack-sdk 3.43.0 -> 3.44.0 (pyproject messaging+slack extras,
  lazy_deps, uv.lock)
- tests: autouse fixture pins the probe to legacy under the mocked SDK;
  5 new tests cover both routing paths for typing, clear, and title
- docs: slack.md scope table + status-line notes mention both methods
2026-09-12 22:15:11 -07:00
Teknium 11d761eed9 fix(stream): flush residual SSE buffer at EOF and widen resilience to the Gemini native adapter
Widening commit on top of the salvaged #9834: the same SSE parsing loop
class drops a final frame that is not newline-terminated (its bytes sit
in `buffer` at EOF and are discarded), and a clean EOF without [DONE]
was presented as a complete answer. Ported from earendil-works/pi#8997
(pi credited: Qiaochu Hu), which fixed the identical class in pi's
streamProxy.

- gateway/run_turn.py::_run_agent_via_proxy — flush the residual buffer
  after the read loop; surface EOF-without-[DONE] (warn + error result
  when nothing was received); extract _consume_sse_line so line parsing
  and the EOF flush share one code path.
- agent/gemini_native_adapter.py::_iter_sse_events — same residual-buffer
  flush via a shared _parse_sse_line helper.
- Tests: 3 invariants (residual flush x2 sites, EOF-without-DONE error),
  proven red on origin/main.
2026-09-12 21:18:24 -07:00
0xbyt4 c2b9a517e0 fix(gateway): harden proxy mode SSE streaming — 3 resilience bugs
Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.

1. `[DONE]` marker didn't break the outer chunk loop
   ---------------------------------------------------
   The `break` on `[DONE]` only exited the inner line-parse `while`,
   leaving the outer `async for chunk in resp.content.iter_any():` to
   keep reading. If the upstream held the connection open after
   `[DONE]` (buggy proxy, crashed server, network hang), the client
   waited up to sock_read=1800 seconds (30 min) for the next chunk.

   Fix: set a `done` flag when `[DONE]` is seen and check it at the
   top of the outer loop.

2. No TCP connect timeout
   -----------------------
   `ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
   the default `None` (no timeout). An unreachable proxy host (DNS
   fail, firewall, remote down) would hang on TCP connect for the OS
   default (minutes) before surfacing an error to the user.

   Fix: add `sock_connect=30` so connect failures surface within 30s.

3. SSE JSON parse exception handling was too narrow
   -------------------------------------------------
   The inner parse caught only `json.JSONDecodeError`. A response like
   `{"choices": [null]}` parsed successfully, then
   `choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
   object has no attribute 'get'`. That bubbled up to the outer
   `except Exception`, aborting the entire stream — any further chunks
   were lost, and the user saw the accumulated partial response
   without knowing why.

   Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
   dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
   to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
   chunk now skips, the stream keeps parsing.

New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:

- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
  chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
  kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
  chunks and verifies both good chunks are captured, bad ones skipped
2026-09-12 21:18:24 -07:00
Teknium e151d0b345 fix(discord): reject a numeric application ID pasted as the bot token during setup
Port from openclaw/openclaw#140531: users paste the application ID from the
Developer Portal's General Information page instead of the bot token (Bot
page); the gateway then fails at runtime with an opaque 401. A real bot token
is dot-separated base64 and never purely numeric, so the setup wizard now
rejects an all-digit answer with pointed guidance and re-prompts once. A
second consecutive numeric answer is kept (user override), and non-numeric
tokens are saved exactly as before.

Adapted to hermes: the guard lives in the Discord plugin's interactive_setup
(the active setup path per the #9983 review — legacy _setup_standard_platform
is not the dispatch target), mirroring the existing Telegram token-shape
validation in hermes_cli/setup_platforms.py.
2026-09-12 21:03:12 -07:00
teknium1 a52e553af4 test: utf-8 encodings in hot-serve test 2026-09-12 08:49:16 -07:00
teknium1 d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
teknium1 b146cf1d0e fix(email): thread attachment sends on the caller's reply_to
send_document() accepted reply_to but never passed it down, so attachments
always threaded from the cached per-address context (or not at all) even
when the caller named the message to reply to. The plain-text path
(_send_email) already honored it; the attachment path now does too.

The metadata half of #10131 (send_image rejecting metadata=) was already
fixed on main by the adapter parity pass. Diagnosis from #10131 and the
explicit-reply_to-wins shape from PR #10321 (which targeted the
pre-plugin path).

Fixes #10131
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
2026-09-12 08:44:00 -07:00
teknium1 67bd2f6571 fix(gateway/config): root-level platform blocks keep their adapter keys
A root-level `webhook:` block (the pre-`platforms:` spelling, still
supported by platform_section) is never copied into platforms_data, so
the removed _PORT_BRIDGE_KEYS table was its only route to `extra` and
the previous commit regressed it (port fell back to 8644). Bridge every
non-typed key of a root block in _bridged_keys with the same typed-key
exclusion and explicit-extra precedence as PlatformConfig.from_dict.

Found by independent review before merge.
2026-09-12 08:42:58 -07:00
teknium1 2f3ccfa3bd fix(gateway/config): promote every top-level platform key into extra
`platforms.webhook.port: 9100` (and `routes`, `secret`, api_server `key`/
`cors_origins`, any adapter setting) was silently dropped unless nested
under `extra:` — PlatformConfig.from_dict only read a fixed set of typed
fields. Two partial bridges (a per-platform port/host/secret table in the
loader and an api_server-only block) covered a few keys and had to be
extended for every new one.

from_dict now promotes every non-typed top-level key into `extra`, with an
explicit `extra:` value winning on a clash and typed fields never leaking
into `extra` on a to_dict/from_dict roundtrip. Both hand-written bridges
are removed.

Same direction as PRs #10208/#10211/#10453 (rainow's #10206 diagnosis) and
#20506; those targeted the pre-loader layout.

Fixes #10206
2026-09-12 08:42:58 -07:00
teknium1 ec8b4b2d6f fix(gateway): recheck the heartbeat guard after an awaited edit
Review finding: the guard ran before `edit_message` was awaited; a restart
notice sent during that await followed by a failed edit produced a fresh
"Working" fallback bubble after the notice. Recheck before the fallback
send; the notifier ends instead.
2026-09-12 08:36:14 -07:00
nightq 0fc3cbac59 fix: suppress stale "Still working..." heartbeats during gateway restart
The gateway's long-running notification task was sending "Still working...
messages even after a restart was requested, causing confusing UX where
users received a restart warning followed by normal heartbeat messages.

Added a check in _notify_long_running() to skip notifications when
gateway is draining or restart has been requested.

Fixes NousResearch/hermes-agent#10990
2026-09-12 08:36:14 -07:00
teknium1 ac54e5e715 fix(api): forked sessions carry _branched_from so they stay listable
Review finding: with the child now created before the parent is ended,
child.started_at < parent.ended_at, so _BRANCH_CHILD_SQL's timestamp
fallback no longer classifies the fork as a branch child and the default
GET /api/sessions dropped it. Persist the explicit marker the CLI /branch
path already writes; test covers listing + the failed-fork parent survival.
2026-09-12 08:33:08 -07:00
teknium1 df8b548797 test(gateway): run the TCP-witness Windows arm natively via windows_only
test_loop_tick_witness_arms_over_tcp_on_windows swapped the module's `os`
binding for a proxy reporting name="nt" so a Linux interpreter would take
the non-POSIX branch. That is the "Don't fake the host OS" case: the test
only passes if the interpreter believes it is on another OS. Mark it
windows_only so the tests-os lane runs it on windows-latest for real;
scripts/ci/list_os_marked_tests.py picks the file up by marker name and
it is skipped (not errored) on Linux/macOS. The create=True tripwire from
the earlier salvage commit is what makes it runnable there.
2026-09-12 08:30:27 -07:00
Rob Mosher f2ee74d739 test(windows): create=True for the start_unix_server tripwire patch
asyncio.start_unix_server only exists where an AF_UNIX event loop does.
On native Windows the attribute is absent, so patch.object raised
AttributeError while arming the tripwire — the TCP-witness test could
only ever pass on POSIX, the platform it pretends not to be. create=True
arms the forbidden-call tripwire on every platform; mock removes the
created attribute on exit, so no cross-test leakage.

Before: AttributeError on native Windows. After: 4/4 pass on Windows
(verified) — and unchanged on POSIX, where the attribute exists.
2026-09-12 08:30:27 -07:00
teknium1 f5a2c70b12 test(api-server): RequestKey isolation test tolerates aiohttp < 3.14
The final assertion read `aiohttp.web_request.RequestKey` directly, which
AttributeErrors on the very aiohttp versions (< 3.14) the guarded import
exists to support. Compare against getattr(..., None) so the test checks
the module rebinds to whatever the real aiohttp provides.

Follow-up to #109157.
2026-09-12 08:27:53 -07:00
David Metcalfe 3aea61dbff fix(gateway): run cron/housekeeping/MCP cleanup before the failure-exit return
`_start_gateway_shutdown_tail()` returned False on `should_exit_with_failure`
before `cron_stop.set()`, the cooperative thread waits, the planned-stop
watcher stop and MCP shutdown, so a failure exit leaked the cron ticker and
housekeeping daemon threads (and open MCP connections) for embedded/library
callers. The verdict is now resolved after the teardown, matching the
startup-abort path which already shuts MCP down first.

Fixes #12175. Salvaged from #55031 by @DavidMetcalfe, re-applied onto the
extracted shutdown tail with one thread-lifecycle invariant test.
2026-09-12 08:27:32 -07:00
teknium1 2175df4034 fix(feishu): keep source.message_id in step with the batched event id
Text and media batching advance the coalesced MessageEvent.message_id to
the newest message but left source.message_id at the first one, so after
a batch the reply anchor and the event id disagreed. Advance both together
at the two enqueue sites.
2026-09-12 08:26:39 -07:00
teknium1 f9dac388ba test(slack): expect the triggering ts in per-turn thread metadata
The Slack source now carries message_id, so _thread_metadata_for_source
stamps it (run.py's Slack branch already did this whenever it had an anchor).
2026-09-12 08:26:39 -07:00
teknium1 156d3c454c test(feishu): pin source.message_id to the inbound message id
Reply anchors, /sethome's synthetic-thread check and the relay fallback
read source.message_id; the salvaged Feishu fix had no test for it.
2026-09-12 08:26:39 -07:00