Commit Graph

4450 Commits

Author SHA1 Message Date
teknium1 735831f776 fix(mcp): reconcile chore lives in run_profile_reconcile; prune lazy + mid-connect servers
The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.

reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.

test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
2026-09-13 06:29:12 -07:00
Teknium 4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
kshitijk4poor 78b98032c5 fix(telegram): key the deaf-ingress report on dispatcher progress, not update age
Follow-up to the cherry-picked #102383 commit. The check as written was neither
sensitive nor specific:

- It aged the newest received update, so a wedged PTB dispatcher was never
  reported while new updates kept arriving more often than every 300s (probe:
  1 update/250s for an hour -> 0 reports).
- `delivered` counted only MessageEvents reaching the gateway handler, while
  `received`/`dispatched` counted every Update; a single handled callback_query,
  reaction, unauthorized user or unmentioned group message produced a false
  ERROR after 300s of quiet.
- `_record_updates_received` skipped the generation/teardown guard
  `_record_polling_progress` applies, and the counters never reset across
  polling generations, so a late response from a fenced poll or a reconnect
  inflated the backlog.

Now `received` and `dispatched` count the same population (every fetched
update reaches the group-99 catch-all) and the report fires when a backlog
persists with no dispatch progress across two 90s heartbeats, once per stall,
re-armed on progress, reset per generation, at WARNING (diagnostic only;
#71240 owns recovery). The delivered counter and the `note_inbound_delivered`
facade method are dropped; the once-per-adapter "no message handler" error on
BasePlatformAdapter.handle_message stays. `_record_polling_progress` returns
whether the round-trip was accepted so the received stamp reuses its gate.
Tests trimmed to two invariants; every guard proven red by mutation.

Refs #102260
2026-09-13 18:55:15 +05:30
joaomarcos db407dd078 fix(telegram): report a healthy-but-deaf ingress instead of nothing (#102260)
Every Telegram health probe measures the transport. A getUpdates round-trip
that returns 200 proves bytes are moving and nothing else: the stall watchdog
(#92991), the pending-update probe (#42909/#55769), the get_me() heartbeat
(#66377) and the polling-progress instrumentation all stay green while updates
arrive and then die downstream. The adapter then publishes "connected", logs
nothing at all, and is indistinguishable from a bot nobody has messaged.

That is #102260: three weeks of telegram.state "connected" plus "polling
confirmed healthy: getUpdates progressing (generation 1)" with zero inbound
reaching the agent, surviving every restart. Two of the issue's three
hypotheses do not hold on this code — _record_polling_progress fires on every
round-trip (not only at start_polling), and _send_path_degraded is cleared on
the first confirmed round-trip — and the reporter's own observation that fresh
messages are received but not processed places the failure downstream of the
transport, in the one stretch with no instrumentation at all.

Add the missing delivered side of the accounting:

- received: updates Telegram handed the process, read from the getUpdates
  envelope the adapter already parses (an empty result proves the transport,
  not arrival, so only non-empty results count).
- dispatched: updates PTB's dispatcher carried through the whole handler
  chain, stamped in the existing group-99 catch-all before its early returns.
- delivered: inbound events that reached the gateway's message handler,
  stamped in BasePlatformAdapter.handle_message for every platform.

_check_ingress_delivery_gap runs on the existing heartbeat and, when updates
arrived but nothing was delivered for 300s, names the broken hop: received >
dispatched means the dispatcher is not draining, dispatched > delivered means
Hermes is dropping what arrives. Diagnostic only — a received update
legitimately reaches no gateway turn, and reconnecting a healthy transport
cannot repair a dropped update, so this never drives recovery.

Also make the silent discard on the shared funnel speak: handle_message
returned with no log when no message handler was installed, so a mis-wired
adapter discarded 100% of inbound while connected and able to send.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018pT5hFJBRfLj8KMqFhm3qz
2026-09-13 18:55:15 +05:30
teknium1 0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
teknium1 008caa88a2 fix(platforms): extra_or_secret keeps blank-string YAML values where the old readers did
dingtalk _extra_get, mattermost _extra_or_env and slack _extra_or_env_flag/_channel_set fell
through to env only on None, so `allowed_channels: ""` / `free_response_channels: ""` meant
"no whitelist" rather than "use the env CSV". The shared reader treated blank as unset and
silently widened those to the env value. New `blank_is_unset=False` knob restores the old
semantics at those seven call sites; the default (blank = unset) stays for the readers whose
old body was `extra.get(k) or env`.
2026-09-13 05:32:38 -07:00
teknium1 c0d7b05faa fix(whatsapp_cloud): allowlist "*" admits DM intake again
The mixin routes DM intake through _is_dm_allowed, and the Cloud override was a bare
wa_id set lookup with no "*" handling, so allow_from={"*"} (documented for
WHATSAPP_ALLOWED_USERS, inherited by Cloud) went from admitting intake on main to
denying it. Cloud now overrides _entry_matches instead: bare-wa_id membership, then
the shared WhatsApp matcher ("*" + aliases) — one predicate for strict DM auth, DM
intake and groups. whatsapp_cloud joins the parity matrix; a new wildcard row covers
every wildcard host on all three paths (sabotage: revert -> [whatsapp_cloud] fails).
2026-09-13 05:32:38 -07:00
teknium1 7370c82c2d fix(wecom): WECOM_ALLOW_ALL_USERS opens DMs again; mixin refuses hosts with no env prefix
WeComAdapter mixed in OwnAccessPolicyMixin without ALLOW_ALL_ENV_PREFIX, so
_allow_all_env_names() read "_ALLOW_ALL_USERS" and every open-DM WeCom deployment
(setup still writes WECOM_ALLOW_ALL_USERS) silently denied all DMs.

- WeComAdapter.ALLOW_ALL_ENV_PREFIX = "WECOM"
- OwnAccessPolicyMixin.__init_subclass__ raises TypeError on an empty prefix so the
  omission cannot ship again (every other host already sets one).
- Parity test now runs a third matrix row that sets each host's own
  <PREFIX>_ALLOW_ALL_USERS; the GATEWAY-only row is why this slipped through.
  Sabotage: prefix removed -> [platform] row fails even with the guard reverted.
2026-09-13 05:32:38 -07:00
teknium1 de114b3af1 refactor(platforms): one scoped-secret reader and spec-driven enablement/YAML-bridge boilerplate across all adapters
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:

- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
  docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
  `external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
  Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
  keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
  (SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
  bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
  slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
  `_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.

Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:

- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
  irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
  matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
  nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
  `yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
  mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.

Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
  neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
  seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
  env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
  does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
  all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
  "Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
  reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
  as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
  value (was list/tuple/set only) — same env text for every real YAML shape.

Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.

Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
2026-09-13 05:32:38 -07:00
teknium1 9b1990583d refactor(gateway): adapters share helpers.cancel_task / MessageDeduplicator / bounded_put
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).

- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
  task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
  definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
  whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
  wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
  max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
  LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
  refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
  `~/.hermes`, so profiles with HERMES_HOME set find them.

Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.

Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
2026-09-13 05:32:38 -07:00
teknium1 3a179fe524 refactor(gateway): every Ogg/Opus voice transcode routes through base.transcode_to_ogg_opus
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).

`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.

Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
2026-09-13 05:32:38 -07:00
teknium1 8185834bc9 refactor(gateway): one OwnAccessPolicyMixin replaces five copied DM/group intake predicates
Weixin, WeCom, QQBot, WhatsApp (common + cloud) and Yuanbao's AccessPolicy each carried
their own `_open_dm_opted_in` / `_is_dm_allowed` / `_is_dm_intake_allowed` /
`_is_group_allowed`, differing only in the platform prefix of the allow-all env var and in
small drifts. The same "allow-all must be scoped / fail closed" fix has landed on this trio
at least four times; a shared rule means it lands once.

gateway/platforms/access_policy_mixin.py::OwnAccessPolicyMixin owns the predicates,
reads every env name through the scoped `get_scoped_secret` and exposes two hooks:
`_entry_matches` (platform allowlist matching) and `_live_dm_allow_from` (env-seeded
lists re-read live). `ALLOW_ALL_ENV_PREFIX` is the only per-adapter datum.

Retained overrides (behavior genuinely differs):
- whatsapp_cloud `_allow_all_env_names` adds WHATSAPP_CLOUD_ALLOW_ALL_USERS;
  `_is_dm_allowed` keeps its bare-wa_id normalisation.
- whatsapp_common `_entry_matches` -> phone/LID alias matching; `_live_dm_allow_from`.
- wecom `_is_group_allowed(chat_id, sender_id)` adds the per-group sender allowlist on
  top of the shared chat-level rule; `_entry_matches` strips `wecom:user:` prefixes.
- qqbot `_entry_matches` (case-insensitive, `*`).
- yuanbao `AccessPolicy.is_group_allowed`: `open` groups still require the allow-all
  opt-in (no runner-side mention gate), so it wraps the shared rule.

Behavior change: Weixin `_is_dm_intake_allowed` now denies a blank/whitespace principal
(the other four already did; safe variant chosen). Weixin's inline group gate now goes
through `_is_group_allowed` with identical verdicts.
2026-09-13 05:32:38 -07:00
teknium1 81d77c280a refactor(platforms): photon rides the base send-retry loop; slack/discord parse Retry-After with the shared parser
Photon forked BasePlatformAdapter._send_with_retry before three fixes landed there: the server's
retry_after is honoured over exponential backoff, a long server penalty (> 60s) returns a typed
failure instead of sleeping inline (#91969), and an exhausted rate-limited send no longer posts the
failure notice inside the flood penalty. The fork got none of them. Photon's two genuine differences
are now hooks on the base — `_send_retry_is_final(result)` (structured auth/target refusals are
returned as-is, no retry, no plain-text resend) and `_send_plain_fallback(...)` (no Markdown banner,
richlink() bypassed) — and the 42-line fork is deleted.

Slack's `_retry_after_from_exc` and Discord's `_extract_discord_retry_after` parsed the header by
hand and only understood the numeric form; both now call agent.retry_utils.parse_retry_after_seconds
(numeric or HTTP-date, either header casing). Discord keeps its `retry_after` attribute path, the
`X-RateLimit-Reset-After` fallback and the 1s floor.

Behavior change: Photon retries now add up to 1s of jitter to the backoff and honour a server
retry_after; Slack/Discord recognise an HTTP-date Retry-After they previously ignored.
2026-09-13 05:21:39 -07:00
teknium1 73eadd54f5 fix(platforms): standalone senders return redacted error envelopes
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.

`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.

Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
2026-09-13 05:21:39 -07:00
teknium1 ad305bead5 refactor(gateway): send_exec_approval is a base template method; 9 adapters only render buttons
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.

BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.

The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.

Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
  fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
2026-09-13 05:21:39 -07:00
teknium1 aac36fd96f refactor(gateway): text-batch flush lives on BasePlatformAdapter; shield + cancel-race fixes reach all 8 adapters
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.

The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.

`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
2026-09-13 05:21:39 -07:00
teknium1 11576390fe refactor(model): one persist writer for /model across CLI, gateway, TUI, dashboard; ACP + dashboard validate through switch_model
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.

Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.

Sites -> canonical:
  hermes_cli/cli_model_switch_mixin.py::_persist_global_switch          -> deleted; _commit_model_switch calls persist_model_selection
  hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
  gateway/slash_commands_model.py::_persist_model_switch_to_config       -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
  tui_gateway/model_switch.py::_persist_model_switch                     -> deleted; _apply_model_switch calls persist_model_selection
  hermes_cli/web_server_config.py::_apply_main_model_assignment          -> apply_model_selection(result) (+ explicit custom api_key)
  hermes_cli/web_server_config.py::_validated_main_model_selection       -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
  hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
  acp_adapter/server.py::_resolve_model_selection                        -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError

Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.

Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.

Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
2026-09-13 05:21:02 -07:00
teknium1 f537c0b4c3 refactor(status): CLI, gateway and TUI /status render the same field set from hermes_cli/status_report.py
The three /status renderers (hermes_cli/cli_session_mixin.py::_show_session_status,
gateway/slash_commands_status.py::_handle_status_command, tui_gateway/methods_session.py
session.status) each hand-built Session ID / Path / Title / Model (provider) / Created /
Last Activity / Tokens / Agent Running with their own getattr(agent, "model") fallback
chain, their own updated_at/last_updated_at/last_activity_at scan and their own timestamp
format. A fix to one (a new last-activity column, a placeholder change) silently missed the
other two.

hermes_cli/status_report.py::build_status_fields now derives the common facts once and
returns them as structured, display-ready data; status_lines() renders the English
"Label: value" form for the CLI and TUI. The gateway keeps translating through its
existing t("gateway.status.*") catalog keys (no locale change); the CLI keeps reasoning /
approvals / context, the gateway keeps free-tier / context / queue depth / Matrix scope,
the TUI keeps its Project line. tui_gateway/methods_session.py::_status_dt and the
CLI's inline updated_at loop are gone; cli_session_mixin._timestamp_or stays for its
remaining history-timestamp caller.

Behavior change: none intended for populated sessions. Unified edge cases: a
SessionDB row with an unparseable started_at now falls back to now() on the TUI as it
already did on the CLI, and the TUI's fallback on a bad updated_at is the created stamp on
both surfaces.

Test: tests/hermes_cli/test_status_report_contract.py drives the three real renderers with
one session (distinctive model, provider, title, stamps, token count) and asserts each
output carries every common value. Sabotage-verified: builder dropping tokens -> red;
TUI hand-formatting the model line -> red; restored -> green.
2026-09-13 05:21:02 -07:00
teknium1 991b23ad8d refactor(sessions): one session-id minter; QQ update-prompt key from build_session_key
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.

Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".

gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.

Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
2026-09-13 05:21:02 -07:00
teknium1 b5f072e2ce fix(agent): manual /compress "nothing to compress" gate stays gateway-only
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.

Review follow-up on #109610.
2026-09-13 05:20:26 -07:00
teknium1 c0be0e0826 refactor(display): one think-tag list; ACP tool titles derive from agent/display previews
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.

acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
2026-09-13 05:20:26 -07:00
teknium1 7114da3de6 refactor(agent): manual /compress runs through one core with --preview/--aggressive on every surface
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.

`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
2026-09-13 05:20:26 -07:00
teknium1 c3de99bbf0 refactor(state): one carrier-aware user-turn rewind behind CLI /undo, /retry, gateway and TUI
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).

The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
2026-09-13 05:20:26 -07:00
teknium1 164c5ea142 fix(agent): LLM traffic honours CIDR and wildcard NO_PROXY entries like the platform adapters do
Three answers to "is this host in NO_PROXY": process_bootstrap used the stdlib
proxy_bypass_environment (no CIDR, no `*.`), gateway/platforms/base.py had a
full matcher (should_bypass_proxy) and a second suffix-only one
(is_host_excluded_by_no_proxy, used by Slack). Live-verified: with
NO_PROXY=10.0.0.0/8 Telegram bypassed the proxy while the LLM call to a 10.x
endpoint went through it.

The full matcher moves to the leaf module agent/proxy_bypass.py (stdlib only,
importable at early boot); both base.py functions are one-line forwarders and
process_bootstrap._get_proxy_for_base_url uses it (passing host:port so
port-qualified entries match). The six-key proxy env scan is also shared.
2026-09-13 05:09:43 -07:00
teknium1 b91088d768 refactor(config): one effective-user-config loader replaces 9 hand-rolled raw→overlay→expand pipelines
Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.

hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.

Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
2026-09-13 05:09:06 -07:00
teknium1 579dbe0c71 fix(cron): sqlite_util imports are late so a running scheduler survives an on-disk upgrade
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.

- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
  add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
  _initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
  importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
  imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
  from the cached sqlite_util, then import each store). The previous repoint deleted names
  from cron.executions, which notepad/incidents do not import from, so it passed regardless.
  Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
  with "cannot import name 'open_db'".
2026-09-13 05:08:29 -07:00
teknium1 576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1 6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
teknium1 226df89f74 refactor(redact): one secret-pattern source; a2a, gateway chat and monitoring egress scrub through redact_for_egress
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.

Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.

Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).

Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
2026-09-13 05:07:50 -07:00
teknium1 9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1 3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
teknium1 2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
Franci Penov f361971eed feat(gateway): fire agent_loop_stopped plugin hook on interrupt
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.

_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.

Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.

Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
2026-09-12 22:19:44 -07:00
0xbyt4 c2b9a517e0 fix(gateway): harden proxy mode SSE streaming — 3 resilience bugs
Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.

1. `[DONE]` marker didn't break the outer chunk loop
   ---------------------------------------------------
   The `break` on `[DONE]` only exited the inner line-parse `while`,
   leaving the outer `async for chunk in resp.content.iter_any():` to
   keep reading. If the upstream held the connection open after
   `[DONE]` (buggy proxy, crashed server, network hang), the client
   waited up to sock_read=1800 seconds (30 min) for the next chunk.

   Fix: set a `done` flag when `[DONE]` is seen and check it at the
   top of the outer loop.

2. No TCP connect timeout
   -----------------------
   `ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
   the default `None` (no timeout). An unreachable proxy host (DNS
   fail, firewall, remote down) would hang on TCP connect for the OS
   default (minutes) before surfacing an error to the user.

   Fix: add `sock_connect=30` so connect failures surface within 30s.

3. SSE JSON parse exception handling was too narrow
   -------------------------------------------------
   The inner parse caught only `json.JSONDecodeError`. A response like
   `{"choices": [null]}` parsed successfully, then
   `choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
   object has no attribute 'get'`. That bubbled up to the outer
   `except Exception`, aborting the entire stream — any further chunks
   were lost, and the user saw the accumulated partial response
   without knowing why.

   Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
   dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
   to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
   chunk now skips, the stream keeps parsing.

New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:

- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
  chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
  kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
  chunks and verifies both good chunks are captured, bad ones skipped
2026-09-12 21:18:24 -07:00
teknium1 6a66a5d481 fix(desktop,dashboard): served profile's api_server/webhook read connected with their /p/<profile>/ URL; shared-gateway restart asks first
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.

- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
  webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
  (`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
  adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
  the real one, not a config guess. `hermes status` lists those URLs beside the other
  shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
  when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
  (statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
  shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
  Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
  start/stop on a served profile renders as an inline notice instead of a raw error toast.
2026-09-12 12:52:19 -07:00
teknium1 2d121aa322 fix(gateway): hot-serve reaches pooled Desktop backends; deleted profiles leave no stale runtime entries
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
  ?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
  process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
  recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
  them as `stopped`.
2026-09-12 08:49:16 -07:00
teknium1 d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
teknium1 67bd2f6571 fix(gateway/config): root-level platform blocks keep their adapter keys
A root-level `webhook:` block (the pre-`platforms:` spelling, still
supported by platform_section) is never copied into platforms_data, so
the removed _PORT_BRIDGE_KEYS table was its only route to `extra` and
the previous commit regressed it (port fell back to 8644). Bridge every
non-typed key of a root block in _bridged_keys with the same typed-key
exclusion and explicit-extra precedence as PlatformConfig.from_dict.

Found by independent review before merge.
2026-09-12 08:42:58 -07:00
teknium1 2f3ccfa3bd fix(gateway/config): promote every top-level platform key into extra
`platforms.webhook.port: 9100` (and `routes`, `secret`, api_server `key`/
`cors_origins`, any adapter setting) was silently dropped unless nested
under `extra:` — PlatformConfig.from_dict only read a fixed set of typed
fields. Two partial bridges (a per-platform port/host/secret table in the
loader and an api_server-only block) covered a few keys and had to be
extended for every new one.

from_dict now promotes every non-typed top-level key into `extra`, with an
explicit `extra:` value winning on a clash and typed fields never leaking
into `extra` on a to_dict/from_dict roundtrip. Both hand-written bridges
are removed.

Same direction as PRs #10208/#10211/#10453 (rainow's #10206 diagnosis) and
#20506; those targeted the pre-loader layout.

Fixes #10206
2026-09-12 08:42:58 -07:00
teknium1 ec8b4b2d6f fix(gateway): recheck the heartbeat guard after an awaited edit
Review finding: the guard ran before `edit_message` was awaited; a restart
notice sent during that await followed by a failed edit produced a fresh
"Working" fallback bubble after the notice. Recheck before the fallback
send; the notifier ends instead.
2026-09-12 08:36:14 -07:00
nightq 0fc3cbac59 fix: suppress stale "Still working..." heartbeats during gateway restart
The gateway's long-running notification task was sending "Still working...
messages even after a restart was requested, causing confusing UX where
users received a restart warning followed by normal heartbeat messages.

Added a check in _notify_long_running() to skip notifications when
gateway is draining or restart has been requested.

Fixes NousResearch/hermes-agent#10990
2026-09-12 08:36:14 -07:00
teknium1 ac54e5e715 fix(api): forked sessions carry _branched_from so they stay listable
Review finding: with the child now created before the parent is ended,
child.started_at < parent.ended_at, so _BRANCH_CHILD_SQL's timestamp
fallback no longer classifies the fork as a branch child and the default
GET /api/sessions dropped it. Persist the explicit marker the CLI /branch
path already writes; test covers listing + the failed-fork parent survival.
2026-09-12 08:33:08 -07:00
vominh1919 d0093a11f0 fix(sessions): /branch and API fork end the parent only after the child exists
Both paths ended the source session as "branched" before create_session
ran, so a failed create left the user on a session already marked ended
with no branch behind it. Create the child first; the parent is ended
only once the branch is real.

Salvage of #11048 (targeted the pre-split cli.py handler; ported to
hermes_cli/cli_commands_mixin.py and the api_server fork sibling);
authored by @vominh1919.

Refs #11030
2026-09-12 08:33:08 -07:00
David Metcalfe 3aea61dbff fix(gateway): run cron/housekeeping/MCP cleanup before the failure-exit return
`_start_gateway_shutdown_tail()` returned False on `should_exit_with_failure`
before `cron_stop.set()`, the cooperative thread waits, the planned-stop
watcher stop and MCP shutdown, so a failure exit leaked the cron ticker and
housekeeping daemon threads (and open MCP connections) for embedded/library
callers. The verdict is now resolved after the teardown, matching the
startup-abort path which already shuts MCP down first.

Fixes #12175. Salvaged from #55031 by @DavidMetcalfe, re-applied onto the
extracted shutdown tail with one thread-lifecycle invariant test.
2026-09-12 08:27:32 -07:00
bixycler 70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1 6716f22396 fix(gateway): mirror_to_session reports False when the transcript write fails
`_append_to_sqlite` caught and debug-logged its own exceptions, so the outer
handler in `mirror_to_session` never fired and every failed SQLite write was
reported as a successful mirror. Callers (cron in_channel seed, send_message)
had no way to know the transcript was never updated.

Let the write helper raise; the caller already warns and returns False.

Fixes #10130
2026-09-12 08:22:42 -07:00
teknium1 21bc1d3775 fix(api-server): trim RequestKey fallback comment, add import-isolation test
The salvaged comment restated the symptom at length; keep only the WHY.
Drop the `# pragma: no cover` markers (the repo does not gate on coverage).
One invariant test: with aiohttp.web_request lacking RequestKey, `web` stays
bound to the aiohttp module and RequestKey is None (red on origin/main).
2026-09-12 08:04:30 -07:00
hoozge 943e04ba2b fix(gateway): keep aiohttp web import alive when RequestKey is unavailable
On aiohttp < 3.14 the RequestKey import fails and the shared except clause
also resets the already-imported web module to None, so every admission
reply (non-streaming POST /v1/runs) raised AttributeError: 'NoneType'
object has no attribute 'json_response' and surfaced as HTTP 500.

Import the two names in separate try/except blocks; RequestKey already has
None-guards at its use sites.
2026-09-12 08:04:30 -07:00
Teknium b0c383cdf7 fix(desktop): served profiles show running and route lifecycle to the multiplexer from a pooled local backend
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
2026-09-12 06:13:44 -07:00
Teknium d76856cc69 fix(migrate): shared-ingress adapters declare serves_profile_prefix so migration reports them as notices
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
2026-09-12 02:02:20 -07:00