The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.
reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.
test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
Follow-up to the cherry-picked #102383 commit. The check as written was neither
sensitive nor specific:
- It aged the newest received update, so a wedged PTB dispatcher was never
reported while new updates kept arriving more often than every 300s (probe:
1 update/250s for an hour -> 0 reports).
- `delivered` counted only MessageEvents reaching the gateway handler, while
`received`/`dispatched` counted every Update; a single handled callback_query,
reaction, unauthorized user or unmentioned group message produced a false
ERROR after 300s of quiet.
- `_record_updates_received` skipped the generation/teardown guard
`_record_polling_progress` applies, and the counters never reset across
polling generations, so a late response from a fenced poll or a reconnect
inflated the backlog.
Now `received` and `dispatched` count the same population (every fetched
update reaches the group-99 catch-all) and the report fires when a backlog
persists with no dispatch progress across two 90s heartbeats, once per stall,
re-armed on progress, reset per generation, at WARNING (diagnostic only;
#71240 owns recovery). The delivered counter and the `note_inbound_delivered`
facade method are dropped; the once-per-adapter "no message handler" error on
BasePlatformAdapter.handle_message stays. `_record_polling_progress` returns
whether the round-trip was accepted so the received stamp reuses its gate.
Tests trimmed to two invariants; every guard proven red by mutation.
Refs #102260
Every Telegram health probe measures the transport. A getUpdates round-trip
that returns 200 proves bytes are moving and nothing else: the stall watchdog
(#92991), the pending-update probe (#42909/#55769), the get_me() heartbeat
(#66377) and the polling-progress instrumentation all stay green while updates
arrive and then die downstream. The adapter then publishes "connected", logs
nothing at all, and is indistinguishable from a bot nobody has messaged.
That is #102260: three weeks of telegram.state "connected" plus "polling
confirmed healthy: getUpdates progressing (generation 1)" with zero inbound
reaching the agent, surviving every restart. Two of the issue's three
hypotheses do not hold on this code — _record_polling_progress fires on every
round-trip (not only at start_polling), and _send_path_degraded is cleared on
the first confirmed round-trip — and the reporter's own observation that fresh
messages are received but not processed places the failure downstream of the
transport, in the one stretch with no instrumentation at all.
Add the missing delivered side of the accounting:
- received: updates Telegram handed the process, read from the getUpdates
envelope the adapter already parses (an empty result proves the transport,
not arrival, so only non-empty results count).
- dispatched: updates PTB's dispatcher carried through the whole handler
chain, stamped in the existing group-99 catch-all before its early returns.
- delivered: inbound events that reached the gateway's message handler,
stamped in BasePlatformAdapter.handle_message for every platform.
_check_ingress_delivery_gap runs on the existing heartbeat and, when updates
arrived but nothing was delivered for 300s, names the broken hop: received >
dispatched means the dispatcher is not draining, dispatched > delivered means
Hermes is dropping what arrives. Diagnostic only — a received update
legitimately reaches no gateway turn, and reconnecting a healthy transport
cannot repair a dropped update, so this never drives recovery.
Also make the silent discard on the shared funnel speak: handle_message
returned with no log when no message handler was installed, so a mis-wired
adapter discarded 100% of inbound while connected and able to send.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018pT5hFJBRfLj8KMqFhm3qz
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.
Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).
Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
dingtalk _extra_get, mattermost _extra_or_env and slack _extra_or_env_flag/_channel_set fell
through to env only on None, so `allowed_channels: ""` / `free_response_channels: ""` meant
"no whitelist" rather than "use the env CSV". The shared reader treated blank as unset and
silently widened those to the env value. New `blank_is_unset=False` knob restores the old
semantics at those seven call sites; the default (blank = unset) stays for the readers whose
old body was `extra.get(k) or env`.
The mixin routes DM intake through _is_dm_allowed, and the Cloud override was a bare
wa_id set lookup with no "*" handling, so allow_from={"*"} (documented for
WHATSAPP_ALLOWED_USERS, inherited by Cloud) went from admitting intake on main to
denying it. Cloud now overrides _entry_matches instead: bare-wa_id membership, then
the shared WhatsApp matcher ("*" + aliases) — one predicate for strict DM auth, DM
intake and groups. whatsapp_cloud joins the parity matrix; a new wildcard row covers
every wildcard host on all three paths (sabotage: revert -> [whatsapp_cloud] fails).
WeComAdapter mixed in OwnAccessPolicyMixin without ALLOW_ALL_ENV_PREFIX, so
_allow_all_env_names() read "_ALLOW_ALL_USERS" and every open-DM WeCom deployment
(setup still writes WECOM_ALLOW_ALL_USERS) silently denied all DMs.
- WeComAdapter.ALLOW_ALL_ENV_PREFIX = "WECOM"
- OwnAccessPolicyMixin.__init_subclass__ raises TypeError on an empty prefix so the
omission cannot ship again (every other host already sets one).
- Parity test now runs a third matrix row that sets each host's own
<PREFIX>_ALLOW_ALL_USERS; the GATEWAY-only row is why this slipped through.
Sabotage: prefix removed -> [platform] row fails even with the guard reverted.
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).
- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
`~/.hermes`, so profiles with HERMES_HOME set find them.
Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.
Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).
`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.
Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
Weixin, WeCom, QQBot, WhatsApp (common + cloud) and Yuanbao's AccessPolicy each carried
their own `_open_dm_opted_in` / `_is_dm_allowed` / `_is_dm_intake_allowed` /
`_is_group_allowed`, differing only in the platform prefix of the allow-all env var and in
small drifts. The same "allow-all must be scoped / fail closed" fix has landed on this trio
at least four times; a shared rule means it lands once.
gateway/platforms/access_policy_mixin.py::OwnAccessPolicyMixin owns the predicates,
reads every env name through the scoped `get_scoped_secret` and exposes two hooks:
`_entry_matches` (platform allowlist matching) and `_live_dm_allow_from` (env-seeded
lists re-read live). `ALLOW_ALL_ENV_PREFIX` is the only per-adapter datum.
Retained overrides (behavior genuinely differs):
- whatsapp_cloud `_allow_all_env_names` adds WHATSAPP_CLOUD_ALLOW_ALL_USERS;
`_is_dm_allowed` keeps its bare-wa_id normalisation.
- whatsapp_common `_entry_matches` -> phone/LID alias matching; `_live_dm_allow_from`.
- wecom `_is_group_allowed(chat_id, sender_id)` adds the per-group sender allowlist on
top of the shared chat-level rule; `_entry_matches` strips `wecom:user:` prefixes.
- qqbot `_entry_matches` (case-insensitive, `*`).
- yuanbao `AccessPolicy.is_group_allowed`: `open` groups still require the allow-all
opt-in (no runner-side mention gate), so it wraps the shared rule.
Behavior change: Weixin `_is_dm_intake_allowed` now denies a blank/whitespace principal
(the other four already did; safe variant chosen). Weixin's inline group gate now goes
through `_is_group_allowed` with identical verdicts.
Photon forked BasePlatformAdapter._send_with_retry before three fixes landed there: the server's
retry_after is honoured over exponential backoff, a long server penalty (> 60s) returns a typed
failure instead of sleeping inline (#91969), and an exhausted rate-limited send no longer posts the
failure notice inside the flood penalty. The fork got none of them. Photon's two genuine differences
are now hooks on the base — `_send_retry_is_final(result)` (structured auth/target refusals are
returned as-is, no retry, no plain-text resend) and `_send_plain_fallback(...)` (no Markdown banner,
richlink() bypassed) — and the 42-line fork is deleted.
Slack's `_retry_after_from_exc` and Discord's `_extract_discord_retry_after` parsed the header by
hand and only understood the numeric form; both now call agent.retry_utils.parse_retry_after_seconds
(numeric or HTTP-date, either header casing). Discord keeps its `retry_after` attribute path, the
`X-RateLimit-Reset-After` fallback and the 1s floor.
Behavior change: Photon retries now add up to 1s of jitter to the backoff and honour a server
retry_after; Slack/Discord recognise an HTTP-date Retry-After they previously ignored.
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.
`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.
Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.
BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.
The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.
Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.
The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.
`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
The three /status renderers (hermes_cli/cli_session_mixin.py::_show_session_status,
gateway/slash_commands_status.py::_handle_status_command, tui_gateway/methods_session.py
session.status) each hand-built Session ID / Path / Title / Model (provider) / Created /
Last Activity / Tokens / Agent Running with their own getattr(agent, "model") fallback
chain, their own updated_at/last_updated_at/last_activity_at scan and their own timestamp
format. A fix to one (a new last-activity column, a placeholder change) silently missed the
other two.
hermes_cli/status_report.py::build_status_fields now derives the common facts once and
returns them as structured, display-ready data; status_lines() renders the English
"Label: value" form for the CLI and TUI. The gateway keeps translating through its
existing t("gateway.status.*") catalog keys (no locale change); the CLI keeps reasoning /
approvals / context, the gateway keeps free-tier / context / queue depth / Matrix scope,
the TUI keeps its Project line. tui_gateway/methods_session.py::_status_dt and the
CLI's inline updated_at loop are gone; cli_session_mixin._timestamp_or stays for its
remaining history-timestamp caller.
Behavior change: none intended for populated sessions. Unified edge cases: a
SessionDB row with an unparseable started_at now falls back to now() on the TUI as it
already did on the CLI, and the TUI's fallback on a bad updated_at is the created stamp on
both surfaces.
Test: tests/hermes_cli/test_status_report_contract.py drives the three real renderers with
one session (distinctive model, provider, title, stamps, token count) and asserts each
output carries every common value. Sabotage-verified: builder dropping tokens -> red;
TUI hand-formatting the model line -> red; restored -> green.
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.
Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".
gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.
Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.
Review follow-up on #109610.
The CLI stream mixin and the gateway think filter each carried a hand-copied think-tag
tuple guarded by a "must stay in sync" comment; adding a tag meant three edits. The
scrubber (agent/think_scrubber.py) now exports THINK_OPEN_TAGS/THINK_CLOSE_TAGS and both
consumers (and strip_think_blocks' regexes) bind to them.
acp_adapter/tools.py::_TITLE_BUILDERS hand-rolled 25 per-tool titles that
agent/display.build_tool_preview already produces (with redaction). ACP titles are now
"<tool>: <preview>"; no ACP-specific overrides remained necessary.
CLI, gateway, TUI and ACP each re-sequenced the same chain (partial split -> estimate ->
_compress_context(force=True) -> lock-skip detection -> rejoin tail -> summary), and the
flag set differed per surface: TUI treated `--preview` as a focus topic, ACP ignored
arguments entirely. For the one command that legitimately breaks the prompt cache that
divergence is a correctness problem, not a style one.
`agent/conversation_compression_manual.py::compress_now` owns the sequence; surfaces parse
their own argv, install `after_messages`, re-anchor session ids and render. TUI and ACP gain
`--preview`, `--aggressive` refusal and `here [N]` parity.
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).
The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
Three answers to "is this host in NO_PROXY": process_bootstrap used the stdlib
proxy_bypass_environment (no CIDR, no `*.`), gateway/platforms/base.py had a
full matcher (should_bypass_proxy) and a second suffix-only one
(is_host_excluded_by_no_proxy, used by Slack). Live-verified: with
NO_PROXY=10.0.0.0/8 Telegram bypassed the proxy while the LLM call to a 10.x
endpoint went through it.
The full matcher moves to the leaf module agent/proxy_bypass.py (stdlib only,
importable at early boot); both base.py functions are one-line forwarders and
process_bootstrap._get_proxy_for_base_url uses it (passing host:port so
port-qualified entries match). The six-key proxy env scan is also shared.
Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.
hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.
Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.
- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
_initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
from the cached sqlite_util, then import each store). The previous repoint deleted names
from cron.executions, which notepad/incidents do not import from, so it passed regardless.
Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
with "cannot import name 'open_db'".
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.
hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.
Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
`ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
hermes_state_messages._scrub_surrogates (0 callers) deleted.
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
plugins/platforms/a2a/security.py::redact_outbound shipped text to a REMOTE peer
through 8 private regexes (sk-, sk-ant-, ghp_ only, xox[bap] only, AKIA, JWT,
Bearer, email) and never called redact_sensitive_text, so every prefix added to
agent/redact.py (hf_, glpat-, xapp-, npm_, Telegram bot tokens, private keys,
DB URLs, env assignments, auth headers, plugin-registered patterns) was absent
on the A2A path. gateway/run.py::_GATEWAY_SECRET_PATTERNS and
agent/monitoring/redaction.py::_TOKEN_RE/_BEARER_RE were two more parallel
"fallback" lists to maintain.
Now agent/redact.py::redact_for_egress is the one egress scrub:
redact_sensitive_text(force=True) + a bearer sweep for prefix-less opaque
tokens, fail-closed ("[redaction-unavailable]"). Gateway user-facing text,
monitoring export and A2A outbound call it; A2A keeps only its e-mail pass.
Behavior changes: a2a egress now masks the full canonical set; the gateway
chat path returns the fail-closed sentinel instead of a raw string when the
redactor raises; honcho plugin registers hch-at-/hch-rt- with
register_redaction_patterns (masked on every surface; mask shape is the
shared head/tail form instead of "hch-at-[redacted]"); proxy_cli token
display uses mask_secret (4 visible prefix chars instead of 12).
Invariant test: redact_outbound masks a synthesized token for every
registered prefix pattern (fails when reverted to the private list).
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.
utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.
Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.
_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.
Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.
Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.
1. `[DONE]` marker didn't break the outer chunk loop
---------------------------------------------------
The `break` on `[DONE]` only exited the inner line-parse `while`,
leaving the outer `async for chunk in resp.content.iter_any():` to
keep reading. If the upstream held the connection open after
`[DONE]` (buggy proxy, crashed server, network hang), the client
waited up to sock_read=1800 seconds (30 min) for the next chunk.
Fix: set a `done` flag when `[DONE]` is seen and check it at the
top of the outer loop.
2. No TCP connect timeout
-----------------------
`ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
the default `None` (no timeout). An unreachable proxy host (DNS
fail, firewall, remote down) would hang on TCP connect for the OS
default (minutes) before surfacing an error to the user.
Fix: add `sock_connect=30` so connect failures surface within 30s.
3. SSE JSON parse exception handling was too narrow
-------------------------------------------------
The inner parse caught only `json.JSONDecodeError`. A response like
`{"choices": [null]}` parsed successfully, then
`choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
object has no attribute 'get'`. That bubbled up to the outer
`except Exception`, aborting the entire stream — any further chunks
were lost, and the user saw the accumulated partial response
without knowing why.
Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
chunk now skips, the stream keeps parsing.
New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:
- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
chunks and verifies both good chunks are captured, bad ones skipped
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
them as `stopped`.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
A root-level `webhook:` block (the pre-`platforms:` spelling, still
supported by platform_section) is never copied into platforms_data, so
the removed _PORT_BRIDGE_KEYS table was its only route to `extra` and
the previous commit regressed it (port fell back to 8644). Bridge every
non-typed key of a root block in _bridged_keys with the same typed-key
exclusion and explicit-extra precedence as PlatformConfig.from_dict.
Found by independent review before merge.
`platforms.webhook.port: 9100` (and `routes`, `secret`, api_server `key`/
`cors_origins`, any adapter setting) was silently dropped unless nested
under `extra:` — PlatformConfig.from_dict only read a fixed set of typed
fields. Two partial bridges (a per-platform port/host/secret table in the
loader and an api_server-only block) covered a few keys and had to be
extended for every new one.
from_dict now promotes every non-typed top-level key into `extra`, with an
explicit `extra:` value winning on a clash and typed fields never leaking
into `extra` on a to_dict/from_dict roundtrip. Both hand-written bridges
are removed.
Same direction as PRs #10208/#10211/#10453 (rainow's #10206 diagnosis) and
#20506; those targeted the pre-loader layout.
Fixes#10206
Review finding: the guard ran before `edit_message` was awaited; a restart
notice sent during that await followed by a failed edit produced a fresh
"Working" fallback bubble after the notice. Recheck before the fallback
send; the notifier ends instead.
The gateway's long-running notification task was sending "Still working...
messages even after a restart was requested, causing confusing UX where
users received a restart warning followed by normal heartbeat messages.
Added a check in _notify_long_running() to skip notifications when
gateway is draining or restart has been requested.
FixesNousResearch/hermes-agent#10990
Review finding: with the child now created before the parent is ended,
child.started_at < parent.ended_at, so _BRANCH_CHILD_SQL's timestamp
fallback no longer classifies the fork as a branch child and the default
GET /api/sessions dropped it. Persist the explicit marker the CLI /branch
path already writes; test covers listing + the failed-fork parent survival.
Both paths ended the source session as "branched" before create_session
ran, so a failed create left the user on a session already marked ended
with no branch behind it. Create the child first; the parent is ended
only once the branch is real.
Salvage of #11048 (targeted the pre-split cli.py handler; ported to
hermes_cli/cli_commands_mixin.py and the api_server fork sibling);
authored by @vominh1919.
Refs #11030
`_start_gateway_shutdown_tail()` returned False on `should_exit_with_failure`
before `cron_stop.set()`, the cooperative thread waits, the planned-stop
watcher stop and MCP shutdown, so a failure exit leaked the cron ticker and
housekeeping daemon threads (and open MCP connections) for embedded/library
callers. The verdict is now resolved after the teardown, matching the
startup-abort path which already shuts MCP down first.
Fixes#12175. Salvaged from #55031 by @DavidMetcalfe, re-applied onto the
extracted shutdown tail with one thread-lifecycle invariant test.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
`_append_to_sqlite` caught and debug-logged its own exceptions, so the outer
handler in `mirror_to_session` never fired and every failed SQLite write was
reported as a successful mirror. Callers (cron in_channel seed, send_message)
had no way to know the transcript was never updated.
Let the write helper raise; the caller already warns and returns False.
Fixes#10130
The salvaged comment restated the symptom at length; keep only the WHY.
Drop the `# pragma: no cover` markers (the repo does not gate on coverage).
One invariant test: with aiohttp.web_request lacking RequestKey, `web` stays
bound to the aiohttp module and RequestKey is None (red on origin/main).
On aiohttp < 3.14 the RequestKey import fails and the shared except clause
also resets the already-imported web module to None, so every admission
reply (non-streaming POST /v1/runs) raised AttributeError: 'NoneType'
object has no attribute 'json_response' and surfaced as HTTP 500.
Import the two names in separate try/except blocks; RequestKey already has
None-guards at its use sites.
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.
Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).