Proxy mode forwards platform messages to a remote Hermes API server via
SSE. The streaming loop introduced in 90c98345 had three robustness
gaps that could hang the gateway or truncate responses on imperfect
upstream behaviour.
1. `[DONE]` marker didn't break the outer chunk loop
---------------------------------------------------
The `break` on `[DONE]` only exited the inner line-parse `while`,
leaving the outer `async for chunk in resp.content.iter_any():` to
keep reading. If the upstream held the connection open after
`[DONE]` (buggy proxy, crashed server, network hang), the client
waited up to sock_read=1800 seconds (30 min) for the next chunk.
Fix: set a `done` flag when `[DONE]` is seen and check it at the
top of the outer loop.
2. No TCP connect timeout
-----------------------
`ClientTimeout(total=0, sock_read=1800)` left `sock_connect` at
the default `None` (no timeout). An unreachable proxy host (DNS
fail, firewall, remote down) would hang on TCP connect for the OS
default (minutes) before surfacing an error to the user.
Fix: add `sock_connect=30` so connect failures surface within 30s.
3. SSE JSON parse exception handling was too narrow
-------------------------------------------------
The inner parse caught only `json.JSONDecodeError`. A response like
`{"choices": [null]}` parsed successfully, then
`choices[0].get("delta", {})` raised `AttributeError: 'NoneType'
object has no attribute 'get'`. That bubbled up to the outer
`except Exception`, aborting the entire stream — any further chunks
were lost, and the user saw the accumulated partial response
without knowing why.
Fix: add type guards (`isinstance(choices, list)`, `isinstance(first,
dict)`, `isinstance(delta, dict)`) and extend the caught exceptions
to `(json.JSONDecodeError, TypeError, AttributeError)`. One bad
chunk now skips, the stream keeps parsing.
New tests in `tests/gateway/test_proxy_mode.py::TestStreamingResilience`:
- `test_done_marker_stops_reading_trailing_chunks` — verifies trailing
chunks after `[DONE]` are dropped (not appended to `full_response`)
- `test_client_timeout_sets_sock_connect` — captures the ClientTimeout
kwargs and asserts `sock_connect` is set to a reasonable bound
- `test_malformed_chunk_is_skipped_not_fatal` — streams good/bad/good
chunks and verifies both good chunks are captured, bad ones skipped
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
them as `stopped`.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
A root-level `webhook:` block (the pre-`platforms:` spelling, still
supported by platform_section) is never copied into platforms_data, so
the removed _PORT_BRIDGE_KEYS table was its only route to `extra` and
the previous commit regressed it (port fell back to 8644). Bridge every
non-typed key of a root block in _bridged_keys with the same typed-key
exclusion and explicit-extra precedence as PlatformConfig.from_dict.
Found by independent review before merge.
`platforms.webhook.port: 9100` (and `routes`, `secret`, api_server `key`/
`cors_origins`, any adapter setting) was silently dropped unless nested
under `extra:` — PlatformConfig.from_dict only read a fixed set of typed
fields. Two partial bridges (a per-platform port/host/secret table in the
loader and an api_server-only block) covered a few keys and had to be
extended for every new one.
from_dict now promotes every non-typed top-level key into `extra`, with an
explicit `extra:` value winning on a clash and typed fields never leaking
into `extra` on a to_dict/from_dict roundtrip. Both hand-written bridges
are removed.
Same direction as PRs #10208/#10211/#10453 (rainow's #10206 diagnosis) and
#20506; those targeted the pre-loader layout.
Fixes#10206
Review finding: the guard ran before `edit_message` was awaited; a restart
notice sent during that await followed by a failed edit produced a fresh
"Working" fallback bubble after the notice. Recheck before the fallback
send; the notifier ends instead.
The gateway's long-running notification task was sending "Still working...
messages even after a restart was requested, causing confusing UX where
users received a restart warning followed by normal heartbeat messages.
Added a check in _notify_long_running() to skip notifications when
gateway is draining or restart has been requested.
FixesNousResearch/hermes-agent#10990
Review finding: with the child now created before the parent is ended,
child.started_at < parent.ended_at, so _BRANCH_CHILD_SQL's timestamp
fallback no longer classifies the fork as a branch child and the default
GET /api/sessions dropped it. Persist the explicit marker the CLI /branch
path already writes; test covers listing + the failed-fork parent survival.
Both paths ended the source session as "branched" before create_session
ran, so a failed create left the user on a session already marked ended
with no branch behind it. Create the child first; the parent is ended
only once the branch is real.
Salvage of #11048 (targeted the pre-split cli.py handler; ported to
hermes_cli/cli_commands_mixin.py and the api_server fork sibling);
authored by @vominh1919.
Refs #11030
`_start_gateway_shutdown_tail()` returned False on `should_exit_with_failure`
before `cron_stop.set()`, the cooperative thread waits, the planned-stop
watcher stop and MCP shutdown, so a failure exit leaked the cron ticker and
housekeeping daemon threads (and open MCP connections) for embedded/library
callers. The verdict is now resolved after the teardown, matching the
startup-abort path which already shuts MCP down first.
Fixes#12175. Salvaged from #55031 by @DavidMetcalfe, re-applied onto the
extracted shutdown tail with one thread-lifecycle invariant test.
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.
Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.
Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.
Fixes#9565
`_append_to_sqlite` caught and debug-logged its own exceptions, so the outer
handler in `mirror_to_session` never fired and every failed SQLite write was
reported as a successful mirror. Callers (cron in_channel seed, send_message)
had no way to know the transcript was never updated.
Let the write helper raise; the caller already warns and returns False.
Fixes#10130
The salvaged comment restated the symptom at length; keep only the WHY.
Drop the `# pragma: no cover` markers (the repo does not gate on coverage).
One invariant test: with aiohttp.web_request lacking RequestKey, `web` stays
bound to the aiohttp module and RequestKey is None (red on origin/main).
On aiohttp < 3.14 the RequestKey import fails and the shared except clause
also resets the already-imported web module to None, so every admission
reply (non-streaming POST /v1/runs) raised AttributeError: 'NoneType'
object has no attribute 'json_response' and surfaced as HTTP 500.
Import the two names in separate try/except blocks; RequestKey already has
None-guards at its use sites.
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.
Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
sms, line, teams, bluebubbles, msgraph_webhook, whatsapp_cloud, wecom_callback and
feishu (webhook mode) build their aiohttp app exactly as before and hand it to
bind_listener(): standalone and default-profile behaviour is unchanged (same host,
port, reuse_address, access_log), while a multiplex secondary publishes the app for
/p/<profile>/ forwarding instead of binding. BlueBubbles registers the /p/<profile>/
URL with its server; LINE builds media URLs off the shared listener's prefix when no
LINE_PUBLIC_URL is set; WeCom skips its own port-in-use probe in shared mode.
Under gateway.multiplex_profiles a secondary profile with Twilio / LINE / Teams /
BlueBubbles / Microsoft Graph / WhatsApp Cloud / WeCom-callback / Feishu-webhook
credentials was refused WHOLE at config load (SecondaryPortBindingConfigError): every
one of its platforms was skipped because these adapters bind their own port and the
default profile owns the one listener.
The refusal is gone. gateway/platforms/shared_ingress.py gives port-binding adapters
a shared-listener mode: the runner stamps `_shared_listener_profile` on a secondary's
port-binder, `bind_listener()` publishes the adapter's fully wired aiohttp app instead
of starting a TCPSite, and the default listener (api_server, or the webhook adapter
when there is no api_server) forwards `/p/<profile>/<tail>` to the served profile's
adapter whose router matches `/<tail>`, under that profile's runtime scope. The request
is therefore verified by the NAMED profile's adapter with its own secret and replies
leave through that adapter; the un-prefixed path keeps serving the default byte for
byte; an unknown profile or a profile without an adapter for the path is a 404, never
another profile's bot. api_server and webhook stay MIRRORS (the default's own adapter
answers /p/<profile>/ for them) and are the only port-binders a secondary must not
enable; `SHARED_LISTENER_MIRROR_PLATFORMS` is that set in gateway/config.py.
The adapter records its public callback URL (`ingress_url`) in gateway_state.json
under `<profile>:<platform>` and logs it once at connect, so the operator knows what
to paste into the vendor console.
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
Ramp Router efforts cache + warm/disk flags, xAI and OpenRouter image catalogs,
Hindsight append-capability verdict, memory-provider skill registry, OpenViking
atexit provider, Honcho loopback flow status, Langfuse client (os.environ-only
credentials) and disk-cleanup's protected cron paths held one profile's
credential- or home-derived value process-wide; YuanbaoAdapter._active_instance
was last-connected-wins across profiles.
Keyed by home key / credential fingerprint under an override, credentials read
through the secret scope, warm threads run under copy_context(); unscoped module
slots stay for the single-profile path and the existing monkeypatch tests.
Review findings on the salvage (all reproduced with a real SessionStore):
1. Primary persisted-agent path skipped the boundary. The agent's turn-start
flush already persists the user row stamped with the inbound platform id, so
`has_platform_message_id` saw THIS turn's own row, took the "duplicate" branch
and skipped the whole block — including the new assistant boundary. The
transcript stayed `[..., 'user']`, exactly the open tail #107070 is about.
Fresh sessions hid it a second way: `session_meta` is appended after the
agent-flushed user row, so a naive "newest row" tail read sees `session_meta`.
2. The exception fallback appended the boundary unconditionally; a redelivery of
an already-closed turn produced `['user', 'assistant', 'assistant']`.
3. The exception fallback wrote the user row + boundary before classifying a
400/500-on-long-session as overflow, growing a session that is already too
large (the #1630 no-grow rule the persist path honours).
Fix: `SessionDB.latest_conversation_role()` (newest active row excluding the
`session_meta`/`system` bookkeeping rows the model never sees) behind
`SessionStore.transcript_tail_role()`, which resolves the same route
`load_transcript` reads via the existing `_compression_tip_for_session_id`.
One `_hmwa_close_failed_turn()` appends the boundary iff that tail is an open
user row; both the persist path and the exception fallback call it, so the
user-row dedupe no longer gates the boundary and a redelivery never stacks.
The overflow verdict in `_hmwa_agent_error_reply` is an early return ahead of
every transcript write. The `failed_turn_notice` kwarg and its dead
`or _hmwa_failed_turn_notice(...)` fallback are gone; the notice is derived
where each consumer needs it.
Tests (each red with the production change reverted, green here): boundary
keyed on the durable tail with the user write deduped (agent-flushed row →
closed; redelivery → nothing); fresh-session agent-flushed failed first turn
closed despite `session_meta` (real store); exception-path redelivery adds no
second boundary (real store, every lineage location, contract asserted from
store state); exception-path overflow persists nothing. Live E2E:
`evals/gateway_failure_ownership/probe.py` (real AIAgent + fixture provider)
20/20; the two `failed provider input` turns that previously left an open user
tail now close with the "not processed" row.
The notice was derived twice per failed turn (reply + persisted row) from
the same input with different gates, so the shown text and the stored row
were not guaranteed identical. Classify once, pass failed_turn_notice into
_hmwa_persist_turn_transcript. Inline the one-line boundary-row staticmethod
at its two sites, drop the no-producer "notice already present" guard, and
assert the notice constants in tests instead of substrings.
The boundary row closes the user row it follows. When Telegram redelivers
the same failed message the dedupe branch skips the user row, so writing the
boundary unconditionally stacked two consecutive assistant rows and the
notice would be concatenated twice on replay. Append it only alongside the
user row; one test replays the retry.
The tool-evidence scan sliced messages[history_offset:] unguarded — after
mid-turn compression that slice is empty and the "not processed / resend"
notice would have been emitted over a turn whose tools DID run. Route it
through media_repair._current_turn_messages, which already falls back to the
last user row. Drop the notice=None default nobody used, compute the notice
inside _hmwa_persist_turn_transcript instead of threading it as a 12th
kwarg, and share one boundary-row builder between the persist path and the
exception fallback.
PR #107088 changed "Try again or use /reset to start a fresh session."
to "Use /reset to start a fresh session if needed." That rewording is
unrelated to the boundary-row fix; the PARTIAL notice appended right
after it already carries the "verify before resending" caveat. Keep
main's text so the diff stays scoped to the transcript boundary.
Follow-up to #107088 (fangliquanflq), refs #107070.
PR #107088 routed the "Session too large" reply through
_hmwa_add_failed_turn_notice with the PARTIAL notice ("some actions may
already have run; verify"). Context overflow is a deterministic request
rejection, not an indeterminate-effect failure — #107567 just tightened
that verdict — so the overflow branch returns main's exact text again.
One invariant test: a 400 on a >50-row history yields the /compact
guidance alone, without the partial notice.
Follow-up to #107088 (fangliquanflq), refs #107070.
- _drop_turn_slot takes the post-bump run_generation from _interrupt_running_turn
and forwards it to _release_running_agent_state: _interrupt_and_clear_session
awaits adapter.interrupt_session_activity between bump and release, so a
successor claiming the slot in that window must not have its sentinel/lease
wiped by the displaced /stop tail (sync eviction path forwards too, for the
same guard).
- _drop_turn_slot then sweeps lease_tokens from generations older than current:
a hung evicted turn's finalizer may never run, and each such generation pinned
its token (and _SessionLease) forever. Identity-checked + idempotent release,
so a live successor's token is untouched.
- _claim_one_turn_restore folds the two spellings of the one-shot "earliest
snapshot wins" rule (/moa direct write, /model --once setdefault) into one
helper; /model --once passes its pre-switch snapshot so the earliest-wins
contract is expressed once.
- The `one_turn_restore["run_generation"]` stamp had no reader once settlement
moved to the invalidate chokepoint (the finalizer guards on its own generation
via _is_session_run_current); delete it and the test lines that set it.
- release-slot-then-evict-cached-agent (#44212 rationale) was duplicated in
/stop and eviction; one _drop_turn_slot owns it.
- Best-effort interrupt uses the repo's _log_suppressed seam like run_shutdown.
- turn_lease.rebind resolves the lease via token.lease like release does
(identity, not a session_id lookup); _held_turn_lease hands back the token
map so release/rebind stop re-peeking session state.
- Test trims: vacuous isinstance, registry-internals asserts.
`/model X --once` then `/model Y --once` before any turn replaced the
pending snapshot with one taken while X was live, so slot cleanup restored
X and made the first temporary model permanent (ehz0ah, review on #106966).
setdefault keeps the snapshot from the first command — the user's standing
override — as the restore target. One producer-driven regression test.
_hm_evict_running_agent copied the head of _interrupt_and_clear_session
(peek → sentinel check → request_hard_interrupt → invalidate), minus the
turn-process reaper the stop path spawns — so tool subprocesses of an
evicted turn were never reaped. Extract the sync core
(_interrupt_running_turn) and call it from both; the raising-interrupt
guard now protects /stop as well.
#107013 made the finalizer's one-shot restore (/moa, /model --once)
generation-guarded so a displaced turn cannot clobber its successor — but
every displacing path (/stop, /new, idle or reaped eviction) bumps the
generation before that finalizer runs, so the guard skipped the restore and
the one-shot model stayed in force for every later message (main restored
unconditionally). Settle the snapshot inside
_invalidate_session_run_generation, the chokepoint all of those paths go
through, so the displaced finalizer then finds nothing to restore.
/moa now records its prior override in the same conversation.one_turn_restore
snapshot /model --once uses (via _snapshot_session_model_override) instead of
per-turn event attributes, which removes _restore_moa_one_shot,
_moa_run_generation and the "stamp only when None" plumbing; the finalizer
checks ownership with the existing _is_session_run_current. One test drives
the /stop-mid-turn settlement and the stale-finalizer no-op; the moa restore
test now exercises the shared path.
Every `TurnLeaseToken` is now constructed by `SessionTurnLeaseRegistry.acquire`
with its concrete `_SessionLease`, so `release()` no longer needs the
`getattr(token, "lease", None) or self._leases.get(...)` fallback that #107013
left in place. Make `lease` a required constructor argument and resolve the
lease from the token alone; the mapping lookup could only ever return the same
object (or a stale alias after rotation, which is exactly the case identity
release exists to avoid).
PR #107013 introduced `TurnState.lease_tokens` (generation-keyed) so a
displaced turn's unwind releases only its own transcript lease, but kept
the older `lease_token`/`lease_generation` pair alongside it, leaving two
sources of truth and a fallback branch in `_held_turn_lease`.
Drop the singleton pair: acquire, release, rebind, and the legacy
`_turn_lease_tokens` view all read and write `lease_tokens[run_generation]`.
Behaviour is unchanged; the fallback that reconciled the two fields is gone.
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:
- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
timeout were bare os.getenv → the default profile's values; a job without a
model silently ran on the default's HERMES_MODEL instead of refusing.
cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
kanban worker inherited the launch profile's non-credential .env settings
and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
HERMES_MODEL) — X's worker ran in the default's docker image on the
default's model. tools/environments/local.py::strip_launch_profile_env drops
them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
assignee (toolset probes call get_secret without a scope → swallowed
UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
writing the completed tick into the DEFAULT profile's state.db and leaving
X's row awaiting_response forever; the --until judge ran with the default's
aux credentials.
- background processes: a secondary's processes.json (scope-relative since
adf23550f5) was never read at startup; its processes were not re-adopted
and notify_on_complete notices were lost. Startup recovers every served
home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
drain for the ambient profile (default's mode for everyone; X's `off`
dropped a sibling's `all` event), recovered watchers used the default's
mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
_deliver_platform_notice used the default's GatewayConfig so a secondary's
notice_delivery: private went public.
Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.
Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.
BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
gateway/run.py bridges the LAUNCH profile's sessions.cjk_fts / search_slow_ms
into HERMES_CJK_FTS / HERMES_SEARCH_SLOW_MS at import (and re-bridged them per
turn), and hermes_state_fts / hermes_state_search read os.getenv — so a served
secondary always got the default profile's values.
hermes_state_common.routed_sessions_setting() reads the routed profile's
config.yaml under a HERMES_HOME override and the env bridge when unscoped; both
consumers use it. The per-turn re-bridge is skipped inside a secondary's scope
so it can no longer write the default's slots from a routed turn.
A profile served by the default multiplexer owns no gateway.pid / gateway_state.json,
so every surface that reads per-profile identity files called it stopped while the CLI
status surfaces (hermes -p X status / gateway status / cron status) said "running via the
default-profile multiplexer":
- `/api/status?profile=X` and `/api/messaging/platforms?profile=X` reported
gateway_running=false / state=None / "gateway_stopped" in the same body that listed X
under gateways[].served_profiles. The shared ladder `resolve_gateway_liveness` gains a
fourth rung for a named profile_dir: the live default multiplexer that records X in
served_profiles IS X's gateway (pid = multiplexer pid, runtime = its record, X's
`<X>:<platform>` entries re-keyed to the standalone shape).
- `POST /api/gateway/stop?profile=X` spawned `hermes -p X gateway stop`, which printed
"No gateway running for this profile" (exit 0) into the action log while the UI flipped
to stopped and the multiplexer kept serving X; `/api/gateway/restart?profile=X` spawned
a `-p X gateway restart` that only exits 78. start/stop now answer 409 with the
multiplexer explanation (one helper shared with the existing start refusal) and restart
targets the multiplexer, the process that actually serves X. A `--force`-started
separate gateway for X (own pid file) keeps normal per-profile management.
- CLI `hermes -p X gateway stop` refuses with exit 78 like run/start/install/restart when
X has no gateway of its own, instead of a contradictory exit-0 "not running".
Docs: multi-profile-gateways.md §1 and §5 describe stop + the dashboard behaviour.
Keep the two invariants (cross-profile isolation, same-namespace dedup) and a short
docstring; the composite key already covers route and principal rotation.
Review on this PR (tracked issue #84256) found the profile-only fix left
two more dimensions of the shared OpenAI-compat Idempotency-Key cache
unscoped:
1. No endpoint/route discriminator. _run_idempotent() is shared by both
POST /v1/chat/completions and POST /v1/responses, but _IdempotencyCache
._store keeps the fingerprint only as a value inside the key's slot —
not part of the key itself. A client reusing one Idempotency-Key across
both routes would have the second route's settled response silently
overwrite the first route's storage slot, so a legitimate retry of the
first request would recompute instead of replaying its own cached
result. Each call site now passes a stable logical route discriminator
("chat_completions" / "responses") folded directly into the cache key,
rather than relying on request.path (which would wrongly split /v1/...
and its /p/<profile>/v1/... alias into different namespaces).
2. No authenticated principal. The sibling durable /v1/runs admission API's
_run_idempotency_scope() already composes
sha256(profile \0 expected-api-key-or-sentinel) for its non-room case.
_run_idempotent() now delegates directly to
self._run_idempotency_scope(request) for its principal/profile scope
instead of re-deriving the profile alone, so both endpoints share one
definition of "who is this request for" and a caller whose
API_SERVER_KEY changes mid-cache-lifetime can no longer replay a
previous principal's cached response.
Adds regression tests proving: route isolation (including the exact
chat -> responses -> chat retry scenario from the review), concurrent
cross-profile inflight isolation, principal/key-rotation isolation, and
that same-route/same-profile/same-principal dedupe is unchanged.
Mutation-verified with three independent negative controls: reverting the
whole diff (TypeError on the now-required route kwarg), keeping route as a
parameter but dropping it from the key (only the two route-isolation tests
fail), and keeping the shared-scope call but reverting it to a bare
profile lookup (only the two principal-isolation tests fail).
_run_idempotent() (chat completions / responses) keyed the shared
_idem_cache purely on the client-supplied Idempotency-Key header plus a
body fingerprint, with no profile identity anywhere in the key. Under
gateway.multiplex_profiles, both the native routes and their /p/<profile>/
mirrors share the same process-global cache, so two different profiles'
clients sending the same Idempotency-Key with a matching fingerprint (a
realistic case for automation/SDKs that derive the key deterministically
from request content) silently share a cached response instead of each
profile running its own turn.
The sibling /v1/runs admission API already scopes its own idempotency
namespace this way (_run_idempotency_scope() folds in
_api_request_profile.get() or "default"); this mirrors that exact
expression into _run_idempotent()'s cache key.
Under gateway.multiplex_profiles a served secondary profile's adapter is built and
connected inside _profile_runtime_scope while os.environ still holds the DEFAULT
profile's .env. Credentials and allowlists were already read through the profile
scope (get_scoped_secret / _platform_gate_env), but the non-credential SETTINGS the
adapters read with bare os.getenv were not, so a served profile silently ran with the
default profile's values: webhook listener host/port/URL (SMS, Teams, LINE, Feishu,
BlueBubbles), Signal's connect URL/account gate, mention gating and reactions (Slack,
Matrix, Signal, Feishu, BlueBubbles, Discord), Matrix thread/session/E2EE policy and
message-length limits, Discord backfill/command-sync/attachment caps, Buzz reply mode
and env enablement seed, A2A agent name/port/description/toolsets, and the
api_server model alias.
Every such read now goes through the existing scoped reader (get_scoped_secret, or the
adapter's own scope-aware helper): under a secondary's scope the profile's own .env is
authoritative and a miss yields the default -- never another profile's value; the
default profile and single-profile gateways keep reading os.environ exactly as before.
Buzz and A2A previously short-circuited to "extra only / built-in default" under a
scope, which also dropped the profile's OWN .env; they now read the scope so a served
profile matches its standalone gateway.
The parity harness (temp HERMES_HOME, default + 2 secondaries with distinct values for
every env var each adapter reads, real load_gateway_config + adapter factory in both
topologies) went from 70 raw process-env bypass sites across 14 adapters to only the
HERMES_<PLATFORM>_* perf knobs and the api_server listener vars, which are process-
global by design (agent.secret_scope._GLOBAL_ENV_*).
resolve_proxy_url() read platform_env_var (TELEGRAM_PROXY, DISCORD_PROXY,
MATTERMOST_PROXY, MATRIX_PROXY) via raw os.environ.get() — a shared
chokepoint used by 7+ adapters. Under a secondary multiplex profile,
os.environ holds the default profile's YAML-to-env bridge output, so a
secondary profile with its own (different or absent) proxy config
would silently borrow the default profile's proxy, or vice versa —
and proxy URLs can embed credentials (http://user:pass@host).
Fix: read platform_env_var through gateway.platforms._shared's
get_scoped_secret(), which is scope-aware (profile's own scope first,
falling back to os.environ only when unscoped/default-profile) and
already used by 10+ other adapters for the same class of setting. The
generic HTTPS_PROXY/HTTP_PROXY/ALL_PROXY fallback stays a raw env read
— those are OS/system-level network settings, not a per-profile Hermes
concept.
This closes a gap PR #100448 explicitly deferred ("proxy_url is
deliberately NOT scoped here ... left for a dedicated fix to that
shared helper").
Salvage of #77660 covers quotes of INBOUND media via the bridge's download
cache. The reported case is the other half: a user quotes an image the bot
sent (a cron-delivered chart) and asks "what is this?" — the bridge cache
knows only inbound messages, and Meta's Cloud webhook ``context`` carries
only the quoted wamid, so the agent received a text-only turn and could not
see the image it had itself delivered.
Widen ``gateway/rich_sent_store`` (already the (chat_id, message_id) → text
index Telegram and the Cloud adapter use for quoted text) with
``record_media`` / ``lookup_media``: both WhatsApp adapters index the local
path + MIME of every media send at send time and every inbound media
receive, and on a quoted reply fold the resolved ``(path, mime)`` into the
event's own ``media_urls``/``media_types`` so the existing vision/audio
pipeline handles it like a direct attachment. ``lookup_media`` drops entries
whose file no longer exists. Baileys: the outbound index is the fallback
when the bridge cache misses; the cache-dir guard stays on bridge-supplied
paths only (our own sends are paths we chose). Cloud: link sends (public
URL, no local bytes) are not indexed.
One invariant test per adapter, red on origin/main.