Under gateway.multiplex_profiles, `_start_one_profile_adapters` skipped
Platform.RELAY / Platform.WHATSAPP for secondaries with a bare `continue`,
and the startup "not being served" WARNING only covered platforms the
PRIMARY skipped. Four secondaries on one live box had WHATSAPP_ENABLED=true
and nothing in the log, status file, or `hermes gateway status` said the
channel was dead.
- `_note_unserved_secondary_platform`: one INFO per (profile, platform)
naming the reason (shared process-level ingress owned by the default) and
the remedy (enable it on the default profile, or disable it here), plus a
`<profile>:<platform>` runtime-status stamp (state=disabled,
error_code=multiplex_shared_ingress).
- `_start_secondary_profiles` folds those platforms into the loud WARNING
when NO profile (default included) runs them.
- `hermes gateway status --profile X` prints
`whatsapp: not served under multiplex (shared ingress owned by default)`
from that stamp; /api/status excludes `disabled` entries from the
platforms degraded verdict (informational, not a fault).
- Docs: multi-profile-gateways.md gets the shared-ingress rule.
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
`live_default_gateway_pid()` (hermes_cli/gateway_multiplex_served.py) read only the
pid record, so it returned None for a gateway that is alive but has no gateway.pid.
Consumers of the helper then reported the gateway as down:
- `hermes -p <profile> cron list` printed "Gateway is not running" with "jobs won't
fire automatically" while the multiplexer was firing that profile's jobs
- `hermes -p <profile> status` dropped its "running (via the default-profile
multiplexer)" line
- `named_profile_served_by_running_multiplexer()` returned False for a profile the
live gateway serves
The rest of the liveness surface already handles a missing pid file: the
`runtime_pid_probe` seam of `resolve_gateway_liveness()` exists for "launch-service
gateways with no live PID file" (hermes_cli/profiles.py, hermes_cli/web_routers/),
and `hermes_cli/gateway_migrate._live_gateway_pid()` reads "pid file, then runtime
status". This probe was the one call site that never got either.
Read the pid record first, then the PID in `gateway_state.json` validated against the
process table, matching `_live_gateway_pid()`. A record naming a dead pid still
resolves to None, so a stopped gateway keeps reporting stopped and cron keeps warning.
Related to #99631.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
hermes -p <name> gateway status, hermes gateway status and hermes status (under
Serves:) list the /p/<profile>/<path> URL per inbound-port platform the live
multiplexer serves, read from the <profile>:<platform> ingress_url in the default
home's gateway_state.json (hermes_cli/gateway_multiplex_served.py). The dashboard's
messaging payload carries the same ingress_url and the Channels page renders it.
The dashboard's 409 guard now covers only api_server/webhook (the mirrored pair):
enabling Twilio/LINE/Teams/... on a secondary is allowed because the gateway serves
it.
Under gateway.multiplex_profiles the default gateway serves every profile, yet
four startup/status paths still reasoned from the wrong source:
* A secondary profile's API_SERVER_KEY (which the docs REQUIRE for /p/<profile>/
auth) auto-enabled api_server in that profile's config, so
_load_secondary_profile_config raised SecondaryPortBindingConfigError and the
whole profile was skipped. gateway/config_env.py::_enable_from_env now leaves
`enabled` alone for port-binding platforms while a multiplexer loads a
NON-default profile (home override + multiplex flag, the same signal
gateway.config uses for scoped reads); the credential still lands in extra so
the shared listener can authenticate the prefix. Default profile unchanged.
* "Is this profile served?" was re-derived from the default config.yaml plus
GATEWAY_MULTIPLEX_PROFILES as seen by the CLI process. `hermes -p coder ...`
loads coder's .env, so an env-only opt-in on the default profile was invisible
(guard never fired, status said stopped) and an allowlist edit flipped the
answer before the restart. named_profile_served_by_running_multiplexer now
reads the pid-verified default gateway_state.json served_profiles (written by
_record_served_profiles) first and falls back to config derivation only when
the key is absent. The record helpers live in hermes_cli/gateway_multiplex_served.py.
* The served-profile guard ran only inside `gateway run`. `hermes -p X gateway
start|install|restart` reached the service manager, whose unit then exited 78
forever (systemd parks it while the CLI prints "started"; launchd KeepAlive
respawns every 30 s). The service verbs now run the same guard up front
(exit 78, same message) and accept --force; the Desktop /api/gateway/start
route returns 409 for a served profile instead of spawning a doomed child.
* Status surfaces disagreed: `hermes -p coder status` said stopped, `hermes -p
coder cron status` said "cron jobs will NOT fire" while `cron list` said fine,
and the default `hermes status` never listed served profiles. Both now route
through the probe / the recorded served set. The -p/--profile matcher in
_scan_gateway_pids and gateway.status._command_line_belongs_to_profile compares
the flag token for equality (`-p ops` no longer claims -- or lets `gateway
stop` SIGTERM -- an `-p ops-2` gateway).
Docs: multi-profile-gateways.md now describes the start/install refusal, the
--force flags, the API_SERVER_KEY behaviour and the single default-home
gateway_state.json (the per-profile runtime_status.json claim was wrong).
Fixes#100397
Addresses #89726#97360#71344
(cherry picked from commit d002c1864a7b6a22c53758b16b7b0cc79aea2edf)