From bf51fee548ceaa28d2c299a41c0f65066c04c061 Mon Sep 17 00:00:00 2001 From: Teknium <127238744+teknium1@users.noreply.github.com> Date: Fri, 11 Sep 2026 07:22:11 -0700 Subject: [PATCH] docs(multiplex): make the multiplexed-gateway page match what the code does MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The "one gateway for all profiles" section had drifted from the runtime. Each claim was re-verified at its defining symbol on current main and rewritten to the behaviour users will actually see: - named-profile guard: only `gateway run` refuses (exit 78 / EX_CONFIG, systemd RestartPreventExitStatus; launchd KeepAlive still retries); `start`/`install` do not refuse themselves and the message lands in the service log; `--force` is a `run`-only flag (hermes_cli/gateway.py::_guard_named_profile_under_multiplexer, _cmd_start) - same (platform, token) in two profiles: the duplicate adapter is parked as fatal/duplicate_credential and the gateway keeps running — it was described as a fail-fast startup error (gateway/run_adapters.py::_refuse_duplicate_claim) - status surfaces: one gateway_state.json under the default home with `:` entries + served_profiles; nothing is written under a secondary home (the page claimed a per-profile runtime_status.json), and `hermes status` does not list served profiles — `gateway list`, `-p X gateway status` and /api/status do (gateway/status.py::write_runtime_status, hermes_cli/status.py::_render_gateway, hermes_cli/gateway.py::_cmd_status) - API_SERVER_KEY in a secondary .env auto-enables api_server and trips the port-binding skip; document the `enabled: false` pin (gateway/config_env.py::_api_server / _enable_from_env) - allowlist: it is a start-time snapshot, and the Desktop backend's cron ticker enumerates every local profile regardless of it (hermes_cli/web_server.py::_start_desktop_cron_ticker) - routed-profile cron via the shared bot: only when the profile has no live adapter of its own, and a route carrying guild_id never matches a cron target because delivery matches on chat_id/thread_id only (cron/scheduler_provider.py::tick_adapters_for, cron/scheduler_preflight.py::SharedRouteAdapters.get) - add a "What is isolated per profile" table (credentials, authorization, endpoints, media denylist, MCP child env, outbound egress, session namespace, logs, terminal) describing behaviour, not PR numbers - configuration.md: the ${VAR} scoping paragraph now says where it applies and links to the table --- website/docs/user-guide/configuration.md | 2 +- .../docs/user-guide/multi-profile-gateways.md | 42 +++++++++++++++++-- 2 files changed, 39 insertions(+), 5 deletions(-) diff --git a/website/docs/user-guide/configuration.md b/website/docs/user-guide/configuration.md index 7ca818974f..003e76d1fa 100644 --- a/website/docs/user-guide/configuration.md +++ b/website/docs/user-guide/configuration.md @@ -136,7 +136,7 @@ delegation: Multiple references in a single value work: `url: "${HOST}:${PORT}"`. If a referenced variable is not set, the placeholder is kept verbatim (`${UNDEFINED_VAR}` stays as-is) and a warning is logged. Bare `$VAR` is not expanded. -Under a [multiplexed multi-profile gateway](/user-guide/multi-profile-gateways), references in a profile's `config.yaml` resolve against **that profile's** `.env` (its secret scope), not the shared process environment — a `${MATRIX_ACCESS_TOKEN}` in profile B stays unresolved unless B defines the variable itself. Single-profile runs are unchanged. +Under a [multiplexed multi-profile gateway](/user-guide/multi-profile-gateways), references in a profile's `config.yaml` resolve against **that profile's** `.env` (its secret scope), not the shared process environment — a `${MATRIX_ACCESS_TOKEN}` in profile B stays unresolved (kept verbatim, warning logged) unless B defines the variable itself. This holds wherever B's config is loaded inside the multiplexer: routed gateway turns, B's adapter startup, and B's cron jobs. Single-profile runs are unchanged. See [What is isolated per profile](/user-guide/multi-profile-gateways#what-is-isolated-per-profile) for the full list. Cursor-style SecretRef syntax is also accepted: `${env:VAR_NAME}` resolves exactly like `${VAR_NAME}` (the `env:` prefix is stripped), so MCP or provider snippets copied from Cursor / Claude configs work unchanged in both `config.yaml` and the `mcp_servers` block. Other SecretRef sources (`${file:...}`, `${vault:...}`, `${bitwarden:...}`) are **not** resolved inline — external secret backends inject their values into the environment at startup via the `secrets:` block, so reference them as `${env:NAME}` instead; unknown prefixes warn once and stay verbatim. diff --git a/website/docs/user-guide/multi-profile-gateways.md b/website/docs/user-guide/multi-profile-gateways.md index 14e2d80d2e..41b7d5c94e 100644 --- a/website/docs/user-guide/multi-profile-gateways.md +++ b/website/docs/user-guide/multi-profile-gateways.md @@ -212,9 +212,14 @@ still aborts gateway startup rather than silently dropping the unsafe profile. Polling/connection platforms (Telegram, Discord, Slack, Matrix, Signal, …) work fine multiplexed, but each profile that enables one must supply its **own** bot token — the same token cannot be polled by two profiles at once. If two profiles -configure the same `(platform, token)`, startup fails fast naming both profiles -(see [Token-conflict safety](#token-conflict-safety) — the rule is unchanged, -it's just enforced inside the one process now). +configure the same `(platform, token)`, the gateway logs an error naming both +profiles and parks the **duplicate** adapter (it shows as `fatal / +duplicate_credential` in runtime status) while the first claimant and every +other profile keep running — the gateway itself does not exit. The default +profile's adapters connect first and claim their credentials, so the parked +adapter is always the secondary's (see +[Token-conflict safety](#token-conflict-safety) — the rule is unchanged, it's +just enforced inside the one process now). #### 4. Session keys are namespaced by profile @@ -319,6 +324,27 @@ only for that profile's events. Shell hooks run with the routed profile's `HERMES_HOME`, without the default profile's secrets in their environment, and their stdin payload carries a `profile` field naming the profile that fired them. +#### What is isolated per profile + +A quick reference for what a multiplexed turn resolves from **its own** +profile and never shares with the default or any sibling: + +| Concern | Resolved from | Behaviour when the profile lacks it | +|---|---|---| +| Provider keys, bot tokens, `${VAR}` refs in `config.yaml` | The profile's own `.env` (its secret scope) | Unresolved / no adapter — never the default profile's value | +| Authorization (`GATEWAY_ALLOW_ALL_USERS`, `GATEWAY_ALLOWED_USERS`, per-platform allowlists and allow-all opt-ins) | The owning profile's `.env` and `config.yaml` | Closed — a default-profile opt-in never opens a secondary's bot | +| HTTP endpoints (`/p//api/...`, `/p//webhooks/...`, platform event callbacks) | The named profile's `API_SERVER_KEY`, `profile:`-bound webhook routes, and its own adapter | `401`/`404`; delivery without an adapter is `502`/`503`, never another profile's bot | +| `MEDIA:` attachment denylist | Every home under `profiles/` plus the default home, enumerated at check time | A turn can never attach another profile's `.env`, `auth.json`, `state.db`, sessions or token stores | +| stdio MCP child environment | Safe baseline + the profile's scoped values for secret-source names + the server's own `env:` | A name the profile lacks is absent from the child — no default-profile fallthrough | +| Outbound egress (`send_message`, shutdown/restart/`/update` notices, `/loop` wakeups, `profile:`-bound webhook delivery, `github_comment` tokens) | The profile's own connected adapter and `.env` | Clear failure; never posts through the default profile's bot | +| Session namespace | `agent::…` (default keeps `agent:main:…`) | Two profiles on the same chat never share history | +| Logs | `agent.log` / `errors.log` / `gateway.log` under the profile's own home | — | +| Terminal sandbox settings (`terminal.*`, SSH targets) | The profile's `config.yaml` | Documented default; unparsable config → execution refused | + +What is **shared** by design: the process, its PID/lock and `gateway_state.json` +(default home), the one HTTP listener, and the `profile_routes` table (declared +on the default profile). + ### Serving selected profiles By default, `gateway.multiplex_profiles: true` serves every valid named profile @@ -347,6 +373,11 @@ started as `hermes -p gateway run` always ticks its own profile's cron st as well. A named profile outside the allowlist may still run its own standalone gateway. +One caveat: the served set is a **start-time snapshot**. A profile created or +added to the allowlist while the multiplexer is running is not picked up until +`hermes gateway restart` (profiles deleted at runtime are dropped from cron +ticking automatically). + ### Routing shared-bot chats to profiles (`profile_routes`) Multiplexing selects a profile per **credential** (each profile's own bot @@ -657,7 +688,10 @@ and reboots. Each profile must use unique bot tokens for each platform. If two profiles share a Telegram, Discord, Slack, WhatsApp, or Signal token, the second -gateway refuses to start with an error naming the conflicting profile. +gateway refuses to start with an error naming the conflicting profile. Under +[multiplexing](#alternative-one-gateway-for-all-profiles-multiplexing) the same +rule parks only the duplicate profile's adapter and the shared gateway keeps +running. To audit: