resolve_runtime_provider returns the bare billing class 'custom' for every
named providers:/custom_providers: entry; the configured id only survives in
requested_provider. All three fallback resolvers (gateway, TUI/desktop, cron)
persisted runtime['provider'] as the agent identity, so an automatic fallback
labeled the session 'custom' in the UI and billing rows, while a manual
/model switch to the same provider showed the configured name.
New shared helper hermes_cli.fallback_config.effective_runtime_provider()
upgrades the bare class back to the entry's configured identity (ad-hoc
provider: custom entries stay unchanged), applied at all three sites —
same class as the delegate_tool fix.
The ChatGPT Codex models endpoint interprets client_version as a Codex
CLI compatibility version and filters out any model whose
minimal_client_version is newer than the value sent. Hermes hardcoded
client_version=1.0.0 at both catalog request sites, so model visibility
was accidentally coupled to a version scheme Hermes doesn't follow —
future models gated behind a higher minimal version would silently
vanish from the account catalog.
The backend accepts the exact sentinel 0.0.0 as an ungated request
returning the complete account catalog (verified live: 0.0.0 and
current versions return identical model sets today, while omitting the
parameter is HTTP 400 and out-of-sequence values like 0.0.1 return no
models). Both request sites (hermes_cli/codex_models.py and the
context-length probe in agent/model_metadata.py) now share one
CODEX_UNGATED_CLIENT_VERSION constant.
Clean-room port of the observed behavior in zed-industries/zed#62729;
no GPL code translated.
OpenAI documents credit_balance_exhausted / *_spend_limit_exceeded /
organization_usage_limit_exceeded as HTTP 429 responses, so the invariant
must hold through _status_429 (which short-circuits _by_error_code), not
only on the status-less body path.
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
The streaming choke point was covered only by an identity check on the
imported alias (a change-detector that would pass with a no-op wiring).
Replace it with a real chunk-loop run asserting STOP -> stop and
MAX_TOKENS -> length on the assembled response, and collapse the
duplicate transport cases into one parametrized invariant.
Some OpenAI-compatible gateways fronting Gemini backends emit the native
uppercase finish reasons (STOP, MAX_TOKENS) instead of the lowercase
OpenAI contract values. Every downstream comparison in Hermes uses
lowercase literals, so an uppercase reason silently fell through: a
clean STOP completion missed the stop handling and a MAX_TOKENS
truncation never entered the length-recovery path.
Adds normalize_finish_reason() as the single owner in
agent/message_sanitization.py (case fold + alias map: max_tokens->length,
end->stop, function_call->tool_calls) and wires it at both wire-intake
choke points: ChatCompletionsTransport.normalize_response and the
streaming chunk-capture loop in chat_completion_helpers. Non-string and
empty values pass through unchanged so existing 'or "stop"' defaults
and the Poolside int-reason path keep their behavior.
uvx.exe (uv's actual Windows shim) and pipx.exe bypassed the MCP OSV
malware check entirely, and backslash-qualified commands only resolved
when running under ntpath. Basename now splits on both separators;
matching stays exact (npx.cmd / uvx.exe / uvx.cmd / pipx.exe) so
lookalikes like npx.exe or npx.cmd.bak remain fail-open.
Port from openclaw/openclaw#123194: a hostile or misbehaving remote MCP
server could stream an unbounded HTTP catalog/tool-result body that the
MCP SDK buffers and JSON-parses before any of Hermes' post-parse limits
(resource cap, tool-result truncation) run.
New _make_mcp_body_cap_transport wraps the owned httpx AsyncClient's
transport on the Streamable HTTP (mcp >= 1.24) and SSE paths:
- finite HTTP bodies capped at 10 MiB (Content-Length rejected up front,
streamed bodies capped chunk-by-chunk);
- each SSE event capped at 10 MiB, with accounting reset at completed
event boundaries so long-lived streams/keepalives are unlimited;
- violations raise httpx.ReadError naming the byte cap, handled by the
existing transport teardown/reconnect path (#66092).
verify/cert now live on the inner AsyncHTTPTransport (client-level TLS
kwargs are inert once a custom transport is passed); the SSE
httpx_client_factory is always injected so the cap applies with default
TLS too. Legacy mcp < 1.24 path (SDK-internal client, no hook) stays
uncapped — same degradation as strict_redirect_headers.
Portable Agent Plugin packages fold the plugin name into the MCP
registry name three times over (slug, digest, and again as the server
key), so mcp__<server>__<tool> routinely exceeds the 64-char function
name limit OpenAI-compatible providers enforce — while the same server
registered via `hermes mcp add` stays well under it. The oversized name
is never rejected loudly; the tool just becomes unreachable. Clamp
mcp_prefixed_tool_name() to 64 chars with a deterministic, collision-
safe hash suffix, mirroring the existing property-key clamp in
schema_sanitizer.py. Dispatch is unaffected since handlers already
close over the original unprefixed tool name.
Fixes#81331
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.
- HermesTokenStorage records hermes_issuer alongside cached tokens
(stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
(no forced re-login for existing installs — deliberate divergence
from Codex, which requires reauth).
Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
Snapshot-then-replace let a voice transcript or interrupt re-queue that landed
between the two steps be dropped by edit/rm/move/clear. Every mutation is now one
critical section under queue.Queue's own mutex.
A leading management word alone no longer hijacks prompts: "/queue clear the logs"
and "/queue edit the config" enqueue as before; only "/queue clear", "/queue edit N ...",
"/queue rm N", "/queue move A B" (numbers present) manage the queue. Handlers live on
CLILoopsMixin next to the existing _cmd_queue (cli.py is a facade), dispatched via a
verb table, and the CommandDef lists the subcommands for tab completion.
Droid v0.203 (Aug 25 2026) added 'edit queued messages' — a queued steering
message can be pulled back and changed before it is sent. Hermes /queue could
only append blindly: no way to see, fix, drop, or reorder queued prompts.
/queue now supports management subcommands in the CLI:
- /queue — list pending prompts (bare prompt still enqueues)
- /queue list — same
- /queue edit N <p> — replace item N (keeps voice sentinel, #65827)
- /queue rm N — remove item N
- /queue move A B — reorder
- /queue clear — drop everything
- /queue add <p> — force-enqueue prompts starting with a management word
Queue mutations hold queue.Queue's mutex and rebuild unfinished_tasks so
join()/task_done bookkeeping stays consistent. Paste references expand on
enqueue and edit, matching the old inline path.
Reimplementation of PR #18833 by @abhinav11082001-stack (commit was authored
under a fabricated 'Hermes Agent' noreply identity that cannot be carried
into history; engineering credit is theirs), hardened for current main:
voice-sentinel-aware previews/edit, paste-reference expansion, queue
bookkeeping asserts, out-of-range no-op tests, and docs.
A bare substring check let a short alias (/q, /v, /bg, /hb) count as
documented whenever a longer command that starts with the same letters
appears anywhere in the doc, so direction 1 of the contract could pass
while the alias's own command had no row.
Two-direction contract test (tests/website/test_slash_commands_doc_parity.py):
every CommandDef must be documented under its name or an alias, and every
doc table row must resolve to a registered command. Ported from IronClaw's
doc-fact contract tests (nearai/ironclaw#7378), adapted from their clap
--help parser to our COMMAND_REGISTRY single source of truth.
Real drift it caught, fixed here: /loop (alias /proactive) shipped with a
full feature page (user-guide/features/loops.md) and CLI+gateway handlers
but never got a row in the slash-commands reference. Added to both the CLI
Session table and the messaging table, plus the both-surfaces note.
AIAgent.interrupt() fans the stop out to a snapshot of _active_children. A child
that is attached after that snapshot — an orchestrator subagent still building
its fan-out siblings, or one that has not yet hit its next iteration check —
started with no signal and ran to completion as an orphan while its parent had
already reported `interrupted`. Live repro (mid orchestrator, stop delivered
between grandchild A and B builds): grandchild B kept its `sleep 20` alive and
stayed in the subagent registry after the mid returned.
_attach_child now mirrors a pending parent stop onto the newcomer, so the whole
spawn tree dies with the node that was stopped. Tests: unit (late attach gets
the stop, normal attach does not) + the real _build_child_agent path with a
stopped orchestrator.
The Desktop "Approvals: off" toggle persists approvals.mode: off. The shell
guards (check_all_command_guards / execute_code) honour it as a bypass, but
_run_approval_gate, the shared gate that computer_use, plugin approval rules,
SSH-config writes and the dangerous-pattern prompt all route through, only
checked _yolo_active() (process --yolo / session /yolo). So with approvals
off, every destructive computer_use action still prompted.
Regressed when 3e066dfedd moved computer_use onto the shared gate: its old
private gate never consulted mode at all, and the shared gate had never been
given the third bypass source. Gate now mirrors the shell guards:
yolo OR approvals.mode == "off" -> approved. Hardline blocks and deny rules
still run before it.
Swallowing FeatureUnavailable returned a bare False, so a hosted operator
saw the same generic "requirements not met / run hermes setup" line the
original report started with. The registry's _probe already logs a raised
exception with its message, so propagating the error is what puts
"target not writable" / "quarantine 404" in gateway.log.
Hosted/Docker images lock /opt/hermes/.venv, so --install-deps writing
site-packages fails with Permission denied and the adapter never starts.
Route Google Chat through lazy_deps (HERMES_LAZY_INSTALL_TARGET) and bake
the extra into the published image so a configured gateway can connect.
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.
The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.
Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.
Closes#89259. Salvages #90649 by @higgs1729.
Settings → Model → Auxiliary gets a reasoning-effort selector per task next to the
provider/model pick (inherit / Off / level), sent as reasoning_effort on
POST /api/model/set and read back from GET /api/model/auxiliary.
(cherry picked from commit f09d008f10a81f57ed2426f835898c8e8ae595d7, resolved onto main; the backend half lives in
hermes_cli/web_server_config.py since the routers split and lands in the next commit)
_profile_scoped_rpc (mcp.servers.*, insights.get, cron.manage, skills.manage,
mcp.catalog, plugins.manage) bound only HERMES_HOME for params.profile.
config.yaml's ${VAR} expansion (config._env_ref_lookup) and the MCP probe's
header/env interpolation resolve through get_secret, which with no scope
installed reads plain os.environ, i.e. the launch profile's values. The
Desktop MCP setup Test-connection (mcp.servers.test) for a secondary thus
sent the default profile's token (or the literal placeholder) and reported
green against the wrong credential - the JSON-RPC twin of the dashboard
REST gap fixed in #110271 (#109901).
Bind the same home + secret + terminal composition a turn binds
(_session_profile_runtime_scope), hydrating the profile's external secret
sources first. os.environ is never mutated. Drops the now-unused
mcp_rpc_helpers.reset_profile.
Under gateway.multiplex_profiles the routed handler runs inside
_profile_runtime_scope, but the adapter's delivery side
(_process_message_background -> _extract_response_content, weixin's own
send()) extracts and validates the reply's MEDIA: / bare-path attachments
after that scope was reset. Docker translation in platforms/base.py
(_docker_sandbox_dir_candidates via get_active_profile_name,
_parse_docker_volume_mounts via the scope-aware TERMINAL_DOCKER_VOLUMES)
therefore resolved a secondary's /output or /root path against the DEFAULT
profile's sandbox and mounts: dropped as "not found on this host", or a
same-named file from the default's mount delivered instead (#109024).
Add GatewayRunner._media_delivery_scope_for_source (home + terminal policy,
no secret hydration: path validation reads no credentials and runs on the
loop) and enter it from BasePlatformAdapter._media_delivery_scope around
the two extraction sites that run outside the turn scope. The streamed path
(_deliver_media_from_response), the background task and cron delivery
already run inside their profile scope.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Port from QwenLM/qwen-code#11723: attaching the configured zone to croniter's
naive wall-clock result resolves the repeated autumn hour to its earlier
occurrence (fold=0), so a base inside the second occurrence received a
next_run_at up to an hour in the past — the fire path would treat it as due
immediately and loop. Try both folds of each candidate and return the earliest
instant strictly after the base; wall-clock jobs still fire exactly once on
the repeated hour (Vixie cron semantics).
The tests patched cron.jobs.get_timezone while _ensure_aware resolves the process's
own zone, so on unfixed main they went red over that disagreement (05:00/04:00) rather
than over the shipped symptom. Configuring HERMES_TIMEZONE through the real resolution
path instead makes the same two tests fail on main with the actual symptom — 08:00 on
spring-forward day, 10:00 on fall-back day — and pass after the fix.
Drops the UTC-stored last_run_at case: 605ba4adea ("interpret naive timestamps as local
time") already normalizes a stored instant into the configured zone, so that half was
fixed before this branch existed and is not this change's to claim. Two invariant tests,
both red on base.
croniter 6.x ignores tzinfo on its start time and instead uses the start's
UTC *offset* as its working offset. compute_next_run passed a tz-aware
last_run_at straight into croniter(expr, base_time), which produced two bugs:
- a last_run_at stored in UTC (+00:00) shifted the next fire to the cron
hour in UTC rather than local time (09:00 UTC = 05:00 America/Toronto)
- on DST transition days the wall-clock hour drifted one hour off (08:00 on
spring-forward, 10:00 on fall-back)
Render the base as the configured zone's naive wall clock for croniter,
then re-attach the zone to the result, so the wall-clock hour stays correct
every calendar day including DST boundaries. Fall back to the base's own
zone only when no timezone is configured.
Adds regression tests covering a UTC-stored last_run_at, spring-forward,
fall-back, and a full-year walk across both DST transitions.
The empty-value row means 'a child inherits the parent agent's effort' for the
delegation task, not a provider default. Wording adopted from #105431 by
@fangliquanflq, which added the same step for delegation alone.
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
Gate the tombstone+notify on "a live default gateway has recorded a served set"
(recorded_served_profiles() is not None) rather than on the per-profile
_served_by_running_multiplexer probe: a multiplexer serves every dir under
profiles/, the signal is cheap, and the narrower probe falls back to config
derivation the CLI process cannot see. Trim the salvaged tests to two
invariants — ordering (unroute while the old home still exists and a stale
mkdir_under_hermes_home of it is refused; hot-serve after the move; no
tombstone left) and no-signal-without-multiplexer. Rollback on a failed move is
kept and covered by the same code path.
Builds on #109269 (xielevi). Fixes#109267.
The salvaged flag was spelled --profile, which collides with the global
-p/--profile that hermes_cli.main scans BEFORE argparse: `hermes webhook
subscribe x --profile compta` would switch this CLI process to compta's
HERMES_HOME and write the subscription into compta's webhook_subscriptions.json
— a file the default gateway's webhook adapter never reads — while the route
still lacked the profile key. #109020 special-cased the scanner for the webhook
subcommand; naming the flag --route-profile removes the ambiguity without
touching _scan_profile_flag: -p picks the gateway whose subscriptions file is
written, --route-profile picks which /p/<profile>/ prefix may hit the route.
Docs: cli-commands reference row, multi-profile-gateways webhook section, the
route `profile` field. Builds on #109020 (fangliquanflq). Fixes#109016.
_profile_home deliberately raises for an explicit target that no longer exists
(an unavailable target must never fall back to the launch profile). Nothing
translated that exception, so a desktop client still holding a deleted profile
turned every profile-scoped RPC (session.create/resume, config.get/set, ...)
into "ws dispatch crash" + -32603. Raise a typed ProfileUnavailableError
(FileNotFoundError subclass, so method-level contracts are unchanged) and map
it once in handle_request to code 4064 — the code the other profile-scoped
surfaces already use for "profile not found". The display-only
_response_profile_name fall-back from #107831 now catches the typed error.
Folds the salvaged test into the existing target-unavailable file (one
invariant test covering both methods and the display helper). Fixes#107829.
get_profile_dir now raises ValueError for a name that is not a valid profile id
(#108346). The tools/MCP scoped-RPC wrapper resolved the profile inside a broad
try that mapped resolve-time errors to the handler's own fail code (or re-raised
for mcp.servers.*), so a traversal-shaped name produced a different error than a
missing profile. Map it to the same 4064 the other profile-scoped surfaces use,
and drop the salvaged change-detector test that froze the profiles-root path.
Widens #108348 (Adolanium) to its remaining sibling surface.
A WS 'profile' param like '../../foo' normalized to a path component that
escaped the profiles root, letting a connected client bind an arbitrary
existing directory as a profile home (state.db opened there, and session
delete chains into per-id file cleanup under <dir>/sessions/).
get_profile_dir now validates the canonical name against the profile id
regex before joining it under profiles/. The regex only, not the reserved
list, so pre-reserved-list dirs like profiles/hermes keep resolving.
Callers that probe existence (profile_exists, _profile_home, the 4064
resolvers) treat ValueError as 'not found'.
Extends the multi-profile MCP paragraph with the identity rules landed in
this branch: OAuth tokens live per profile so each profile's calls run as
its own account; client_cert/client_key are part of the connection
identity; trust and supports_parallel_tool_calls are the consuming
profile's policy even when it shares another profile's connection.
Wording of the per-profile OAuth account guarantee follows the docs draft
in PR #109574 (its token-file fingerprint code path was not taken).
Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
_parallel_safe_servers was keyed by the raw server name while _servers
moved to (scope, name) connection keys (ceaf622c6d). With profile A's `x`
serial and profile B's same-named `x` opted into parallel calls, B's
discovery pass flipped A's tool to parallel-safe and the batch planner put
two A calls in one parallel segment against a server that never opted in
(review of #108352, finding B).
The opt-in is now recorded under the discovering profile's own key and
is_mcp_tool_parallel_safe() looks it up under the calling profile's key,
so one profile's policy never reaches another's same-named server.
Single-profile processes keep the bare-name key, byte for byte.
Under gateway.multiplex_profiles a profile whose mcp_servers entry has the
same route and credentials as another profile's live connection adopts that
connection instead of opening its own. _same_server_route() compares only
the connection identity, so a `trust: untrusted` profile adopted a
`trust: full` profile's connection; _trust_gate_check() then resolved the
OWNER's connection key, read `full`, and let the untrusted profile run
write-capable tools without the approval prompt its config demands
(review of #108352, finding A; regression from ceaf622c6d, where the
name-keyed ledger let the last registrant's trust win instead).
`trust` is the consuming profile's policy, not a property of the
connection: _server_trust_levels is now keyed by the calling profile's own
key (recorded at its own registration and at adoption, dropped when its
overlay is removed), while readOnlyHint stays under the connection key
because it describes the server's tools. Sharing the connection is still
allowed — only the gate is per profile.
Follow-up to the #109930 salvage (#109901). The probe endpoint was the reported
site, but the same class covers every router path that expands a secondary
profile's `${VAR}` refs while only a home override is installed:
`GET /api/mcp/servers` (a `${VAR}` in `url` expanded from this process's env)
and the `/auth` config read, whose expanded entry is handed to the OAuth worker.
Hoist the PR's inline wrapper into one `_profile_secret_scope` context manager
(mirrors `_run_dashboard_mcp_oauth`'s wrapping) and use it at all three sites.
Policy unchanged: scope miss still falls through to os.environ outside
multiplexing; under multiplexing a miss is a miss, never another profile's value.
Tests: the salvaged probe test now uses monkeypatch.setenv (no raw os.environ
mutation); one invariant test for the list endpoint, red on origin/main.
The /api/mcp/servers/{name}/test endpoint reads config and probes with no
profile secret scope installed, so config.yaml's ${VAR} expansion
(_env_ref_lookup) and the probe's interpolation resolve against the
dashboard process's own os.environ — the default profile's values (or
nothing) on a shared remote dashboard. A secondary profile whose
credential comes only from an external secret source (Bitwarden/
1Password) never resolves and the probe sends the literal placeholder,
so the server answers 400 while a fresh profile-scoped CLI process
works (#109901).
Wrap both the config read and the probe in _config_profile_scope +
hydrate_profile_secret_sources + set_secret_scope so refs resolve
against the requested profile's .env plus its per-home hydrated secret
sources, matching the multiplexed turn path (#84079 semantics).
`write_runtime_status` re-stamps the previous writer's `gateway_state.json` in place and
only `_record_served_profiles` (multiplex on) ever wrote `served_profiles`, so a
multiplexer's list survived into a later non-multiplex run of the same home. Every
`hermes -p X` surface then kept treating X as served by that live default gateway: exit
78 on start/install, "running via the default-profile multiplexer" on status (review of
#108352, finding D, second half). The secondary-profile phase now writes an empty list
when multiplexing is off; an empty list is the authoritative "serves nobody else" the
readers already honour.