- PUT /api/messaging/platforms on a pooled `hermes --profile X serve` arrives without
?profile= (Desktop local topology, #109088): resolve the hot-serve target from the
process's own profile so the multiplexer is pinged and the UI skips the restart banner.
- A profile deleted while the reconcile lock was held by its own adapter connect was
recorded back into served_profiles; re-check the live set before recording.
- Drop a deleted profile's `<name>:<platform>` runtime-status entries instead of leaving
them as `stopped`.
- /p/<profile>/line/webhook is verified with the NAMED profile's channel secret under
its runtime scope; another profile's secret is 401 at that URL; the bare path is
untouched; unknown profile / profile without the adapter is 404.
- A shared-listener adapter binds no port and records its /p/<profile>/ ingress_url
in runtime status; LINE media URLs use the shared prefix.
- Runner: a secondary's port-binders are constructed in shared-listener mode instead
of refusing the whole profile; api_server/webhook are skipped as mirrors.
- Dashboard: only the mirrored pair is refused (409) on a secondary.
Preserve the two contributor fixes, slim them to two behavioral invariants, and enter explicitly requested homes even inside a nested scope. Real native remote Desktop changes Disabled to gateway_stopped for default and named profiles; direct API controls preserve explicit disable and empty-profile isolation. Unit A/B and regression suites remain queued under the shared campaign lock.
The desktop app always sends profile=default on GET /api/messaging/platforms.
_is_current_profile() recognized only None/""/"current" as the dashboard's
own profile — NOT the string "default" — so a single-profile install (the
standard `hermes gateway setup` flow: token in .env, no platforms: section in
config.yaml) entered the profile-scoped branch of _config_profile_scope().
That branch derives platform enablement from config.yaml only and never calls
load_gateway_config()'s env-override pass (which enables the platform when the
token is in the environment). Result: a platform connected via .env reported
enabled=false, state="disabled" while it was actually running. The unscoped
GET (no profile param) correctly reported enabled=true, state="connected".
Fix: classify by resolved path, not by string. After _is_current_profile()
fails, _config_profile_scope() now resolves the requested profile dir and
compares it against get_process_hermes_home().resolve() — the same comparison
_is_other_profile() already uses. When they match (profile=default on a
default-home process), yield None (no override), taking the unscoped path that
calls load_gateway_config(). A named-profile process (`-p worker`) has a
different HERMES_HOME, so its profile=default resolves to a different directory
and still scopes correctly — cross-profile secret isolation is preserved.
The scoped branch's config.yaml-only enablement is DELIBERATE:
load_gateway_config()'s env pass reads os.environ and would leak the root
install's tokens into a genuinely different profile's state. This fix only
reclassifies requests that name the process's OWN home; it does not touch the
scoped branch or gateway/config_env.py.
Refs #104614
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.
Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
(DEFAULT_CONFIG smart-approval leaked in) — pinned approval
mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
silently probed live provider catalogs (~2s/test) — stubbed
cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
MagicMock agents — interrupt.side_effect now clears _running_agents
(22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
(24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
with rationale comment — the watchdog-dump assertion needs the loop
blocked well past deadline+grace under parallel load (flaked once in
the 40-worker verification run at a 0.2s margin).
Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
The sidebar strip and the Channels page could contradict each other on
the same page load — "Gateway running" next to "The gateway is not
running." /api/status and /api/messaging/platforms each open-coded their
own liveness ladder: status probed GATEWAY_HEALTH_URL and scoped its
PID/state reads to the requested profile, messaging did neither and used
the uncached raw PID probe.
Three deployments hit the split: a cross-container gateway (no local PID,
only the health probe can see it), a profile-scoped dashboard (messaging
borrowed a DIFFERENT profile's runtime state, reporting a false
"connected" that hides a real outage — #71211), and a launch-service
managed gateway with no PID file.
Adds resolve_gateway_liveness() in gateway/status.py as the single ladder
(cached PID -> HTTP health probe -> runtime-status PID with
expected_home) and routes both endpoints, /api/messaging/platforms/{id}/test,
and the kanban dispatcher-presence probe through it. Probe callables are
injectable so the existing monkeypatch seams keep working, and
GatewayLiveness.probe_error distinguishes "down" from "couldn't tell" so
the kanban warning keeps failing OPEN instead of crying wolf.
Closes#71211.
The #62803 branch predates PR #64636's Telegram token-shape validation
on the messaging platform PUT endpoint; align the new guard test's
fixture with the validated format.
The Channels API (PUT /api/messaging/platforms/{id}) accepted and persisted
enabling a port-binding platform on a secondary profile while
gateway.multiplex_profiles is on — a config the gateway only rejects on its
next start, aborting startup with MultiplexConfigError for every multiplexed
profile.
Validate before any .env/config.yaml write and return 409 for the enable
attempt. Disabling and clearing env stay allowed so an already-invalid
profile can be repaired. The port-binding platform set moves to
gateway/config.py (PORT_BINDING_PLATFORM_VALUES) as the single source of
truth shared by gateway startup validation and the dashboard, so the two
policies cannot drift. Platform config mutations now get a names-only audit
log line.
Fixes#62791
Two halves of the same community report (dashboard Profile Builder):
1. A fresh dashboard/CLI-created profile got no .env file unless cloned,
so it silently inherited API keys and messaging tokens from the shell
environment / root install. create_profile() now seeds a placeholder
.env (0600) for non-clone profiles, matching the SOUL.md seeding.
2. The Channels endpoints (/api/messaging/platforms GET/PUT/test) were
not profile-scoped: they read/wrote the dashboard process's own .env
via load_env()/save_env_value() regardless of the global profile
switcher. They now accept the standard optional profile param (body
beats query on the PUT, matching other scoped writes) and run inside
_profile_scope(). When scoped, the payload no longer falls back to
os.environ or load_gateway_config()'s env-override layer — both carry
the ROOT install's credentials and would misreport them as the
profile's. /api/messaging/platforms added to PROFILE_SCOPED_PREFIXES
so the sidebar switcher scopes the Channels page automatically.