`hermes gateway migrate --multiplex` ran its one fallible step LAST (install +
start the default gateway) with nothing around it. On a fleet whose secondary
ran a system unit as root (#110850) that step raised, leaving the flag on, the
secondary's unit removed and no gateway anywhere, and the re-run hit the
"already multiplexing (flag on)" short-circuit over an empty fleet.
- apply_migration(): the default bring-up runs inside a rollback. On failure the
manifest written before the first destructive step restores the flag and
reinstalls every recorded per-profile gateway with its recorded User=.
- MigrationPlan.interrupted: flag on + manifest present + no live default
gateway is a half-applied migration, not "already multiplexed"; the re-run
resumes from the manifest (target manager and User= read from it, since the
units themselves are gone) instead of refusing. Flag off + leftover manifest
refuses to overwrite it and points at --standalone.
- ProfileGateway.services records EVERY installed unit (user and system) and the
manifest carries them; apply stops/uninstalls all of them and rollback
reinstalls all of them, so a second owner is never left live beside the
multiplexer. The unattended hook treats a two-unit profile as an ambiguous
topology and refuses (review finding on #110205).
- gateway_identity(): an unresolvable User= on a system unit stays None instead
of borrowing the profile directory's owner; the unattended hook treats the
unknown principal as a boundary (review finding on #110205).
- auto_migration_opted_out(): reads the effective config (load_config_readonly
under the default home), so a managed `false` wins over a user `true` and a
YAML string "false" is an opt-out, not a truthy value (review finding on
#110205).
Builds on KoNit-K's #110854 (run_as_user threaded through install, preserved
from the removed system unit).
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
Reshape the two salvaged commits onto current main (#109954):
- Move the boundary guard out of the gateway_migrate facade into a new sibling
hermes_cli/gateway_migrate_guards.py as a table of guard functions
(_AUTO_MIGRATION_GUARDS: service domain, UNIX user, HERMES_HOME tree) plus the
identity resolver. The facade grows by ~20 lines only (uid/runtime_home on
ProfileGateway, one seam, the hook wiring).
- Compare uids, not strings: live pid owner via /proc (ps fallback only on
macOS, where /proc does not exist), else the system unit's User= via
_read_systemd_user_from_unit (root when absent), else the home directory's
owner. None means unknown and never blocks.
- The home-tree guard reads the HERMES_HOME the installed unit pins, not the
directory the plan enumerated: that is where the gateway really runs and is
exactly the "stale copies under profiles/" shape from the report.
- When the default is detached, a service-managed secondary is a different
domain for the AUTO path (it must not elect the secondary's manager); the
explicit command keeps electing it as before.
- The explicit command surfaces the same findings as notices (dry run shows
them) and is never blocked by them; only the update hook refuses.
- Rename the opt-out key to gateway.auto_multiplex_migration (nested only, no
top-level alias) and read it before a plan is built, so false prints nothing
and touches nothing. The explicit command ignores it.
- Tests trimmed to the invariants: one parametrized boundary test that exercises
the real hook end to end (refuses, touches nothing, dry run shows the notice),
one "same user / same scope still migrates" control, one opt-out test.
- Docs: boundary table + renamed opt-out section in multi-profile-gateways.md;
one line in hermes_cli/AGENTS.md.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Athena <athena@olympus.local>
`hermes update` folds an eligible multi-profile install onto one multiplexed
gateway on its own, and there is currently no way to say no. The only lever,
`gateway.multiplex_profiles: false`, is also the default: `_read_multiplex_flag`
returns `False` for "absent" and for an explicit `false` alike, so an operator
who has already decided to stay on per-profile gateways has no way to record
that decision. The migration runs again on the next update.
Add `gateway.auto_migrate` (bool, default `true`). Read from the default
profile's config, it gates the automatic path only:
- absent or `true` -> today's behaviour exactly, no change
- `false` -> `maybe_auto_migrate_after_update()` returns before
building a plan; no output, no changes
`hermes gateway migrate --multiplex` is an explicit request and still migrates
regardless of the flag, so it stays the supported way to opt back in.
One early return, one schema entry with the reasoning inline, one invariant
test (opt-out blocks the hook, absent/true do not, explicit command still
applies), one section in the multi-profile gateways guide.
Follow-ups from the #109962 review that landed after the merge:
- A hand-edited manifest (`"secondaries": null`, a record without profile/home,
`"default": []`) crashed with a traceback AFTER the flag had already been
flipped. It is now refused up front, with the same message from the dry-run
plan and the real run.
- The dry-run plan said "restore its recorded standalone gateway" for a record
with neither service nor pid, where the run did nothing; both now say so.
- `--standalone --dry-run` with no manifest exits 1 like the real path and
prints the same hint.
- One "Rollback incomplete" message; a failed default restart now tells the
user the multiplexer is still running and how to stop it.
- Drop the dead `prefixes and` guard; an empty profile name can no longer
produce a bare ":" platform prefix.
The manifest is kept after a partial rollback so it can be re-run, but the
detached-gateway branch spawned unconditionally, so every re-run doubled the
secondary (service start/restart is idempotent, Popen is not).
Tests: dedupe the fixture-local profile-name helper to module scope.
Rollback started each per-profile gateway while the still-live multiplexer's
gateway_state.json listed it as served, so `hermes -p X gateway run` refused
with exit 78 and RestartPreventExitStatus parked the unit for good — the exact
symptom #109473 reports, one step later. Clear served_profiles (and the
profile's platform entries) first; recorded [] is authoritative for the guard.
A failed secondary previously skipped the default restart with the flag already
off, leaving config saying standalone while the process kept multiplexing. The
default now restarts regardless (it is the last operation either way, so a
cgroup-kill of this process can no longer strand later steps); the manifest is
kept for a re-run only when something failed.
Tests: the fake service layer now runs the real served-by-multiplexer guard at
each secondary start, and a failed-secondary contract test; both red on the
prior stack.
An interruption during the first secondary stop/uninstall left no recovery
metadata on disk. Write the complete manifest first; the dict never changes
afterwards, so the per-secondary and post-flag rewrites were byte-identical
and are dropped.
Based on #109490 by @JoaoMarcos44.
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.
`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.