Commit Graph

15 Commits

Author SHA1 Message Date
teknium1 8c82e93132 fix(gateway): migrate --multiplex rolls back or resumes when the default cannot come up; every installed unit counts
`hermes gateway migrate --multiplex` ran its one fallible step LAST (install +
start the default gateway) with nothing around it. On a fleet whose secondary
ran a system unit as root (#110850) that step raised, leaving the flag on, the
secondary's unit removed and no gateway anywhere, and the re-run hit the
"already multiplexing (flag on)" short-circuit over an empty fleet.

- apply_migration(): the default bring-up runs inside a rollback. On failure the
  manifest written before the first destructive step restores the flag and
  reinstalls every recorded per-profile gateway with its recorded User=.
- MigrationPlan.interrupted: flag on + manifest present + no live default
  gateway is a half-applied migration, not "already multiplexed"; the re-run
  resumes from the manifest (target manager and User= read from it, since the
  units themselves are gone) instead of refusing. Flag off + leftover manifest
  refuses to overwrite it and points at --standalone.
- ProfileGateway.services records EVERY installed unit (user and system) and the
  manifest carries them; apply stops/uninstalls all of them and rollback
  reinstalls all of them, so a second owner is never left live beside the
  multiplexer. The unattended hook treats a two-unit profile as an ambiguous
  topology and refuses (review finding on #110205).
- gateway_identity(): an unresolvable User= on a system unit stays None instead
  of borrowing the profile directory's owner; the unattended hook treats the
  unknown principal as a boundary (review finding on #110205).
- auto_migration_opted_out(): reads the effective config (load_config_readonly
  under the default home), so a managed `false` wins over a user `true` and a
  YAML string "false" is an opt-out, not a truthy value (review finding on
  #110205).

Builds on KoNit-K's #110854 (run_as_user threaded through install, preserved
from the removed system unit).
2026-09-14 16:16:06 -07:00
KoNit-K 18a45464f3 fix(gateway): preserve system service user during migration 2026-09-14 16:16:06 -07:00
teknium1 9188e708b3 fix(gateway): served_profiles bind to a verified gateway identity, not bare PID existence
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.

One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
2026-09-13 15:41:01 -07:00
teknium1 0abfd1105c fix(migrate): hermes update refuses to fold cross-user / cross-scope gateways; auto_multiplex_migration opt-out
Reshape the two salvaged commits onto current main (#109954):

- Move the boundary guard out of the gateway_migrate facade into a new sibling
  hermes_cli/gateway_migrate_guards.py as a table of guard functions
  (_AUTO_MIGRATION_GUARDS: service domain, UNIX user, HERMES_HOME tree) plus the
  identity resolver. The facade grows by ~20 lines only (uid/runtime_home on
  ProfileGateway, one seam, the hook wiring).
- Compare uids, not strings: live pid owner via /proc (ps fallback only on
  macOS, where /proc does not exist), else the system unit's User= via
  _read_systemd_user_from_unit (root when absent), else the home directory's
  owner. None means unknown and never blocks.
- The home-tree guard reads the HERMES_HOME the installed unit pins, not the
  directory the plan enumerated: that is where the gateway really runs and is
  exactly the "stale copies under profiles/" shape from the report.
- When the default is detached, a service-managed secondary is a different
  domain for the AUTO path (it must not elect the secondary's manager); the
  explicit command keeps electing it as before.
- The explicit command surfaces the same findings as notices (dry run shows
  them) and is never blocked by them; only the update hook refuses.
- Rename the opt-out key to gateway.auto_multiplex_migration (nested only, no
  top-level alias) and read it before a plan is built, so false prints nothing
  and touches nothing. The explicit command ignores it.
- Tests trimmed to the invariants: one parametrized boundary test that exercises
  the real hook end to end (refuses, touches nothing, dry run shows the notice),
  one "same user / same scope still migrates" control, one opt-out test.
- Docs: boundary table + renamed opt-out section in multi-profile-gateways.md;
  one line in hermes_cli/AGENTS.md.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Athena <athena@olympus.local>
2026-09-13 14:48:09 -07:00
Athena 8ed2fec94c feat(gateway): let an install opt out of the automatic multiplex migration
`hermes update` folds an eligible multi-profile install onto one multiplexed
gateway on its own, and there is currently no way to say no. The only lever,
`gateway.multiplex_profiles: false`, is also the default: `_read_multiplex_flag`
returns `False` for "absent" and for an explicit `false` alike, so an operator
who has already decided to stay on per-profile gateways has no way to record
that decision. The migration runs again on the next update.

Add `gateway.auto_migrate` (bool, default `true`). Read from the default
profile's config, it gates the automatic path only:

- absent or `true`   -> today's behaviour exactly, no change
- `false`            -> `maybe_auto_migrate_after_update()` returns before
                        building a plan; no output, no changes

`hermes gateway migrate --multiplex` is an explicit request and still migrates
regardless of the flag, so it stays the supported way to opt back in.

One early return, one schema entry with the reasoning inline, one invariant
test (opt-out blocks the hook, absent/true do not, explicit command still
applies), one section in the multi-profile gateways guide.
2026-09-13 14:48:09 -07:00
KoNit-K 10654711fa fix(gateway): guard auto migration service boundaries 2026-09-13 14:48:09 -07:00
kshitijk4poor 422bc9bde9 fix(migrate): rollback refuses a malformed manifest before mutating; plan and run agree
Follow-ups from the #109962 review that landed after the merge:

- A hand-edited manifest (`"secondaries": null`, a record without profile/home,
  `"default": []`) crashed with a traceback AFTER the flag had already been
  flipped. It is now refused up front, with the same message from the dry-run
  plan and the real run.
- The dry-run plan said "restore its recorded standalone gateway" for a record
  with neither service nor pid, where the run did nothing; both now say so.
- `--standalone --dry-run` with no manifest exits 1 like the real path and
  prints the same hint.
- One "Rollback incomplete" message; a failed default restart now tells the
  user the multiplexer is still running and how to stop it.
- Drop the dead `prefixes and` guard; an empty profile name can no longer
  produce a bare ":" platform prefix.
2026-09-13 20:06:00 +05:30
kshitijk4poor fadcff9227 fix(migrate): rollback re-run skips a detached secondary that is already running
The manifest is kept after a partial rollback so it can be re-run, but the
detached-gateway branch spawned unconditionally, so every re-run doubled the
secondary (service start/restart is idempotent, Popen is not).

Tests: dedupe the fixture-local profile-name helper to module scope.
2026-09-13 19:49:17 +05:30
kshitijk4poor d8cfb149b1 fix(migrate): clear the multiplexer's served record before starting secondaries; always restart the default last
Rollback started each per-profile gateway while the still-live multiplexer's
gateway_state.json listed it as served, so `hermes -p X gateway run` refused
with exit 78 and RestartPreventExitStatus parked the unit for good — the exact
symptom #109473 reports, one step later. Clear served_profiles (and the
profile's platform entries) first; recorded [] is authoritative for the guard.

A failed secondary previously skipped the default restart with the flag already
off, leaving config saying standalone while the process kept multiplexing. The
default now restarts regardless (it is the last operation either way, so a
cgroup-kill of this process can no longer strand later steps); the manifest is
kept for a re-run only when something failed.

Tests: the fake service layer now runs the real served-by-multiplexer guard at
each secondary start, and a failed-secondary contract test; both red on the
prior stack.
2026-09-13 19:49:17 +05:30
joaomarcos f9e47aa6fe fix(migrate): write the rollback manifest before the first destructive operation
An interruption during the first secondary stop/uninstall left no recovery
metadata on disk. Write the complete manifest first; the dict never changes
afterwards, so the per-secondary and post-flag rewrites were byte-identical
and are dropped.

Based on #109490 by @JoaoMarcos44.
2026-09-13 19:49:17 +05:30
KoNit-K 043286ae68 fix(gateway): make standalone rollback recoverable 2026-09-13 19:49:17 +05:30
teknium1 205645ee42 fix(gateway): register gateway.multiplex_profiles; explicit migrate --multiplex flips it with no standalone secondary
`hermes config set gateway.multiplex_profiles true` warned "not a recognized
config key" although gateway/config.py reads it: the key (and profile_routes)
were never in DEFAULT_CONFIG["gateway"]. Both are registered with their doc
comment; the CLI loaders deep-merge new keys, so no _config_version bump.

`hermes gateway migrate --multiplex` with two or more profiles but no
secondary running its own gateway printed "nothing to migrate" and left the
flag OFF. The explicit command now applies the one remaining step — flag on,
default gateway (re)started, the same rollback manifest (empty secondaries)
for --standalone. `hermes update`'s automatic hook keeps treating that case
as a no-op: it never flips modes on an install where nothing was running.
2026-09-12 18:35:21 -07:00
teknium1 d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
Teknium d76856cc69 fix(migrate): shared-ingress adapters declare serves_profile_prefix so migration reports them as notices
#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; #108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
2026-09-12 02:02:20 -07:00
Teknium bcdb49ac7b feat(gateway): adapters declare serves_profile_prefix for /p/<profile>/ ingress
The multiplexer skips a secondary profile that enables a port-binding
platform, unless the default listener already answers that platform under
/p/<profile>/. Which adapters do is now a class attribute on the adapter
(api_server and webhook today) instead of knowledge scattered in prose, so
the migration preflight can tell "URL changes" from "profile would be
skipped" and stays correct as new HTTP-inbound adapters gain the prefix.
2026-09-12 01:49:28 -07:00