Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.
- _restart_launchd_gateway_after_update() (his extraction, adapted):
plist-exists is the ONLY gate; launchd_restart() owns the
bootout/bootstrap/kickstart ladder for every plist-present state;
every failure path is loud and names the manual recovery command.
The gate-error 'except: pass' (the second silent variant) now counts
the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
supervised PID), composing his fix with the returned-is-not-supervised
guard.
- His regression suite adapted to the (restarted, failed) contract; the
old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
bug.
A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).
_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.
That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.
Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.
A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.
Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.
tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.