cf60ebbdfd264c3fc067b9e3230df8752c50d2ec
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1bf93660f1 |
fix(update): verify launchd is supervising the gateway after a restart
On macOS, `hermes update` printed "Update complete!" and exited 0 while the ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848). _restart_macos_launchd_gateways already disagrees with itself about what "restarted" means. Sibling profiles are only appended to restarted_services once _wait_for_launchd_service_pid confirms launchd is running the job on a fresh pid. The invoking profile was appended on "launchd_restart() did not raise" alone. That is a weaker claim than it looks. launchd_restart() returns as soon as the restart has been REQUESTED: the _request_gateway_self_restart branch hands the work to the running gateway and returns immediately, and a plist reload is handed to a detached helper. Both are asynchronous, so a helper that dies before its first bootstrap, or a `launchctl bootstrap` that exits 0 without registering (measured by the reporter on macOS 26.6.1), were both invisible to the caller. The systemd branch of the same phase has never drawn that inference: it polls _wait_for_service_active before recording the unit. Verification is domain-agnostic via a new gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid. The sibling helper needs an explicit domain, and the invoking profile's gate deliberately avoids a domain locate because it fails on macOS-26 hosts whose per-user domains reject service management even though launchd_restart() owns that fallback. The new helper judges by a live supervised pid rather than an exit code (the predicate _launchctl_label_supervising_process already existed; this only adds the wait), and returns True immediately when the detached fallback marker is present, because a gateway running unsupervised there is the designed state and not the silent failure this guards against. A label that restarts but is never supervised now lands in failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes the update exit non-zero instead of reporting success over a gateway that is down. Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with no platform gate, driving the real _restart_macos_launchd_gateways through mocked launchctl outcomes. Reverting the verification to an unconditional append fails 2 of them, including the #88848 regression case. tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new verifier so its 27 existing cases keep asserting on routing rather than on a real launchctl probe; unstubbed, each case would poll the full supervision budget. |
||
|
|
f29ee96dd3 |
fix(update): restart all macOS launchd gateways on hermes update
The macOS branch of the update's fleet-restart step only restarted the invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile> services kept pre-update modules cached in sys.modules and died on their next agent turn (ImportError on new lazy imports, or TypeError/ AttributeError with garbled tracebacks on wider version gaps). The systemd branch already iterates every hermes-gateway* unit; this brings launchd to parity: - _restart_macos_launchd_gateways(): the invoking profile keeps the existing launchd_restart() path; every other gateway of this install is drained via SIGUSR1 (same as systemd siblings), then hard- kickstarted unless KeepAlive already respawned it, then verified on a fresh PID. TimeoutExpired is isolated per label (#68523 parity) and counts toward failed_or_stale_units — including timeouts during liveness discovery, which must not read as "unloaded". - Install-scoped fleet enumeration: launchd_gateway_labels_for_install() derives labels from THIS install's profiles (get_default_hermes_root), not by globbing the shared per-user ~/Library/LaunchAgents — a sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side installs) must never enumerate, let alone restart, another install's fleet. This also keeps the hermetic test suite blind to a dev machine's real gateways. - Domain-explicit sibling handling via _locate_launchd_gateway_service(): liveness, kickstart, and fresh-PID verification all use the domain the service was actually located in (gui/<uid> vs user/<uid> probed per label via `launchctl print`). This addresses the #41403 review defect: the process-wide _launchd_domain() cache resolves the current profile's domain and must never be reused for a sibling. _launchd_domain() itself becomes a thin caching wrapper; behavior unchanged. - _get_service_pids(all_profiles=...): the update path's manual-process sweep excludes every gateway service PID (mirror of the systemd hermes-gateway* pattern) so it cannot mistake a freshly respawned sibling service for a stale manual gateway. Default-scope callers (gateway status, cron checks, stop_profile_gateway's orphan reaper — which kills what it is fed) keep the current-profile-only contract. - _warn_incomplete_gateway_fleet_restart() prints launchctl recovery hints for launchd labels alongside the systemctl ones. Supersedes and completes #41403, addressing its review feedback (per-label domain resolution + mocked regression tests). Co-authored-by: David Neyra <vyr.agent@vyrgs.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |