The early-EOF reaper made _finish_reader return without publishing when
wait() raises, so the session stays tracked for later reconciliation.
That is right for the pipe path (_reconcile_local_exit can still reap via
session.process), but PTY sessions have no session.process: when
ptyprocess.wait raises (waitpid ECHILD after isalive() already reaped the
child) the exitstatus is known, yet poll() reported "running" forever.
Only leave the session tracked when exit_code() is still None; otherwise
record the known status and finish as before.
Review finding: PTY session whose pty.wait raises stays in _running forever (fail-open regression vs main)
Drop the wait-timeout call-shape assertion; the invariants are that the
session records the real exit code and that a failed reap does not publish a
false completion.
Authoring standard 5 wants `# <Skill> Skill`, then When to Use,
Prerequisites and Procedure; the port kept the upstream layout with the
trigger sentence in the intro and no prerequisites section. Body text is
unchanged; the docs page is regenerated for this skill only.
Standard 7 asks for tests/skills/test_<skill>_skill.py: two invariants —
frontmatter/section structure, and generation routed through the native
`image_generate` tool with no residue of the upstream harness.
The per-step leases inside maybe_auto_archive / maybe_auto_prune_and_vacuum
(archive, prune, sweep, vacuum) cover every long step of the construction-time
block, and each renews right before the step it protects, so the extra
report_startup_progress(900) at the top of GatewayRunner._init_session_db
added nothing but a stale phase label ("gateway_startup_state_maintenance"
would outlive the archive step and mask the phase name in the fired record).
Dropped; gateway/run.py is back to origin/main.
Tests: the two contributor tests monkeypatched report_startup_progress in the
module and asserted phase names (change-detectors on the strings). Replaced by
one test that arms a REAL StartupWatchdogHandle and asserts the maintenance
block renews it four times with lease_until in the future — the property the
poller's `lease_until > now` branch actually needs (#111092). Red on
origin/main: lease_count stays at the schema-init lease.
Docs: HERMES_STARTUP_WATCHDOG / HERMES_STARTUP_WATCHDOG_TIMEOUT_S existed only
in the module docstring; add them to website/docs/reference/environment-variables.md
next to the respawn-storm variables (existing env vars only, no new surface).
A detached or service-managed gateway can have a blocked stderr. The watchdog exit escort then hard-exits after ten seconds before the file-based traceback is reached, leaving only the metadata record. Write the durable dump first and pin the ordering with a regression test.
Construction-time maybe_auto_archive / maybe_auto_prune_and_vacuum ran
synchronously with no report_startup_progress lease. A multi-minute
VACUUM of a large state.db accrues near-zero CPU, so the startup
watchdog misread it as a parked deadlock and killed the attempt with
exit 75, live-locking gateway restarts. Renew the lease per long step
(prune, orphan sweep, VACUUM, archive) since leases clamp at 900 s,
plus one lease around the gateway maintenance block.
Fixes#111092
carry_unadmitted_user_message appends the interrupted turn's user row to the
early result's in-memory history only; it is never persisted because that turn
never owned the lease. When the follow-up turn also has to wait for the lease
(the common case: the other process that made the first turn wait is usually
still busy), admit_durable_turn_lease reloads conversation_history from the DB
after admission and replaced the caller's list wholesale, so the tagged row was
dropped from both the model input and state.db. Re-append the tagged rows that
have no _row_id after the reload so the follow-up turn sees and flushes them.
Review finding: waited-reload branch of admit_durable_turn_lease discarded the
_persist_after_admission_interrupt row carried from the aborted turn.
Move the pre-admission carry-forward out of run_conversation (facade) into
agent/turn_facade_lease.py::carry_unadmitted_user_message next to the early
result it repairs, and drop the extra turn-start flush in build_turn_context:
the follow-up turn's normal turn-start persist already writes the marked row
because _db_flush_collect no longer stamps it as durable (live probe: exactly
one A row in state.db after two flushes). Trim to two invariant tests
(carry-forward with metadata; flushed exactly once); the hard-stop negative is
covered by the E2E probe in the PR body.
The dynamic-shell-word rules fired on any `-del*`/`-exec*` substring after a
`find` anywhere in the segment, so quoted predicate arguments the shell never
expands were flagged: `find . -name 'log-del*'`, `find . -name 'pre-exec*.sh'`,
`find src -path '*-exec[0-9]*'`, and `echo find . -{delete,print}`.
Three changes close that:
- `find` must be the command word (_CMDPOS anchored, same as mkfs/rm/dd) and
the dynamic word must start a whitespace-delimited token (`(?<!\S)`), for
both the find rule and the rg/sort/ag/man program-option rule.
- Both rules now scan the quote-masked variant (_QUOTE_MASKED_DANGEROUS_
DESCRIPTIONS, same _mask_quoted_prose used by the positionless hardline
rules) so glob characters inside quotes are data, not expansion.
- _iter_shell_command_starts no longer treats the `{` inside a brace-expansion
word (`-{delete,print}`) as a brace-group opener; it split the word across a
marked start so `echo x; find . -{delete,print}` matched nothing. A brace
group opener is `{` as its own word (after whitespace or a separator).
The quoted-name cases join the inert parametrized test; the separator case
joins the dangerous one. Still two parametrized functions.
Three paths still resolved the raw model/remote-supplied string before the
guard could refuse it, so on Windows the NTLM-leak trigger (resolving the
path) ran anyway: the file-checkpoint helper stats write_file/patch targets
before the tool executes; the ACP file bridge resolves fs/read_text_file and
fs/write_text_file paths before its read/write denylists; and @file:/@folder:
references resolve their target before the reference allow-check. Each now
checks the raw string first and refuses. The GLOBALROOT form now requires
its path separator so a GLOBALROOT-prefixed local name is not misclassified.
The rationale comment names the vector instead of another product's
changelog, and the security docs say the row is enforced on reads as well
as writes, since it sits under the write-guard table.
search_tool resolved its root via _resolve_path_for_task BEFORE the
NT-namespace check saw it (review finding): on Windows the resolve is the
SMB-auth trigger, on POSIX the task-base join hides the prefix from the
resolved-path denylist. The raw-string guard now runs first there too,
and the guard rides the existing top-level agent.file_safety import
instead of two function-local imports.
The tool-layer chokepoint test now covers all four entries and proves
none of them touched Path/_resolve_path_for_task/realpath before the
refusal; the 12 form-matrix tests collapse into one blocked/allowed
invariant over both classifiers.
Claude Code v2.1.234 (Aug 17, 2026) hardened its pre-approval file
accesses to reject Windows NT-namespace (\??\) paths against the NTLM
credential-leak vector. Port the same guard into Hermes file safety:
- agent/file_safety.py: is_nt_namespace_path() / get_nt_namespace_error()
raw-string check (never resolves — resolving IS the leak trigger).
Wired as the first check in get_read_block_error() and the write
denial classifier.
- tools/file_tools.py: raw-string guard at read_file_tool entry and in
_check_sensitive_path (covers write_file_tool + patch_tool), before
the task-base join can anchor the prefix under a POSIX base dir.
- Blocks \??\, \\.\, \\?\UNC\, \\?\GLOBALROOT. Extended-length
local drive paths (\\?\C:\...) and plain UNC shares stay allowed.
- tests/agent/test_nt_namespace_guard.py: 10 blocked forms, 11 allowed
forms, no-resolve proof, tool-layer chokepoint coverage.
- docs: protected-paths table in user-guide/security.md
The fix-pass made _pause_windows_gateways_for_update read s.profile from
every discovered service gateway; two quarantine tests build services with
SimpleNamespace stubs that lacked the field the real WindowsGatewayService
always has.
The per-profile cold-start probe took its running set from the socket-paused
ordinary gateways only. A profile whose gateway is alive under an SCM service
is skipped by the socket pause, so it looked "not running"; with an empty
current-PID list its live start attestation read as dead and the profile
landed in cold_start_profiles. On resume the service was restarted AND a
second, unsupervised gateway was spawned for the same profile.
Build the running set from every profile that had ANY live gateway at
discovery time: paused profiles, profile-mapped processes, and service
gateways' profiles.
Review finding: service-supervised running profile was cold-started beside
its restarted SCM service (double gateway).
_spawn_detached now forwards home= to _build_gateway_argv so the post-update
cold-start can launch another profile's gateway; the two windows_only spawn
tests stubbed _build_gateway_argv with a zero-arg lambda and raised TypeError
on the Windows runner.
- Run `_cold_start_attested_profiles` AFTER the paused profiles are relaunched,
so a sibling that fails to cold-start can never keep the profiles that were
running from coming back.
- A profile that stays down raises like the active-profile cold-start does;
the merged outcome then marks the update incomplete instead of printing
success with a gateway still offline.
- Trim the salvaged tests to the two invariants (dead-attested default beside
a live beta is cold-started under its own home; every dead-attested profile
is cold-started when nothing runs, active first). The "no attestation →
token unchanged" case is the pre-existing behaviour already covered by
test_pause_skips_cold_start_plan_when_desktop_owns_lifecycle.
Review follow-ups on the per-profile obligation:
- Order: the active profile's cold-start guard is fleet-wide (any live
gateway ⇒ done), so a sibling spawned first would have left the active
profile down. Per-profile spawns now run after it.
- A token is built even when the active plan owes nothing (clean exit,
autostart not installed), so a dead-attested sibling still rides on it.
- Readiness for a per-profile spawn is probed in THAT home's identity files
(`_live_gateway_pids(home=)` → `get_running_pid(home/"gateway.pid")`), not
the fleet: a still-running sibling no longer vouches for a dead spawn. The
wait/confirm helpers share one probe function.
- The new PID is attested in the profile home (`_write_start_attestation(...,
home=)`) so a death after the CLI exits reaches the next update, and an
already-live profile is not spawned twice.
`_pause_windows_gateways_for_update` built the attested cold-start plan only when the
ALL-profile running PID list was empty, so a default gateway that died after a ✓ beside a
still-running `beta` never received a cold-start obligation and stayed down after the
update (#110959, fifth review thread on #110020).
- gateway_windows: thread `home` through `_start_attestation_path` → `_read_start_attestation`
→ `_attested_pid_exited_cleanly`/`_attested_dead` → `attested_death_generation(pids, home=)`
and `_consume_start_attestation(gen, home=)`; `_spawn_detached(home=)` builds the argv/env
for that profile home (`--profile` derived by `_launcher_settings`). Defaults unchanged.
- update_cmd_windows: `_record_attested_cold_start_profiles` evaluates every
`profiles_to_serve(multiplex=True)` profile that is not running (Desktop-owned installs only,
active profile left to the existing plan so nothing is spawned twice) and records
`token["cold_start_profiles"] = {name: generation}` — a sibling key, so `profiles[name]`
stays an int for relaunch/verify. `_cold_start_attested_profiles` runs on resume before the
ordinary relaunch, spawns under the profile home, waits for readiness, consumes exactly that
generation; one profile's failure never aborts the others.
- tests: default dead-attested + beta running → token carries the obligation, resume spawns
under the default home and consumes its marker while beta is relaunched; no attested
profile → token and resume unchanged.
Why: the rebase conflict resolution in tests/tools/test_model_tools.py
deleted the unrelated TestBridgeDispatch class (3 tests from 73163e3);
it is restored verbatim from main with TestBrowserRetrievalHints after it.
The fix had five tests for one invariant: the OPENAI_MODEL_EXECUTION_GUIDANCE
check duplicated test_phantom_tool_references, and the two static-schema
"toolset-neutral" checks are now folded into test_silent_without_web_tools,
which runs _apply_dynamic_schemas over the real browser_navigate/browser_cdp
schemas so the rendered descriptions are what is asserted.
execution_guidance_text() no longer takes valid_tool_names: the guidance is
toolset-neutral, so the parameter was ignored; the single caller in
agent/system_prompt.py and its test are updated.
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (3733e4aff5) matched
nothing and were dead; the function now returns the neutral text for
every toolset and its phantom-tool test asserts "no web tool named"
instead of the removed sentence. model_tools ports the PR's hint layer
into main's _DYNAMIC_SCHEMA_REWRITERS table (browser_navigate +
browser_cdp) rather than a second pass after it.
Tests: the two browser_cdp registry tests were re-added by the PR but
main pruned them in 39975613b13b4; replaced with one schema-neutrality
invariant. Exact-wording assertions ("lightweight retrieval tool",
"appropriate permitted retrieval/search tool") were change detectors and
are dropped. tools-reference.md row updated to the new schema text.
rename_profile moves the profile directory while this process may still
hold the cached per-profile mcp-stderr.log handle (left behind by a
completed probe or a running server). On Windows a directory containing
an open file cannot be renamed, the same WinError class delete_profile
now avoids. On other platforms the stale handle stayed cached under the
old home key, so a later probe on the renamed profile opened a second
handle and a new profile re-created under the old name wrote its MCP
stderr into the renamed profile's log. Release the scoped handle next to
the multiplexer unroute, mirroring delete_profile.
Review finding: rename_profile missed the sibling surface of the
delete_profile handle release.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
_resolve_hermes_bin_for_desktop_entry returned the resolver's None
outright when neither argv[0] nor PATH yielded a launcher (a cold
relaunch under `python -m` with a stripped PATH), skipping the
known-wrapper probe its own docstring describes. The persisted
`Exec=` then flipped to the bare module form, while a DE- or
terminal-launched context renders the wrapper form.
Every flip rewrites ~/.local/share/applications/hermes.desktop on the
next launch. gnome-shell 50.x removes a ShellApp from id_to_app on any
.desktop content change without checking its state; if the app was
still STARTING, the last strong reference drops and a later GC-triggered
dispose trips shell-app.c's `state == STOPPED` assertion — taking the
whole Wayland session down (observed crash, 50.4-1.fc44).
Probe the installer's known wrapper locations on the None path too, so
the entry converges on the durable wrapper wherever one exists and only
falls back to the module form when none does. A missing entry being
(re)created can still change the file regardless — this removes the
gratuitous rewrites, not the rescan trigger itself.
The "named once" guarantee only held when open() itself failed. In the reported
case open() succeeds and write/seek/flush raise EIO, so every record went
handleError -> stream=None -> next emit reopens -> _open() reset
_unavailable_reported -> the path line was printed once per record (25 for 25).
Reset the flag in emit() only after a record actually reached the stream; that
is the moment the destination has recovered.
Adds the open-succeeds/I-O-fails case to the test file: red on the previous
head, green here.
Review finding: EIO on write after a successful reopen printed the path per record.
bounded_probe_run collapses spawn failure and timeout into one None, so a venv
interpreter that exists but cannot be executed (PermissionError, ENOEXEC, fork
failure) was reported as "timed out before reporting import health" -- a bogus
warning on the git path and sys.exit(1) on the ZIP path -- and the
`except OSError: return {}` branch beneath it was dead. Let the Popen error
propagate (opt-in flag, default unchanged for the other probes) so "we could not
run our own probe" says nothing about the checkout and no longer blocks the
update, while a spawned child that hangs is still a verdict.
Restores the non-fatal test the earlier rewrite dropped, now driving the real
bounded_probe_run/Popen: red on the previous head, green here.
Review finding: spawn failure of the import probe reported as a fatal timeout.
Builds on the salvaged EIO suppression: the reporter asked that logging
"degrade gracefully and identify the affected path". Print one stderr line
naming the file and errno when the stream first fails, reset the flag when
`_open()` succeeds again so a later failure is reported anew. The contributor's
test is replaced by one invariant: five records through a stream raising EIO
produce zero tracebacks and exactly one path mention, and the next emit
recovers into the real file.
The critical-module import probe (`_critical_module_import_failures`) imports
`run_agent`, whose module-level `load_hermes_dotenv()` resolves every enabled
external secret source. The main-process skip only lived in `hermes_cli.main`'s
own dotenv call (`load_external_secrets=sys.argv[1:2] != ["update"]`), so the
probe child ran op/bws/command helpers with a 120s per-source budget inside the
120s probe and a healthy install reported "critical-module probe still fails to
import after updating: timed out before reporting import health" (#110823).
Move the argv check into `_early_recovery._should_skip_external_secret_sources`,
which every dotenv load already consults, and stamp `sys.argv = ['hermes',
'update']` into the probe so its imports inherit the updater contract.
Invariant test spawns the real probe against a home whose configured helper
touches a marker: red on main, green here.
_purge_stale_hermes_modules evicts hermes_cli.* but not root modules
(hermes_constants, utils, toolsets, ...), so the post-purge
`from hermes_cli.config import ...` re-executes the NEW config.py against
OLD cached root modules. When a pull adds a root-module symbol that
config.py imports at module level, that raises ImportError. The import
sat outside the step's try/except, so the error escaped
_run_post_update_maintenance and aborted the fleet restart -- where
origin/main printed the "Could not check config version / run hermes
config migrate" fallback and continued. Move the purge, reload and
import inside the existing try so the step stays fail-open.
Test: a cached hermes_constants lacking get_process_hermes_home (the
symbol hermes_cli/config.py imports at module level) must print the
fallback and return instead of raising. Red before, green after.
Review finding: fail-open -> fail-closed flip; post-purge hermes_cli.config import outside the try let ImportError abort post-update maintenance.
The pinned-list reload (tools_config) fixes the reported symbol; any
future symbol a pull adds to any module a migration imports at call time
(agent.skill_utils, hermes_cli.toolset_scope, ...) would fail the same
way. Evict every cached Hermes module at the migration entry point with
_purge_stale_hermes_modules — the class fix the fleet-restart phase
already relies on — so the migration graph is rebuilt from the new tree.
Tests: the reporter's exact shape (cached tools_config lacking
_configurable_keys, on-disk v44 config with an explicit platform
toolset list) must still migrate v44 -> v45 through the updater's own
entry point. Existing entry-point tests stub the purge like the rest of
the update suite does.
The updater process is the pre-pull process: hermes_cli.tools_config cached in sys.modules lacks symbols the pull added, so a migration that imports them at call time fails with ImportError and the config silently stays at the old version (#111271).
Review follow-up on #91651: _recover_terminal_input_modes and the termios
drift check used the same `now - 0.0 < interval` idiom as the repaint
throttles. Their windows (0.5s / 1.0s) are unreachable in practice, but
converting them to the None sentinel retires the bug class instead of the
instance, so nobody "simplifies" a None back to 0.0 later. Also adds the
missing regression test for _invalidate, the throttle the PR title is about.
time.monotonic() counts from an arbitrary epoch (boot on Linux). Two CLI
repaint throttles used 0.0 as the never-fired sentinel, so on a freshly
booted VM (CI runners, containers) now - 0.0 < min_interval suppressed
the FIRST repaint ever requested:
- _schedule_focus_regain_redraw: min_interval=60 suppressed the first
focus-regain redraw whenever uptime < 60s — the exact failure in CI
run 32494557030 (test_focus_regain_redraw_is_rate_limited, both
attempts red on a fresh runner, green everywhere else).
- _invalidate: same 0.0 sentinel; a first spinner/stream repaint inside
the first 250ms of uptime was droppable the same way.
Both now use None as the never-fired sentinel. Regression test pins
monotonic()=3.0 with min_interval=60 and asserts the first redraw fires.
_reload_dynamic_routes runs from _handle_webhook and admitted routes through
_dynamic_route_allowed without validate_coalesce_config, so a dynamic route
with a malformed coalesce block (no key, non-numeric window) raised inside
the request handler on its first event instead of being rejected up front
like a static route. The admission check now runs the same validator and
skips (warns on) the offending route. The parametrized rejection test covers
the dynamic path too; upstream-source reference dropped from the module
docstring.
Rapid distinct events on the same logical entity (five pushes to one PR,
a burst of ticket edits, a flapping alert) each carry a fresh delivery ID,
so the idempotency cache cannot suppress them and every event wakes a
separate agent run. Roomote solved this for PR review tasks by keeping one
durable review task per PR and superseding stale heads; this ports the
same debounce-and-supersede pattern to the generic webhook adapter.
New opt-in per-route 'coalesce' block: events group by a payload-derived
key, each new event replaces the pending one and re-arms a quiet-window
timer (window_seconds, default 30), bounded by max_wait_seconds (default
300) past the group's first event so a steady stream cannot starve
dispatch. The settled group dispatches ONE agent run on the latest
event's payload/prompt/delivery templates, with a note when earlier
events were superseded. Pending groups flush on disconnect. Startup
validation rejects missing keys, non-positive windows, and the
deliver_only+coalesce combination.
Rebase onto the decomposed webhook adapter (salvage, #92066):
- Coalescing lives in a topical sibling, gateway/platforms/webhook_coalesce.py
(WebhookCoalescer + validate_coalesce_config); webhook.py only wires it in
(__init__, _validate_route, disconnect, _handle_webhook) and splits main's
_dispatch_agent_run into the HTTP-response wrapper plus _spawn_agent_run,
shared by the immediate and coalesced paths.
- Review finding (unresolved key fields collapsed unrelated entities into one
group): an event whose rendered key still contains a {placeholder} is now
dispatched immediately instead of coalesced; documented.
- Review finding (flush-on-disconnect vs process exit): disconnect() awaits
the handoff of flushed runs; the docs claim is scoped to adapter disconnect
and states that a hard kill loses the current window's buffer.
- cron_job + coalesce is rejected like deliver_only + coalesce (cron_job
landed on main after the PR branched).
- Tests trimmed from 17 to 4 (validation parametrized; debounce/supersede/
independent groups/duplicate-first in one behavioural test; max-wait +
unresolved-key; flush-on-disconnect).
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".
Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.
Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
The Restart Gateway button spawned `hermes gateway restart` as the dashboard's
own user. On a systemd *system* install the CLI refuses every lifecycle verb
below root (`_require_root_for_system_service`), so the button could never work:
the refusal landed in `~/.hermes/logs/gateway-restart.log` while the endpoint
reported a started action. `start`/`stop` shared the defect.
`_spawn_hermes_action` now prefixes `sudo -n` for gateway restart/start/stop when
the action resolves to the system unit. Scope comes from the CLI's own picker
(`_select_systemd_scope`) evaluated against the profile the action addresses, so
a host with a user unit installed keeps the unprivileged path and root never
shells out to sudo. With no passwordless path the request fails with an
actionable message instead of reporting a started action whose child refuses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gateway_declares_external_supervisor treated the control-socket identify answer
as final whenever `supervisor` was non-empty and only accepted "external".
_detect_supervisor() answers "systemd" as soon as INVOCATION_ID is set (and
"launchd" under the XPC name), so a gateway run by a custom, non-canonical
systemd unit / launchd agent with `ExecStart=... gateway run --external-supervisor`
— the documented contract — was classified non-external and `hermes gateway
restart` still took the stop + foreground run_gateway path that stamps the CLI's
PID and wedges every respawn (#110637). The update path's argv check on the same
gateway already said external-supervisor, so the two restart paths disagreed.
Any self-declared supervisor other than "manual" now means hand back (the
declaring supervisor owns the respawn); otherwise the argv/state-file marker
decides as before. The drain-failure message no longer claims a timeout when
SIGUSR1 could not be sent at all.
Review finding: socket-first detection classified a custom systemd unit's
--external-supervisor gateway as manual and took the wedge path.
The handback logic was appended to the hermes_cli/gateway.py facade; it now lives in a
topical sibling. Supervisor detection also reads the gateway's own declaration (control
socket `identify` -> supervisor: "external", then the live argv marker, then the argv the
gateway stamped into gateway_state.json) so a gateway whose command line cannot be read via
psutil is still handed back rather than SIGTERMed and shadowed by a foreground run.
Tests trimmed to the two invariants (handback with fresh-PID success; either failure branch
never takes ownership) plus the plain-manual control. Docs: `hermes gateway restart` is now
part of the --external-supervisor contract.
Address review feedback on the SIGUSR1 handback:
- A graceful exit no longer reports success by itself: the pidfile is
polled for the supervisor's replacement (bounded wait, identity via
get_running_pid's lock+PID liveness check, freshness via != old PID)
before the restart is called done. An unloaded/broken supervisor now
surfaces as a failed restart instead of success printed over a dead
gateway — the same contract launchd_restart enforces through
_wait_for_launchd_service_pid.
- A drain timeout no longer falls back to SIGTERM + foreground run: that
fallback recreated the competing-owner wedge (#110637) this handback
exists to remove. Both failure branches keep the supervisor as the
sole restart owner and fail loudly (exit 1) instead.
Credit: both lifecycle branches flagged by @ehz0ah in review.
A custom launchd agent (plist outside the canonical ai.hermes.gateway path)
is invisible to _installed_service_kind_for, so 'hermes gateway restart'
fell through to the manual stop + foreground run_gateway fallback. The
foreground run stamps the restart CLI's own PID into gateway.pid, and every
KeepAlive respawn of 'gateway run --external-supervisor' then refuses with
'Gateway already running (PID <restart>)' — the gateway stays down until
the restart process is killed (#110637).
Mirror the update path (_prepare_profile_gateway_update_restart): trust the
--external-supervisor argv marker on the live gateway and SIGUSR1 it
(_graceful_restart_via_sigusr1) so it drains, exits, and lets the supervisor
relaunch it. Drain failure falls back to the existing stop/restart path.
Fixes#110637
_marker_only_restart_obsolete cleared the marker when every row from
collect_fleet_versions() was current. At CLI startup the probe runs without
pre_restart_pids, so a gateway the restart phase stopped and never brought
back produces no row at all (dead-pid records are skipped, no DOWN
classification possible). With alpha current and beta down the marker was
discharged and beta's catch-up restart never happened.
Before clearing, also require every gateway identity latest.json owes
(plan.runtimes / fleet entries) to be covered by a current row; the
rows-only rule applies only when the receipt names no gateways. The
identity extraction is shared with _live_fleet_covers_receipt.
Review finding: marker discharged with a DOWN sibling gateway absent from the startup probe.
collect_fleet_versions now classifies a live pid from the state file only when the
home's identity resolver verifies it (#110420); these two tests write the record about
the pytest process itself, so they verify it explicitly. Dead-pid rows are unchanged.