Review follow-ups on the identity binding:
- The sentinel's `start_time` is `time.time()` at `record_startup`, seconds
after the process was born once imports finish, so comparing it with
psutil's create_time within 2 s would have read every real gateway as
undecidable and silently stopped the #109538 cold-start. `record_startup`
now stamps `create_time` (psutil birth via the existing
`process_identity._process_create_time`), `mark_exited` carries it, and
the attestation compares birth to birth. A sentinel from a gateway older
than the stamp falls back to the PID-only rule.
- A resume token written by pre-generation code and resumed by this code
probes the marker again instead of skipping the spawn.
- Horizon allows a 60 s backwards clock step; the unused `now` parameter is
gone; the create-time tolerance is a named constant; the read-then-unlink
in `_consume_start_attestation` is documented as best-effort.
The resume token only recorded `cold_start_if_installed: bool` and execution re-read the
mutable one-shot start-attestation marker to decide whether a Desktop-owned install still owed
a cold-start. A concurrent `hermes gateway status`/`start` (`check_start_attestation`) consumes
that marker between plan and execution, so the spawn was skipped and the token cleared with no
gateway running.
The marker now carries a `generation` nonce. The plan records the generation whose unclean
death authorized the cold-start on the token (`attested_generation`); execution authorizes the
spawn from the token, still re-checks live gateway PIDs, and consumes the marker only while it
is still that generation — a newer marker written by a concurrent start keeps its own report.
`attested_gateway_died` becomes `attested_death_generation` (`None` = undecidable, fail closed).
Follow-up to #110020 review thread (a).
_cold_start_windows_gateway_after_update cleared the dead attestation as soon
as _spawn_detached() returned a PID, before _wait_for_gateway_ready() proved
the gateway survived. When readiness failed, the RuntimeError registered the
retry, but the retry then saw Desktop lifecycle ownership with no marker and
returned success without spawning anything — the silent outage of #109538
came back through the retry path (#110020 review). The marker is now
consumed after readiness is confirmed, so a failed spawn leaves the retry
its recovery obligation.
attested_gateway_died() re-ran find_gateway_pids() (current profile only)
although both callers had just proven the process table empty with
all_profiles=True, and it re-implemented check_start_attestation's
liveness rule. Callers now pass the liveness they hold (current_pids=[])
and both probes share _attested_dead(), so the consuming and read-only
twins cannot drift.
attested_gateway_died() is deliberately read-only so the CLI-start warning
still fires, but that left the dead marker in place after the update path
acted on it. If the restored gateway never became ready (or died again
before the next CLI start consumed the marker), the same stale crash marker
would re-authorize another cold start against Desktop ownership on the
next update. Clear it via the existing _clear_start_attestation() path
right after _spawn_detached() succeeds - the marker has done its job at
that point; a new one is written once the spawn is confirmed ready.
A Desktop self-update hand-off exits the app before the updater runs and can
kill the messaging gateway in those same seconds (#109538), so the updater's
discovery finds no live PID while the one-shot start attestation still
vouches for the dead one. Both Desktop-ownership checks then read "nothing
running" as "nothing to restore" and the bot stayed down until a manual
start.
Consult the attestation non-destructively before Desktop-owned lifecycle
suppresses a cold-start: a vouched-for PID gone without a clean ledger exit
keeps the plan and is restored; no attested death preserves the #76129 skip
unchanged.
Canonicalising _is_backend_argv onto _hermes_holder_subcommand dropped the
'-m hermes_cli.main' entry-shape discriminator the old predicate had. That
widened _orphaned_desktop_backend_pids to tree-kill a standalone hermes.exe
serve / hermes dashboard whose console parent died. The Desktop's only spawn
shape is -m hermes_cli.main (apps/desktop/electron/main.ts), so _is_backend_argv
now forwards to _looks_like_desktop_control_plane — the same canonical-subcommand
AND entry-shape predicate — instead of a second copy. Test matrix gains a
'Desktop backend?' column with plain hermes serve / hermes.exe dashboard rows.
Two kill/relaunch predicates decided identity by argv substring, the bug class root AGENTS.md
forbids: hermes_cli/dashboard_procs.py::_is_desktop_local_serve_cmdline (`"serve" not in cmd`,
on the orphan-reap KILL path) and hermes_cli/update_cmd_windows.py::_is_backend_argv
(`" serve" in argv_low`, in the very file that defines _hermes_holder_subcommand). Both now ask
the canonical token classifier; host/port are read as flag values, not substrings.
hermes_cli/profiles.py::_check_gateway_running open-coded rungs 1/3 of
gateway.status.resolve_gateway_liveness and skipped the multiplexer rung; it is now that ladder
scoped to the profile dir (pid probe keeps cleanup_stale=False so a probe for another profile
never unlinks its PID file). The gateway/status.py ladder itself is untouched.
Behavior change: `hermes kanban --preserve-cache --host 127.0.0.1 --port 0` and
`-m dashboard serve`-style argv are no longer classified as serve backends (never killed /
relaunched as one); a named profile served by the live default multiplexer now reads as
running from _check_gateway_running (previously only via the separate
_served_by_running_multiplexer OR at some call sites).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.