The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.
model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).
Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
CI's deep pytest tmp_path pushes state/gateway.loop-tick.<pid>.sock past
the sockaddr_un limit; the producer swallows the OSError into
loop_tick_socket=False and the POSIX arm test fails falsely. Use a
mkdtemp under the temp root, same as test_update_wedged_gateway.py.
asyncio.start_unix_server does not exist on Windows, so the ungated
call raised AttributeError on every gateway start — a warning +
traceback in the logs each time, with the witness permanently absent
(the documented deliberate fail-safe). Gate the server creation inside
the existing os.name == "posix" block, mirroring the stale-socket
sweep above it, and log the non-POSIX skip at debug. POSIX behavior
and the two-witness liveness contract are unchanged.
Closes#96956
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
A wall-clock step-back (NTP) can leave the marker's mtime in the future,
which would push the idle window out by the step size. Clamp with
min(mtime, now); _last_inbound_at has the same exposure but is at least
bounded by process uptime. Review comment on #100830.
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.
The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
Collapses the three test-refinement commits from #100614 (423f7f8703,
157db13c76, ffa72dd67f): model recovery pressure by provider-call state,
assert the rebuilt-oversized retry fails closed, keep the compacted
history compressible (user+assistant summary rows).
A scoped projects.tree / projects.project_sessions response is built from
ONE profile's state.db, so the request scope is authoritative even for
legacy rows whose persisted profile_name is NULL. Without the stamp those
rows reach the renderer ownerless, and every owner lookup off them — the
branch path included — falls back to whichever backend is active.
Co-authored-by: evan-bradford <evan-bradford@users.noreply.github.com>
Review follow-up on #93959:
1. Partial-failure window: if the row commits but the transcript copy or
title write fails, the durable-but-empty child defeated the lazy
first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
the renderer fail-latched on a transcript-less session again. The seed
block now compensates: delete just this child so the lazy path can
retry cleanly. Disk-full is exempt — deleting data on a full disk makes
things worse.
2. Silent degradation: the best-effort catch now logs at WARNING with
exc_info instead of DEBUG, so a regression in this user-facing path is
observable without enabling debug logs.
Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.
A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:
- Row created in the PARENT's profile-scoped state.db, stamped with
_branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
from pending_title — the branch lands in the parent's lineage instead
of falling back to a message-preview name.
Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.
Fixes#93959
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
A concurrent WAL checkpoint / reset / frame-flush can surface SQLITE_IOERR
to a reader on a perfectly healthy database: a mode=ro connection cannot
perform the WAL recovery the read needs, because recovery writes the -shm
index and read-only mode refuses. The window is millisecond-scale.
Today that one-shot error escapes the SessionDB read-only constructor, and
GET /api/sessions turns it into a 500 the desktop reads as an authoritative
empty list.
Retry it, bounded, in the constructor so every read-only opener is covered —
the sidebar poll, cross-profile aggregation, recall, browse — rather than at
one route. A persistent IOERR still exhausts the budget and propagates.
Remaining transient failures answer 503, so the client keeps the list it has.
On the write path, BEGIN IMMEDIATE can hit the same transient IOERR before
the callback runs. That one is safe to retry on the same connection because
nothing has been mutated; once the callback starts, settlement is unknown and
the error propagates. Never close()+reopen to heal it — close() cancels this
process's POSIX advisory locks on the file for every sibling connection, and
a list poll's reader must stay disposable so a replaced state.db is observed
and the pre-repair forensic backup stays reachable.
Fixes#100436
Co-authored-by: rkfshakti <rkfshakti@users.noreply.github.com>
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
Review finding: dashboard_client_last_seen() discarded the marker once it
was >= 45s old, so after the client disconnected the gateway fell back to
its own (much older) _last_inbound_at and suspended ~45-75s later, not
idle_timeout later as the PR claimed. The staging release leg had in fact
shown 46s.
The marker mtime is a timestamp of real inbound; is_idle already judges
recency. Drop the staleness cutoff entirely and always fold the raw mtime
into the inbound clock (max with _last_inbound_at). An old marker is
harmless: it is outside idle_timeout just like an old _last_inbound_at.
Also:
- touch the marker immediately after ws.accept(), before the ready/skin
setup, so a client waiting on a slow ready frame is still visible
- make the newer-message-wins test discriminating (timeout 10s, chat 5s,
marker 40s: choosing the marker would read idle)
- add < idle_timeout / >= idle_timeout / ancient-marker boundary tests
Re-validated live on hermes-agent-stg-test-6698: last client frame
03:18:11Z -> going dormant 03:20:17Z (126s = 120s timeout + watcher tick)
while the gateway's own inbound clock was 368s stale. Mutation check vs
origin/main files: run.py hunk reverted -> 4 fail, ws.py reverted -> 3
fail, "prefer marker over max" -> 1 fail.
* fix(linux): install Lanczos-resized panel icons, not a PNG in scalable
Cinnamon's panel is ~24px. v2026.8.31 dropped the 1024px asset into
hicolor/scalable (SVG-only), so the Mint panel nearest-neighbor scaled
it into a mangled blob. Decode the PNG and write 24/32/48/256 rasters;
undecodable bytes still copy into one indexed dir. Drop the leftover
scalable file.
* test(linux): cover resized hicolor panel icons and stale scalable cleanup
Pin that a decodeable PNG lands as 24×24/256×256 rasters (not scalable),
a leftover scalable copy from v2026.8.31 is deleted, and truncated
PNGs still fall back to an indexed copy.
The gateway's idle predicate only stamped _last_inbound_at for messaging
inbound, so a scale-to-zero instance suspended under an open desktop app /
web dashboard / TUI. The client's reconnect loop then re-poked the
Fly-proxied hostname, autostart resumed the box, and the instance flapped
suspend -> proxy-wake every ~60s (13 of 72 active opted-in prod instances
on 2026-09-02, with [PC05] connect timeouts visible to the user).
The dashboard runs in a separate process on hosted instances, so the
signal crosses over as a marker file under HERMES_HOME/state:
- tui_gateway/ws.py touches state/dashboard_clients.heartbeat on every
/api/ws connect and inbound frame (clients gateway.ping every 15s),
throttled to one write per 5s per process.
- gateway/scale_to_zero.py: dashboard_client_last_seen() reads the mtime;
a marker older than 45s (the clients' heartbeat deadline) is "client
gone", a missing marker is "no client" (not fail-awake, or nothing
would ever sleep), an unreadable marker fails awake.
- gateway/run.py folds that into seconds_since_last_inbound, so an
attached client gets exactly the same idle_timeout grace after it
disconnects as a chat message does. No new conjunct, _last_inbound_at
itself is not mutated.
Validated as a hot patch on hermes-agent-stg-test-6698 with a real
/api/ws client pinging every 15s: machine held awake for 6.5 min of zero
proxied traffic (was 2 min), then suspended within ~50s of the client
exiting. Tests exercise the real seams (temp HERMES_HOME, the runner's
_scale_to_zero_is_idle composition, tui_gateway.ws.handle_ws) and were
mutation-checked: reverting either half or dropping the staleness check
fails 2-3 of them.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
prompt.submit now retries a completed failed build (installing a fresh
unset agent_ready before rebuilding), so a no-op _start_agent_build stub
leaves the patient wait blocking forever — the turn thread outlived the
test and the whole file hit the per-file timeout on CI. Stub a faithful
failing build instead: set agent_error, fire the session's current
ready event. The pinned contract (visible failure, no silent drop) is
unchanged.
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).
Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.
Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes#97994.
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.
Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
Problem B of #71047: with streaming + reply_to_mode='first', the streamed
preview is a reply-quote of the user's message. When the turn-final edit
hits flood control, the empty-tail fresh-commit resend either (a) also got
flood-capped -> the consumer reported 'failed', the gateway's normal final
send fired, and the never-deleted preview + the fresh final left TWO
visible bubbles, or (b) succeeded but as a plain non-reply message that
didn't match the preview's anchor.
- preserve the turn's reply anchor (initial_reply_to_id) on the
empty-fallback fresh-commit resend so the replacement message quotes the
user's message exactly like the preview and the non-streaming path
- retry a flood-rejected preview deleteMessage once (delete_message
returns False rather than raising) so the stale preview doesn't linger
next to the fresh final; still best-effort, and the preview is only ever
deleted AFTER the replacement send succeeded
- regression tests for the anchor, the delete retry, and the
flood-capped-resend single-bubble suppression decision
Builds on @fangliquanflq's PR #96097 ('preview' verdict for flood-rejected
fresh commits), cherry-picked as the previous commit with the conflict
against fd998120c1 resolved (record the payload AND keep
_delivery_ambiguous only for real timeouts).
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:
- hermes_cli/auth.py Codex OAuth clients (token refresh at
auth.openai.com/oauth/token, device-code login, token exchange, usage
probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
every connect eats the full timeout per AAAA before IPv4 is tried, so
auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
had no explicit racing wired.
Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
installs the existing _HappyEyeballsSyncBackend on a ready-built sync
httpx.Client's direct transports (default transport + mounts), skipping
proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
five Codex OAuth/probe endpoints. Best-effort: falls back to default
serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
no custom backend needed. Documented in build_keepalive_http_client and
pinned by tests (contract test on the anyio signature + a live
regression test where a blackholed 100::1 IPv6 addr hangs and local
IPv4 wins in ~250ms instead of the serial connect timeout).
network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.
Refs #13834; follows #94388 (9cce8725).
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).
Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
the shared client's pooled sockets (force_close_tcp_sockets — FD release
stays with the owning worker), abort+poison of the cached per-request
openai/anthropic wire clients, Codex app-server request_interrupt(), and
the inline _active_request_abort hook. Never client.close(), never
socket.close().
- delegate timeout path: after registering the deferred-close callback,
run one immediate drain plus one 5s re-sweep (covers a connection opened
between the interrupt and the first sweep). The settled read (EOF/EPIPE)
lets the worker unwind, which triggers the deferred close on the worker's
own thread — the only safe FD-release boundary. A worker that still never
settles retains its resources rather than risking a cross-thread close.
Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.
Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.
Closes#94248
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
provider label never carries the vendor token to the alias host (#28660);
route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
Regression for #58576: _profile_scope holds _SKILLS_PROFILE_LOCK across
the payload build, which can block up to 15s on a models.dev cache miss
and starve concurrent /api/config on the same lock. The test records
which scope the handler enters for a selected profile and asserts only
the config-only (contextvar) scope is used.
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.
Three surgical changes:
1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
The inlined watcher now routes the respawned gateway's stray
stdout/stderr to the same sidecar log gateway_windows._spawn_detached
uses (DEVNULL only as fallback), so a gateway killed moments after
respawn leaves a trace. Direct implementation of the 4th repro's
hardening suggestion (1).
2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
canonical _spawn_detached, so the respawned gateway's exit-diag /
lifecycle records show whether it escaped the parent Job Object — a
job-teardown kill is no longer indistinguishable from any other silent
death.
3. Post-update resume verifies liveness before vouching
(hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
runs the same provisional-hit + 2s-confirmation liveness poll every
other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
with all_profiles= for the fleet) before printing ✓, writes the #91675
start attestation for the verified PIDs, and fails the resume with a
"restart could not be verified" warning + recovery hint when no stable
gateway appears. Suggestion (2) of the 4th repro; closes the last
silent-success hole in the family (#84185 fixed the cold-start leg,
#91675 the direct-start leg; this is the relaunch leg).
Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.
Fixes the Bug-1 relaunch-trust leg of #48820.