Review follow-up. The two-phase switch closed the runtime-id leak for a
single switch, but two shapes still broke the wipe→publish ordering:
Overlapping: click A commits (wipes) and waits for its activation behind
the profile-store mutex; click B finishes its dial and wipes too; then A's
activation lands and publishes A AFTER B's destructive wipe, and A's
endGatewaySwitch() dropped the (boolean) barrier while B was mid-commit.
Now the wipe runs INSIDE the serialized section via a `beforeActivate`
commit hook on ensureGatewayAgent, synchronously right before the socket
is activated; the hook re-checks the click revision and declines when
superseded, so a queued-then-superseded switch neither wipes nor
activates. The barrier is token-owned (latest switch wins), and the mutex
wait is a loop so waiters that wake together can't run interleaved.
Stalled: a wedged spawn / ticket mint / IPC (the #93454 class) left the
spinner up, swallowed later clicks on the same source, and — inside
ensureGatewayAgent — latched the mutex and the barrier. Every await is now
bounded (dial, activation, descriptor lookups, setLastUsed) via a shared
withTimeout helper extracted from use-gateway-boot. A commit that timed
out after the new source was already published counts as committed; one
that stalls after the wipe lowers the barrier and repaints the source
that is still active.
Tests: queued-then-superseded switch, dial that never answers (and retry
not swallowed), activation stalled after/before publication, barrier
ownership, commit-hook ordering inside the mutex, bounded descriptor
lookup releasing the mutex. All fail against the previous revision.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Sessions sidebar switcher (selectConnection) activated the target
gateway first and wiped session bindings afterwards — across an IPC
round-trip — so route/session effects saw the new source while
$activeSessionId still named the previous backend's runtime id, sent it
to a backend that never minted it, and got "session not found". The
Settings → Gateway apply (softSwitch) never had this problem: it raises
$gatewaySwitching, runs beforeConnectionSwitch and wipes before dialing.
Give both doors one commit point. store/gateway-switch gains
beginGatewaySwitch()/endGatewaySwitch(): barrier up, the registered
machine-context reset, session wipe — in one synchronous step.
useGatewayBoot registers its beforeConnectionSwitch/refreshSessions with
the store and softSwitch goes through the same entry point.
selectConnection becomes two-phase: dial the target WITHOUT activating
(openGatewayAgent → openGatewayForAgent with an activation lease, so a
live-work recompute can't prune it mid-spawn), then beginGatewaySwitch()
and activate the already-open socket with no await between the wipe and
the publication. A dead target fails with the current workspace intact;
a superseded click never activates its target; an activation that fails
after the wipe repaints the still-active source instead of leaving the
sidebar skeleton.
Tests: real-store end-to-end regression (real useGatewayBoot + gateway
registry + selectConnection over fake sockets) — red on the previous
code with the leaked runtime id observed at publication — plus store
ordering/failure-path, commit-point and prune-lease coverage.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-ups on the #85416 salvage:
- _store_bin_names()/_sibling_names() derive the versioned interpreter name
from the RUNNING interpreter instead of a hardcoded 3.11-3.13 list, with a
sorted python3.* glob fallback for stores/aliases built by a different
Python minor (major bumps, fixtures) — no future version bump can silently
leave an alias resolving back into the versioned uv store.
- _repoint_alias_symlinks unions expected alias names with the versioned
symlinks actually on disk before repointing.
The sidebar row's Tip trigger test derives its fixture from the wall clock:
const startedAt = Math.floor(Date.now() / 1000) - 5 * 60
...
expect(age.getAttribute('aria-label')).toMatch(/^5m, Today at /)
"Five minutes ago" is only today when the run does not straddle local
midnight. Between 00:00 and 00:05 the timestamp falls into the previous
day, formatMessageTimestamp correctly returns the yesterday label, and
the test fails on a day boundary it was never written to exercise:
AssertionError: expected '5m, Yesterday at 11:56 PM'
to match /^5m, Today at /
This is a real CI failure, not a theoretical one - it took down a
check:test:ui shard on an unrelated desktop PR at 00:01 UTC, and it will
do so for any PR whose shard happens to land in that five-minute window.
Pin the clock to local noon before deriving the timestamp so the fixture
can never cross a day boundary. Only Date is faked (toFake: ['Date']),
so the component's own timers - the running arc and the tooltip open
delay - keep running for real; the neighbouring tooltip tests that
advance timers are unaffected. The describe block gains the
useRealTimers teardown its sibling already had.
The production formatter is not changed: rendering "Yesterday at 11:56
PM" for a session started five minutes before midnight is correct, and
the assertion is about the label's composition, not about which day it
names.
The #87196/#87720 conflict resolution kept the bounded-drain helper and
its constant; the windows_only kill-tree test still asserted the dropped
_CUA_INSTALLER_REAP_TIMEOUT name. Same 2-communicate contract, surviving
constant.
Re-enables the routine confirmed-upgrade path on Windows that #95008
deferred wholesale, now that every unattended-hostile surface is closed:
- stdin=DEVNULL (salvaged #79871): upstream's Read-Host consent prompt
can't block a hidden console.
- Bounded post-kill drain (salvaged #87720): a kill-surviving descendant
holding the stdout pipe can't strand the update past its ceiling.
- 120s background ceiling (salvaged #87196): safe now that a legitimate
600s lock wait can't occur on this path.
- NEW lock preflight: upstream's install lock held by a live process ->
skip in ~0s instead of eating its 600s stale-lock window (the actual
11-minute hang observed 2026-08-25; UAC was a red herring — base
install is no-admin by upstream design).
- NEW 5s network preflight: github.com unreachable -> skip immediately.
- Windows unattended runs pass -NoAutoStart, skipping the ONLY
install.ps1 branch that self-elevates (autostart task re-registration).
- Timeout diagnosability: partial installer output is logged on kill so
the next hang names its stage instead of dying silently.
Contract repairs and fresh installs stay interactive-only (SmartScreen /
first-time elevation legitimately need a human).
On Windows, `hermes update` can hang past its own 660s cua-driver timeout
until the user kills an orphaned PowerShell by hand. The timeout ceiling is
not the problem; the code that runs after it is.
`_run_cua_driver_installer` handles `TimeoutExpired` by killing the process
tree and then draining the pipes with a bare `proc.communicate()`. The kill
is best-effort by construction: every `psutil.Error` in `_kill_installer_tree`
is logged at debug level and stepped over, on the reasoning that a partly
killed tree beats none. That is the right call, but it means the drain has to
survive a partial kill, and an unbounded drain does not.
The concrete case is the one reported. `install.ps1` self-elevates through
`Start-Process -Verb RunAs`, so the descendant runs at High integrity and a
medium-integrity `child.kill()` raises `AccessDenied`. The per-child handler
logs it and continues. That survivor is still holding the `stdout=PIPE` write
handle it inherited, so the following `communicate()` waits for an EOF that
arrives only when somebody kills that process manually. A bounded 660s wait
becomes an unbounded one, after the warning has already printed.
Bound the drain instead. A kill that landed closes the pipe immediately, so
this costs nothing on the normal path; a kill that did not costs 15s rather
than forever. The original `TimeoutExpired` is re-raised either way, so the
existing manual re-run hint still prints and the update unwinds. Losing the
tail of a timed-out installer's log is the cheaper half of that trade, and it
is only lost in the case where the run already failed.
The drain deliberately does not close the pipe handles. `communicate()`'s
reader threads are still blocked on them and closing underneath them races;
they are daemon threads, so abandoning them does not hold the interpreter
open.
Both timeout handlers (streaming and captured) now go through one helper.
The streaming child inherits the console rather than a pipe, so it is much
harder to stall there, but the two branches should not drift on a rule this
small.
Tests: 5, in a new `TestInstallerTimeoutDrainIsBounded`. Two fail without the
fix, including the reported scenario end to end (a child kill refused with
`psutil.AccessDenied`, asserting the drain still carries a deadline). The
deadline is asserted as a kwarg rather than by timing, because a test that
proved the hang by hanging would be the same defect wearing a test's name.
Scope note: this does not touch the `stdin` inheritance that lets
`install.ps1`'s `Read-Host` block in the first place. That is #79684 and open
PR #79871 already carries the one-line `stdin=DEVNULL` fix; the two are
independent and neither subsumes the other, since `DEVNULL` cannot unblock a
UAC elevation dialog.
Fixes#87703
When `hermes update` runs the cua-driver installer non-interactively,
stdout is captured (PIPE) but stdin is inherited from the parent process.
The upstream install.ps1 prompts `[Y/n]` when it detects a stale
cua-driver daemon, but the prompt goes to captured stdout (invisible)
while stdin waits for input — causing an 11-minute hang until the
watchdog timeout.
Fix: redirect stdin to subprocess.DEVNULL on the non-verbose path so the
installer reads EOF immediately instead of blocking. The installer exits
quickly with a non-zero code, and the existing handler shows the manual
re-run hint with the installer's captured output.
Fixes#79684
profiles.list opened every profile state.db as a writable SessionDB,
which waits out write-lock patience while that profile's backend is
mid-turn. The desktop RPC timed out and Bot Mode's infinite React
Query retry kept the sidebar on a spinner.
Inspect those DBs read-only and bound roster retries so names still
paint.
Two follow-ups on top of the #86391 salvage:
- check_macos_tcc_grants: a certificate-anchored DR (hermes desktop
--setup-tcc-identity, or a notarized release) now reports as stable in its
own class instead of falling into the identifier-pinned message; the
identifier-pinned message points at --setup-tcc-identity for the strongest
anchor.
- collect_relay_plugin_cutover_findings: only merge process-level env vars
when env_map is None (run_doctor's live path). An explicit env_map is a
complete environment description — merging os.environ on top made
report_deprecated_config_and_env non-hermetic on boxes exporting legacy
relay vars (10 findings vs the expected 2 in
test_report_does_not_count_as_blocking_issue).
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
so a hanging codesign degrades to the unreadable-DR warning, never crashing
the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
GPT-OSS review: an empty codesign output would fall through to the
'stable identity' branch and false-positive. Guard with and
cover the empty-string case. Flash review: the non-macOS silence test
mocked the bundle to None, so it never exercised the platform guard;
mock a real path instead.
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.
- hermes doctor: new check_macos_tcc_grants() reports the desktop
bundle's DR class (cdhash-pinned → grants reset on every update;
identifier-pinned → stable) and prints the exact stale-grant repair
(tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
documents the one-time re-grant for pre-fix grants.
Closes#86385
The heal decision is extracted into should_heal_self_marker_refusal()
so the contract is testable: heal ONLY on exit 2 + a marker naming this
process. Five tests pin it — self-owned heals, foreign owner (real live
sibling process) never heals, missing/garbage marker never heals,
non-exit-2 never heals, and the full acquire -> refuse -> drop-claim ->
retry-precondition lifecycle with a real UpdateMarkerGuard.
windows-rust-e2e.yml mirrors the wine2e pattern: fires only on
wine2e-rust/** pushes, runs the crate's cargo test --lib on
windows-latest (the shipping platform). The permanent Linux lane stays
authoritative for the unix-gated pipe-drain fixtures.
Main's live_marker_owner has adopted self-owned markers since
160586ff8/dbc2a9c8e (#74761), so the cherry-picked comment's claim that
it 'maps self-ownership to None' is stale. The raw read is still the
right tool — the heal needs the single fact 'does the marker name our
PID' without age/liveness policy folded in.
An updater binary spawns 'hermes update' while holding the update marker
with its own PID. A checkout that predates the HERMES_UPDATE_HANDOFF_PID
env fix (8c76fe19f) and the ancestor-pid fallback runs its pre-pull
update_lock.py, reads that marker as a live foreign update, and exits 2
— and the updater deliberately skips its retry for exit 2, so the
refusal loops forever: the update being refused is the one that ships
the fix, and the failure screen's Retry re-enters the same state.
Detect the case with a raw marker read (live_marker_owner deliberately
maps self-ownership to None, so it cannot answer this), drop our own
claim, and retry the child once with the marker absent. The guard
re-removes on Drop (idempotent) and the desktop is already gone at this
point, so nothing races the brief marker-free window.
Fixes#75788
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Background processes started by subagents (task_id sa-*) route their
notify_on_complete / watch_pattern notifications to the parent
conversation (b95ec1cb5) because anything outliving the child needs a
durable consumer. In practice these 'npm ci finished' walls are noise
mid-conversation — the child's consolidated delegation result is the
deliverable.
- New config key delegation.surface_child_process_notifications
(default false = suppress). Flag true restores the previous behavior
exactly (delivery with subagent attribution line).
- drain_notifications drops (never requeues) completion/watch_match/
watch_disabled events whose task_id starts with 'sa-' when the flag
is false, logging at debug with session_id+task_id for diagnosis.
Requeueing would pin them forever — children never drain notifies.
- async_delegation events are NEVER suppressed (they ARE the result).
- watch_disabled emitters now carry task_id so sa- sessions' safety
events follow the same suppression as their other events.
- Config read errors fall back to the default (suppress) and never
crash the drain loop.
- Docs: delegation.md + configuration.md.
Follow-up to the salvaged #94296: the two guards covered the repair and
confirmed-update branches, but when cua-driver is enabled yet not
installed at all, control still reached _run_cua_driver_installer() and
an automatic 'hermes update' would launch the interactive install.ps1
anyway. Add the same defer before the installer run, keep POSIX
behavior unchanged, and give the confirmed-update message a natural
fallback when latest_version is unknown.
Two interaction seams between the #92693 salvage (merged as #95050) and
this branch: the source-label indexing test now compares in token space
(the stemmer shortens 'catalogsource' to 'catalogsourc'), and the
unregistered-core-name describe test forces the unregistered condition
via monkeypatch instead of depending on which sibling test file imported
model_tools first.
The parallel determinism test warms _stem's lru_cache after ~11 distinct
stems, so almost no iterations reach the underlying stemmer and a shared
(non-thread-local) instance survives it. New test bypasses the cache with
per-iteration unique tokens via _stem.__wrapped__, so thousands of stems
run concurrently: a shared stemmer's mutable parse state fails it within
2,000 calls (verified — the mutant dies 8/8 runs; healthy runs stay green).
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.
tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.
The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).
New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.
No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
RoutinesPane resolved its cron owner from a bare $lastRoster.get()
snapshot. BotsHomeView owns the roster fetch, so whenever the pane
mounted before that fetch landed (fresh boot ordering, renderer reload
resetting the atoms) it captured an empty roster forever: the pane
stayed pinned on "Cronjobs are unavailable until this agent appears in
the roster." and Create Cronjob silently no-oped until some unrelated
atom happened to re-render it (#94483).
Subscribe via useValue($lastRoster) instead, matching every other
consumer of the shared roster. Scoping intent is unchanged: a complete
focused owner without an exact roster row still fails closed rather
than routing cron reads/mutations through a stale selection or an
unscoped profile name (contracts in routines-selected-bot.test.mjs).
The source contract in focused-bot-highlight.test.mjs pinned the bare
.get() shape; it now pins the subscription form while keeping the
socket-home-atom prohibition that motivated it.
Fixes#94483
CreateRoutineDialog receives routineCreateTarget() output, which is an
owner OBJECT for roster-scoped bots; wrapping it in {name: bot} rendered
'[object Object]' and broke the meta lookup keyed by object. Resolve the
label through the object-aware botRosterMeta() path instead.
(Salvaged from #93572; the defensive coercion inside displayName was
dropped in favor of fixing the call site only.)
Review follow-up, comments and one type annotation. No behaviour change.
The unmount effect now clears pending timers, but its leading comment is
still entirely about the focus bus, so the cleanup reads as unrelated code
that happens to be in the same block. Say why it lives there: both concerns
are "this composer is going away", they unmount together by definition, and
a sibling unmount-only effect would only be a second place to forget.
In the regression test, the clearTimeout mock declared id as number. Nothing
that reaches it is a number: jsdom under node returns a Timeout object, which
is why the scheduled and cleared arrays are unknown[] and compared by
identity. The annotation documented a shape the test never sees, so it is now
unknown with the cast moved to the one call that genuinely wants a number.
Confirming an inline edit unmounts UserEditComposer while its 200ms submit
latch is still pending, so the callback resumes on an unmounted tree and calls
setSubmitting. Five window.setTimeout calls in the component, none cleared on
unmount; the latch is the one with a delay long enough to reliably survive.
In the app the stray state update is an invisible no-op. In vitest it can land
after the test file's jsdom environment is gone, and React reaches for a window
that no longer exists:
Test Files 570 passed (570)
Tests 5434 passed (5434)
Errors 1 error
ReferenceError: window is not defined
at resolveUpdatePriority react-dom-client.development.js:1308:7
at dispatchSetState react-dom-client.development.js:9126:14
at Timeout._onTimeout user-edit-composer.tsx:606:9
That is an unhandled error, so the job exits 1 with every test passing, on
whichever PR happens to be running rather than on anything related. It reads
exactly like infrastructure noise, which is the reason to fix it rather than
re-run it.
The component already knows about this hazard. Its unmount effect opens with a
comment calling it "the one composer that routinely unmounts", and two of the
other timer callbacks carry defensive try/catch for the composer core being
torn down underneath them. The timers themselves were simply never cancelled;
this cancels them instead of surviving them, which removes the race rather
than tolerating it.
All five sites now go through scheduleTimeout, which records the id and drops
it when the callback runs. The existing unmount effect clears whatever is
left. Two useCallback dep arrays gained scheduleTimeout, which is stable.
One regression test, in the file the CI error originated in: spy on
setTimeout/clearTimeout, submit an edit, unmount, assert the latch id was
cleared. It fails when the clear loop is removed. The 200ms delay is matched
explicitly so unrelated library timers cannot make it pass by accident, and
ids are compared by identity because jsdom under node returns a Timeout object
rather than a number.
Fixes#92462