A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.
Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.
Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.
No migration shims: this IS the system.
Partially reverts the newer-visible-session preference from #91791
(salvage of #91258), which made the pinned canonical Bot Chat
unreachable. Fixes#92040.
Canonical Bot Chats are ALWAYS hidden from the Sessions sidebar:
session.create passes hidden:true unconditionally and
hideOwnedBotSessions() sweeps any that were born visible (asserted in
tests/hide-bot-chats.test.mjs). The bot row is therefore the ONLY
entry point to a bot's forever-chat, so preferring the profile's
freshest visible session did not re-order two equivalent doors — it
removed the only one. Reported symptom: a 106-message bot-building
conversation with no reachable entry point anywhere in the UI, while
the row previewed one session and opened another (a regression of the
preview/click identity #88200 established).
The report behind #91791 was real but has a non-destructive answer:
scratch sessions started via "New chat with this agent" are not
plumbing-titled, so neither hideOwnedBotSessions() nor
sweepBotProfileSessions() hides them (the sweep matches the exact
titles 'Bot Chat' / 'Agent Inbox' / 'Group: …'). They stay listed in
the Sessions sidebar and are reachable there; they simply are not what
the bot row targets, which is by design.
Changes:
- openBotCanonicalChat: when the pin is alive and verified, open it
directly. The newerVisibleBotChat preference is removed from that
branch only; the helper stays for the dead-pin recovery path.
- Drop the now-unused latestVisible parameter and its argument at the
BotRow call site. The second call site already passed three args.
- tests/bot-row-opens-latest.test.mjs ->
tests/bot-row-opens-canonical-chat.test.mjs: the two tests that
asserted the newer-session behaviour are rewritten rather than
deleted, so the reasoning survives in the suite. Adds a source-level
guard ("the healthy-pin branch never prefers a newer visible
session") so this cannot silently regress. The deleted-newer-session
fallback test covered a path that no longer exists; replaced with one
asserting a failed open of a verified pin propagates instead of
forking the forever-chat.
The keepAllProfilesScope: false half of #91791 is untouched.
Plugin suite: 392 pass, 0 fail.
The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.
1. Every pytest subprocess shared one temp root.
pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.
scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.
Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.
2. The config read guard walked directories that other tests were writing.
tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.
The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.
The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.
3. A PTY test waited for a file to exist, and not for its content.
tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.
The test now waits for the expected bytes, with a bounded deadline.
Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.
4. A dialog close timer outlived the test that started it.
ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:
ReferenceError: window is not defined
at resolveUpdatePriority (react-dom-client.development.js:1308)
at dispatchSetState
at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574
The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.
Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.
Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.
Verification:
- The affected Python files and the tests of the runner itself pass under
scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
"window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
before this change. The two cleanup effects carry an eslint-disable line for
the ref-mirror rule. They write a timer handle, and not a mirror of a
reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
no python3 outside the nix store, and the test uses the literal `python3`.
The child exits 127 there. The fix rests on the 25-run measurement above and
on CI.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
Addresses @helix4u's review on #92020:
- Consent notice now matches the real --nous contract: full logs up to
512KB each, likely conversation content/tool outputs/file paths, viewable
by Nous staff AND allowlisted Discord moderators (all 5 locales).
- Client-supplied text (error_context + extra_files) rides _redact_log_text
— the same upload-safe redactor as backend logs (secrets + email masking),
not the weaker bare secret pass; regression test covers both.
- ok:true without view_url or id becomes a structured failure; a returned
id without a link renders an upload-ID fallback the user can quote.
- Generation guard in the store: dismissal is immediate in every phase
(incl. mid-upload); a stale completion can no longer resurrect or
overwrite the dialog. Cancel button never disabled.
Chromium's native selection copy serializes the selection as text/html
with every element's computed color inlined. Copied from a dark theme,
body text lands on the clipboard as near-white (the app ink computes to
color(srgb 0.902 0.929 0.953 / 0.94)); pasted into a light-background
target such as an email, it is invisible.
The renderer never writes rich text itself, so this payload can only
come from Chromium's serializer — which runs after copy handlers decline,
meaning clipboardData reads back empty inside the event. The new guard
therefore decides from the live DOM: it scores the computed ink of the
selected text against the rendered theme mode, and only when they are
opposite schemes does it own the payload, writing text/plain plus a
tag-structured text/html with no paint declarations.
Structure (headings, lists, tables, links, bold/italic, code layout)
survives; colors come from the paste target's defaults. A generic
font-family anchor (sans-serif, monospace inside code) keeps receivers
that convert HTML to rich text on their own compose font instead of the
Times browser default. Same-scheme copies and selections starting inside
editable fields pass through untouched.
The composer middleware is now identification-only: it resolves the
user's @tags against the live roster and annotates the draft with who
they refer to (profile, friendly title, device for cross-connection
rows). The agent decides whether to contact them and does it through
its message_agent tool — one send path, composed messages only.
Deleted the renderer's entire parallel delivery transport:
deliverRemoteRosterMentions / pollRemoteDmReply /
ensureRemoteCanonicalChat and the injected shellout instructions
('[@mention handoff — run hermes -p …]' and 'Desktop is delivering …
over Connections'). This retires the whole invocation bug class at the
source instead of sanitizing it: no verbatim user text is ever
forwarded by the renderer (#91397), and no shell command is ever
composed from prompt text (#91304, #91339 shape).
Tests: mention-identification.test.mjs replaces the two delivery-era
files — identification note shape, no-shellout/no-delivery containment
(sabotage-verified: re-adding a renderer delivery call fails 2 tests),
poisoned-title inertness, pass-through for unknown @s, and a source
contract pinning the deleted machinery. hide-bots + roster-cache-key
harnesses re-pinned to the new contract. 390/390 green.
The two waitForHermesReady cloud-503 tests froze now() at 0, so the
readiness loop never crossed its deadline — the vitest electron project
hung for the full 20-minute CI budget. Advance the clock per poll like
the sibling readiness tests do.
Follow-up on the #85373 salvage: the portal and Discord URLs move out of
the localized hint prose into dedicated action buttons (URLs live in code,
translations can't drift them), matching the layered error card's
action-row idiom from #91493. Overlay test updated to the button contract;
all five locales updated.
eslint --fix output: blank lines before statements and the import-order
spacing in connection-config.test.ts that the check:lint gate rejects.
Formatting only — no logic change.
The electron boot path now classifies a Nous Cloud 502/503/504 at both the
OAuth ticket-mint and readiness boundaries and carries isCloudBackendDown /
statusCode through DesktopBootProgress, but the renderer never consumed the
structured signal — a cloud-backend failure fell into the generic remote-
failure recovery copy.
Make BootFailureOverlay branch on isCloudBackendDown: lead with the
cloud-specific title/description, drop the local-only Repair action, and
surface the actionable portal / Local-mode / Discord guidance (the electron
factory's full message is still shown in the error box).
Adds the cloudDown i18n keys (en + ar/ja/zh/zh-hant) and a regression test
asserting the cloud-down recovery renders and Repair is dropped.
The original implementation classified 502/503/504 only inside the readiness
loop, but for OAuth-backed Cloud connections the WebSocket-ticket mint runs
before waitForHermesReady. A server fault there was wrapped by
gatewayTicketFailure into a generic message and the Cloud-down classifier was
never reached. This closes that boundary and fixes a latent regex defect.
- isServerSideHttpError: structured-first (err.statusCode for 502/503/504),
legacy 'NNN:' prefix as fallback, non-Error inputs rejected. Also fixes the
committed '\d' (double-escaped, matched a literal backslash) that made the
function never detect a status prefix.
- makeNousCloudBackendDownError: single factory for the actionable Cloud-down
error (isCloudBackendDown/statusCode/detail/cause), shared by both the
ticket-mint boundary and readiness exhaustion.
- main.ts: run the Cloud classifier at mintGatewayWsTicket before the
gatewayTicketFailure wrap; 401/403 still route to reauth.
- connection-config.ts: gatewayTicketFailure preserves an integer statusCode
from the source error; auth semantics unchanged.
- boot-progress/IPC: carry isCloudBackendDown and statusCode through
DesktopBootProgress so the renderer overlay (a PR-body promise) can key on
the structured result rather than re-classifying the message string.
Tests: backend-health (structured detection, non-Error rejection, factory
shape/cause/guards, legacy fallback), connection-config (statusCode preserve,
401/403 reauth, integer-only copy), and an OAuth ticket-mint integration
regression (Cloud 503 -> actionable Cloud-down; 401 -> reauth). Connection-
config suite 80/80 green; backend-health sync tests green; the async readiness
loop tests cannot run on this host (pre-existing local-run limitation) and are
the CI gate. PR #85373 (#85335).
When a Hermes Desktop connects to a Nous-managed cloud agent
(*.agents.nousresearch.com) and that backend returns HTTP 502/503/504,
the previous error message was the opaque generic 'Hermes backend did
not become ready: 503: ...' with no guidance that the cloud server
itself is down.
Add isServerSideHttpError and isNousCloudAgentUrl helpers and use them
in waitForHermesReady to detect this exact scenario. When triggered,
throw an error with the hostname, status code, and recovery paths:
check the Nous Portal, switch to Local mode, or reach out on Discord.
Also adds a isCloudBackendDown flag and statusCode property on the
thrown error so the renderer overlay can render specialized UI if desired.
Addresses @helix4u's review on #91493:
- conversation_loop now stamps failure_retryable (the real ClassifiedError
verdict) next to failure_reason; error_surface prefers it and only falls
back to the reason set for older results. Fallback set corrected to match
classify_api_error (auth, format_error, billing_unverified now
non-retryable).
- The descriptor carries the failing session's provider/model captured at
classification time; Copy error details prefers them over the foreground
composer atoms.
- Open logs is labeled 'Open Desktop logs' on remote/cloud connections —
the local folder holds transport logs, not the remote runtime's.
- API-exception module allowlist widened to botocore/boto3/google/grpc/
requests/aiohttp so other adapter SDKs don't misclassify as gateway.
Sessions running on provider 'nous' get a 'Nous support' action on the
failed-turn card, opening the portal help hub
(https://portal.nousresearch.com/help — docs, Discord, GitHub) in the
external browser. All five locales + docs updated.
useNavigate() throws outside a <Router>; streaming.test.tsx renders the
thread bare. Move the Settings deep-link into a SwitchProviderAction child
gated on useInRouterContext(), which is safe in any tree.
Turn errors now carry a structured {layer, code, retryable} descriptor
(agent/error_surface.py) built from the same classifier the retry loop
uses. The tui_gateway stamps it on terminal error frames, retained
failed-turn snapshots, and resume replay; the Desktop error card renders
the layer title (provider / endpoint / streaming / auth / billing /
gateway / runtime / disk) plus matched actions: Retry, Switch provider,
Open logs, Copy diagnostics.
Older backends that omit the descriptor keep today's behavior (generic
title, string-sniff fallbacks) — the field is advisory on both sides.
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
The OpenCode Zen wire slug for the Ox Alpha stealth model is opaque
(x-preview-f-free); users searching the picker for 'ox' or 'ox-alpha'
found nothing. Adds the search alias across all four synced alias
tables (CLI, desktop, web, TUI) plus tests. Wire id is unchanged and
still what renders and gets sent to the provider, matching the k3 →
kimi-k3 precedent. No canonical-dedup collision with opencode-go's
keyed ox-alpha-free slug.
Clicking a bot in the roster always reopened its pinned canonical Bot Chat.
Start a new conversation with bot A, click bot B, click back to A — the new
conversation was gone, replaced by the pinned transcript. A bot row is a
workspace entry point, so it has to land on the live conversation.
Two independent causes, both fixed here:
1. The pin overrode newer work.
`openBotCanonicalChat` opened the pin unconditionally. It now prefers the
bot's freshest VISIBLE session — but only AFTER `profiles.list` has
verified through `preferred_session` that the pin is alive and is a real
canonical Bot Chat. That ordering matters: with a dead or unverified pin,
adopting the profile's latest row would claim an unrelated user
conversation as the bot's chat, and the hide sweep would then hide it.
The existing "no pin" / "dead pin" safety tests cover exactly that and
still pass. The pin keeps owning plumbing (creation, hide sweep, DM
delivery); it just stops shadowing newer conversations.
Guards on the candidate (`newerVisibleBotChat`): the canonical chat can
never shadow itself, an empty draft never displaces a real conversation,
and a gateway that omits `message_count` is treated as real history
rather than discarded.
2. The workspace did not follow the bot.
The three `host.openSession` calls on the bot path relied on the SDK
default `keepAllProfilesScope: true`, so `$activeGatewayProfile` stayed on
whatever profile was active before the click. Sessions created afterwards
were then filed under the previous bot's profile — measured: four new
chats started from three different bots all persisted into one profile's
state.db. Clicking a bot IS a profile switch, so these pass `false`.
Note on the call shape: `previewSession` is `bot.preferred_session || last`,
so on a pinned bot it resolves to the PIN (preview identity must match click
identity). Feeding that as the "newer" candidate makes the whole preference
dead code — it always sees the pin and short-circuits on "same id". The
freshest visible session therefore arrives as its own argument. The first
attempt at this fix had that bug and passed its tests, which is why
`bot-row-opens-latest.test.mjs` mirrors the production call site argument for
argument rather than constructing a convenient one.
Tests: 362 pass (was 348). Each new guard was verified by sabotage — reverting
any one of the three behaviours above makes the suite fail (1, 3, and 1 tests
respectively), so none of them is a test that passes either way.
PaneTab gated its hover close button on two independent inputs: the
onClose verb, and a showCloseButton prop that TreeGroup fed from a
showCloseButton flag on the pane contribution. The middle-click and
Meta-click gestures read only onClose. A tab could therefore close on a
pointer gesture and advertise no control for it.
The flag had no user that hideOnly did not already cover. Both setters
also set hideOnly: true, which removes every close gesture:
- the sessions pane (app/contrib/controller.tsx),
- the Bots pane (plugins/hermes-bots/plugin.js).
The flag was an opt-out marker with no reachable effect, so this change
deletes it instead of teaching it to track the gestures. onClose alone
now decides both shapes. A tab that closes shows the button. A tab
without the verb shows nothing. To make a tab uncloseable, give it no
close verb.
hideOnly and uncloseable keep their meaning. They gate the verb, and
both shapes follow the verb together.
The DialogContent and SheetContent prop of the same name is a different
prop and stays. It has no close verb to derive from, and one caller
changes it while the dialog is open.
Tests: the new tab-close-affordance test renders the real TreeGroup and
asserts that button presence equals middle-click closure. It covers
hideOnly chrome, a plain side pane, the uncloseable workspace, and a
session tile. It reads closure from the layout tree, not from a spy, so
a wired-up mock cannot pass it. A regression that hides the button on a
closeable tab fails two of the four cases. The compiler rejects the
deleted prop, so the test carries no fixture for it. The pane-tab unit
test moves off the deleted prop.
Verified with the full apps/desktop vitest suite, npm run typecheck, and
npm run lint. Two electron process-spawn tests fail on this machine.
They also fail on a clean tree, and they do not touch the pane shell.
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
The strip could only be hidden by an undiscoverable double-tap, and once hidden
the zone had no chrome left to click — no tab, no ✕, no menu holding "Show".
This puts it on the same footing as the status bar, whose hide has never
stranded anyone: ⌥⌘T, a ⌘K row, the shell context menu, and the zone menu, which
now prints the keystroke on the row that takes the strip away so the way back is
stated at the moment it matters. All four resolve their target zone the same way
the other tab verbs do (hovered, else focused, else the workspace) and describe
themselves from what is on screen rather than from a stored value, so "toggle"
always means the opposite of what the user is looking at.
Adds an app-wide default alongside it, in Appearance next to Session List
Density — auto, always, or never, matching VS Code's `workbench.editor.showTabs`
and Zed's `tab_bar.show` for people who want one answer everywhere instead of a
per-zone choice they repeat. A zone that has stated its own preference still
wins, and neither value can strand a pane.
`headerHidden` carried two meanings at once. `true` was either "the user hid
this" or "a double-tap nobody meant hid this"; `false` was either "the user
wants a strip" or "insert / tab-cycling / dock-enforce / adoption pinned one to
escape a dead end". Because the layout wrote the same field the user did, a
repair silently overwrote a preference and neither could be read back — and
since hiding also unmounted the tab, the ✕ and the menu offering "Show header",
a zone that got hidden by accident stayed that way across restarts.
Replaces it with `tabStrip?: 'always' | 'never'`, where absent is auto and only
the user ever writes it, and moves the decision into one resolver that TreeGroup
and the store both call, so the strip on screen and the toggle command cannot
disagree. Reachability moves into that resolver as an invariant that outranks an
explicit `never`: a closeable tile keeps its ✕ and a lone tool panel keeps its
chip, because "hide the chrome" is never a request to make a surface
unreachable. With that guarantee held centrally, the four repair writes are
gone. Persisted `headerHidden` is dropped rather than translated — nothing on
disk distinguishes a deliberate hide from an accidental one, and carrying the
accidents forward would re-strand exactly the people who reported being stuck.
The double-tap hide goes with it, along with the synthesized double-tap detector
it was the only consumer of. It fired from ordinary double-clicks on a tab,
nothing announced it, and its undo lived behind the chrome it had just removed.
`data-zone-no-header` goes too: it marked full-page views for a body
double-click toggle that no longer exists, and nothing has read it since.
Supersedes the tab-side half of the fix from abundantbeing and yoniebans, whose
commits this builds on.
Narrows #86278 to exactly the defect. Tabs pass no double-tap context on any press path (generic pane drag, multi-tab selection drag, chrome.tabDrag), so a double-click on a tab can no longer hide the strip; the strip background keeps its documented hide gesture unchanged.
The body double-tap reveal from #86278 is dropped: the zone body deliberately carries no double-click gesture (virtualized content recreates its nodes between clicks, per the standing ruling in tree-group.tsx), and recovery surfaces for a deliberately hidden header are being decided separately across #84458 / #81638 / #89225. The DOUBLE_TAP_MS export is reverted since no consumer remains outside drag-session.
Test file trimmed to the two assertions that pin the grammar: a tab double-tap must not hide the strip (red on main), the strip background double-tap still hides. Taps release on window between presses so the drag-session synthesized double-tap path is the one exercised.
The synthesized double-tap that hides a zone's tab strip rode every tab's
pointerdown (generic pane drag and each pane's tabDrag), so a routine
double-click on a tab (select a title, retry a click) vanished the whole
bar and stranded the zone with no tab, no close X, and no way back but a
right-click. Keep the documented hide gesture on the strip background
only, and add its inverse as recovery: double-tap a hidden zone's body
restores the strip. Regression tests pin both sides of the grammar.
The full-screen boot surfaces (connecting, onboarding, boot failure, root
crash fallback) paint their backdrop with --ui-chat-surface-background,
which the glass field turns transparent so <body> can be the one painter
(0483133842). That was harmless while glass shipped off; once it shipped
on by default (be3166607e) every boot overlay became a window onto the
shell behind it.
These overlays mask the whole app, so they declare data-glass-opaque —
the existing contract for surfaces that paint over siblings — which pins
the token back to opaque chrome under glass and changes nothing when
glass is off.
updateGroupChat's inline durable-map builder (the local-mutation persist
path) skips tombstoned rooms and carries roomId — but durableGroupChatRooms,
the SEPARATE builder persistGroupChatRooms uses for the remote-merge path
(every pullGroupChatServerState / gateway-swap sync), has neither.
Two independent gaps in the same function:
1. Tombstone resurrection. Disband sets a runtime-only tombstone
({tombstone: true, log: [], ...}) while a drive may still be mid-turn,
with no roomId. mergeRemoteGroupChatSnapshotIntoRooms spreads
...existing before its explicit field overrides (none of which touch
tombstone), so if a remote gateway hasn't received the delete yet
(plausible now that sync fans out to every reachable default-profile
gateway with independent per-connection backoff) and still has a live
copy of the room, the tombstone flag survives into the merged room.
That merged map is handed straight to persistGroupChatRooms, which
wrote it to storage because durableGroupChatRooms had no tombstone
check. On the next cold hydrate the persisted tombstone reads back as
an empty, non-tombstoned room, resurrecting the original bug
(recreating a room under the same name silently becomes "<name> 2")
through a path the earlier tombstone fix didn't cover.
2. roomId loss. mergeRemoteGroupChatSnapshotIntoRooms correctly carries
roomId into the merged in-memory room, but durableGroupChatRooms never
included it in the persisted snapshot. Every room merged in via the
remote-sync path therefore loses its immutable room identity on the
next cold hydrate (comes back with roomId: null) and falls back to
legacy name-keyed identity — breaking id-based rename/merge resolution
and member-session titling ("Group: <roomId>").
Fix: durableGroupChatRooms now mirrors updateGroupChat's inline map
exactly — skip tombstones, carry roomId.
Tests: durableGroupChatRooms unit tests for both gaps, plus an
end-to-end reachability test (tombstone -> mergeRemoteGroupChatSnapshot-
IntoRooms -> persistGroupChatRooms -> storage) proving the merge really
does forward the tombstone and the fix really does keep it out of
storage. Mutation-verified against pre-fix code (all 3 new tests fail).
Full hermes-bots plugin test suite (60+ files) green.
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the deleted stylesheet-text assertions with tests that run the
thing they claim to cover.
e2e/glyph-spinner.spec.ts drives a real browser, where the CSS actually
executes: the strip's animation resolves to steps(N) for N frames, runs
infinitely, and travels a resolved length rather than a percentage (a
percentage translate is layout-dependent and Chromium refuses to
composite it). Both pause gates are covered — the per-spinner
`data-paused` attribute and the global renderer-pause attribute that
window blur / minimize / document-hidden arm — along with the layer
promotion being scoped to running spinners. A sampling test confirms the
transform visits a bounded number of distinct values across one cycle
(steps, not a linear sweep) and that nothing mutates the DOM while it
animates, which is the property the whole change exists to deliver.
status-invalidation-scope.test.tsx pins the scoping itself as a render
count. `useTapbackDoubleClick` is called by AssistantMessageBody and by
nothing else in the tree, which makes it an exact render counter for the
message root without exporting internals. A settle and a delta flush must
both leave that count untouched while the leaves update. Verified by
mutation: reinstating a root-level status subscription fails the settle
test (2 renders where 1 is required).
It also pins node identity across the settle transition, so the
inter-agent collapse cannot go back to swapping element types at the
message-root position and remounting the row under the scroll anchor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Review follow-ups on the compositor spinner and the invalidation scoping.
Spinner CSS:
- Clip each frame to its own box. Braille renders from a system fallback
face (JetBrains Mono has no U+2800 block), whose metrics are not
guaranteed to fit the 1em frame, so neighbouring ink could bleed into
the viewport.
- Name descendants explicitly in the selection guard. The competing
`[data-selectable-text='true'] *` rule has the same (0,1,0)
specificity, so relying on inheritance made the winner depend on
stylesheet order.
- Scope the compositor promotion to spinners that are actually running.
A permanently promoted layer per parked spinner is pure memory at
fan-out breadth, where many sit mounted and paused at once.
- Give every var() the braille default as its fallback, so a missing
custom property degrades to a working spinner rather than an invalid
declaration.
Spinner component: replace the bare `as CSSProperties` cast on the inline
style with an exported GlyphSpinnerVars contract, so a typo in a custom
property name is a compile error rather than a silently dead declaration.
Assistant message:
- Render the inter-agent collapse as a CHILD of the normal body instead
of a competing root. The settled case previously returned a different
element type than the running case, so settling unmounted the whole row
and mounted a fresh one — discarding the DOM the scroll anchor held.
One component, one root, children vary; the truth table is unchanged,
including the collapsed row carrying no tapback listener.
- Collapse AssistantStatusSlot's separate subscriptions into one selector
returning a stable string. The inputs always move together on a status
flip, so reading them separately just multiplied the wake-ups.
- Give StreamingMarker a stable `data-slot` and assert on that rather
than on `span.hidden`.
Repro script: count settled rows by subtracting streaming markers from
message roots instead of `:not(:has(...))`. The selector walked every
row's subtree on each evaluation, inside the very latency window the
probe measures.
Comments: drop the stale translateY(-100%) description, name both pause
triggers, replace hard-coded line-number citations with selector/symbol
ones, note that only the primary window arms the renderer-pause
attribute, and move the forensic trace numbers out of source comments
into the PR.
Delete the three tests that asserted on stylesheet TEXT. AGENTS.md bans
reading source in tests outright, and they demonstrated exactly why: a
var()-fallback edit that changed no rendered pixel broke one of them.
Replacements that exercise the CSS in a real browser follow.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Three items from the adversarial review of the compositor-only spinner.
1. COMPOSITOR PROMOTION. The keyframes travelled translateY(-100%), which
resolves against the strip's own box and so makes the animation
layout-dependent: instrumentation recorded a non-zero compositeFailed on
184/184 records (131072 / 131104) while a sibling transform animation using an
absolute length composited clean. The travel is now
calc(frames * -1 * frame-height), an absolute length for the same distance, and
the strip gets will-change: transform. steps(var(--glyph-spinner-frames)) and
the 1em frame metric are unchanged.
The frame height is now a custom property on .glyph-spinner, used by the clip
viewport, each frame box and the keyframe travel, so those three cannot drift.
On the em-resolution question: the keyframes apply to .glyph-spinner__strip and
nothing below .glyph-spinner declares a font-size, so the strip's em and the
frame's em are the same length -- the property makes that a single declaration
rather than a coincidence to re-verify.
2. SELECTION. .glyph-spinner takes user-select: none (plus -webkit-). These sit
inside [data-selectable-text] subtrees, where the strip contributed all N
glyphs to a transcript copy against the old implementation's one. None is right
for a decorative aria-hidden element.
3. SAME-CLASS SWEEP. chat-swap-overlay.tsx ran its own 80ms setInterval +
setState braille ticker -- the exact mechanism this fix removes. Its setFrame
drove only the glyph (setLabel is independent), and its frame set and cadence
are exactly the `braille` variant, so it now renders GlyphSpinner.
`justify-start` (tailwind-merge lets the caller win) keeps the glyph
left-aligned in its w-3 box as the bare span was. GlyphSpinner gains a `paused`
prop for it: the overlay stays mounted through its fade-out, and the old
cleared-interval behaviour was to stop animating once the swap was done.
Tests: guards for the two invisible-in-jsdom regressions -- keyframes must not
return to a percentage translate, and the selection guard must stay -- plus the
paused prop, and a new chat-swap-overlay.test.tsx pinning that no timer comes
back, the label still survives the fade-out, and the glyph freezes with it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Scheduler attribution on the incident trace (FINDINGS.md round-5-Opus receipt)
puts 133 of 138 wide document-scale recalcs on GlyphSpinner's ticker, and 0 of
254 cheap ones. Replacing glyph.textContent every interval is a structural
text-node mutation, so each tick scheduled a style recalculation that resolved
against the whole document -- with N spinners mounted in a streaming
transcript, that is the incident.
Every frame is now in the DOM from mount as a vertical strip, scrolled by a
transform translateY keyframes animation. Transform animations run on the
compositor: no JS timer, no text mutation, no per-frame style recalc, layout or
schedule.
The strip is N frames tall and each frame is exactly 1em, so translating -100%
travels N frames; steps(N) (jump-end) samples that at 0, 1/N .. (N-1)/N, i.e.
it parks on frame 0..N-1 for one interval each and wraps -- the same sequence
and cadence the setInterval produced. Frame count and duration (N x interval)
arrive as inline custom properties, so all 18 spinner names / 16 distinct
frame-interval shapes share one keyframes rule with no generated or colliding
per-variant CSS.
Sizing, colour and alignment are unchanged: the outer cell keeps its exact
classes, and the 1em clipping viewport is centred by the same items-center that
used to centre the single glyph -- so consumer classNames that set a box
(size-3, size-3.5) or a font-size still land the way they did.
Gating semantics preserved, per the original "N mounted tabs each ticking burns
CPU for pixels nobody can see":
- kept-alive hidden tab -> data-paused -> animation-play-state: paused. Kept
explicit rather than relying on the pane's content-visibility:hidden, since
that containment has a runtime kill switch and older pane layers only set
visibility:hidden, which does not stop an animation.
- window blur / minimize / document hidden -> the strip joins the existing
:root[data-renderer-animations-paused] allowlist in styles.css, driven by
main.tsx's installRendererAnimationPauseState(). That is the mechanism every
other continuous decorative animation here already uses, so the per-spinner
createRendererLoopPauseController goes away.
- reduced motion is now honoured, which the ticker never did: the blanket
@media (prefers-reduced-motion: reduce) rule freezes the animation. This
also makes E2E screenshots deterministic, which that rule exists for.
The frames are marked aria-hidden. role="status" is a live region and the old
implementation rewrote its text ~12x/second, which announced a new glyph on
every tick.
Tests rewritten, deliberately: all five previous cases asserted the ticker
itself (vi.getTimerCount(), per-tick textContent), which no longer exists. The
replacements pin what jsdom can see -- frame order, the custom properties
feeding steps()/duration per variant, zero timers ever created, the hidden-tab
gate, aria-hidden, and that the strip is still named in the global pause rule.
The blur/minimize/document-hidden contracts now belong to that global mechanism
and are covered by lib/renderer-loop-pause.test.ts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The last status-dependent read at the message root, and the most expensive
one: the completedText selector flipped between '' while running and a full
messageContentText(content) join once settled, so every running <-> settled
transition re-ran the join for the whole message AND re-rendered the root. At
stream breadth N that is N joins plus N root re-renders per flip.
completedText and the previewTargets memo it feeds now live in a new
AssistantPreviewEmbeds leaf, mounted at the same position inside
[data-slot='aui_assistant-message-content']. Verified before moving that
previewTargets fed nothing else at the root -- its only consumer was its own
render block. The leaf renders the same wrapper div with the same classes, or
null when there are no targets, so the DOM is byte-identical; a component
boundary adds no node, so unlike StreamingMarker this needed no placement care
around the :first-child/:last-child rules.
The '' branch is preserved deliberately: it is the streaming-side optimization
that keeps the selector referentially stable so per-token flushes skip the
regex scan.
AssistantMessageBody now holds no status-dependent subscription at all -- what
remains is messageId, hasVisibleText, isInterim and turnDurationS, none of
which move on a pending flip.
Adds preview-embeds.test.tsx. The embed had no coverage, and the two cases are
written as a matched pair on the same selector -- present once settled, absent
while running -- so neither can pass vacuously.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Removes the two residual invalidators left by the previous commit.
data-streaming: gone from the message root. The flag is not dead -- it is the
settled-row signal for scripts/run-short-session-hang-repro.mjs -- so it moved
to a permanently-mounted, display:none leaf that is a ROOT-LEVEL sibling, and
the repro now matches on the descendant. Placement is load-bearing three ways:
a node inside [data-slot='aui_assistant-message-content'] would steal
:last-child from the stall indicator and change inter-bubble margins mid-stream
(styles.css:1995-2003); keeping it mounted and toggling only the attribute
keeps the per-flip write on a childless node instead of making it a DOM
structure change; display:none costs no layout or paint while querySelectorAll
and :has() still match it.
Renamed to data-message-streaming rather than reusing data-streaming: shiki
puts that exact attribute on deferred code cards, which are descendants of the
message root, so a descendant-matching selector sharing the name would report
any message holding a deferred code card as still streaming.
root isRunning: gone from the standard path. AssistantMessage now dispatches on
interAgentSender, so the collapse gate's live status subscription lives in
InterAgentAssistantMessage and only the rare inter-agent case pays it. The
enter animation captures its enabled flag once off the runtime, non-reactively,
because use-enter-animation.ts parks the value in a ref behind a useCallback([])
identity and consults it only when the callback ref fires at mount -- a live
subscription fed a value the hook already ignores.
Adds inter-agent-collapse.test.tsx: the collapse gate and the marker contract
both had zero coverage, and nothing in the app reads the marker, so a delete
would otherwise look free and silently regress the repro's response gate.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Prong A of the wide-recalc fix (FINDINGS.md round-4 protocol): read status in
leaf components inside message content, hoist MessagePrimitive.Parts so status
flips cannot re-render the parts subtree. Behavior-identical; invalidation
scope only.
The third primary edit -- dropping the data-streaming root attribute -- is NOT
in this commit. The design doc calls it dead based on a CSS grep, and that grep
is correct (every [data-streaming='true'] rule targets [data-slot='code-card']).
But it has a live non-CSS consumer: scripts/run-short-session-hang-repro.mjs
:928 and :1023 count settled assistant rows via
[data-slot="aui_assistant-message-root"]:not([data-streaming="true"]) and gate
the assistant-response wait on that count growing. Deleting the attribute makes
that selector match every row, so the gate would pass at stream start instead of
completion. Held pending a decision rather than silently weakening the repro.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test harnesses that vi.mock('@/hermes') without setApiRequestProfile make
the session-states transitive graph unloadable; the deferred reconcile
import then rejected unhandled and failed unrelated suites in shard 2.
Catch and skip — the production graph always loads.
Static gateway.ts -> session-states.ts import closed a module cycle that
left $activeGatewayProfile undefined at session-states init (TypeError:
Cannot read properties of undefined (reading 'get')) — the CI red across
all three UI shards. Dynamic import defers the edge past module init;
reconcile semantics unchanged. Proven by re-adding the static edge:
hud/pet suites reproduce the exact CI failure.
A respawned backend re-mints runtime ids, so a pre-reconnect busy state
never receives its terminal busy:false publish and its session stayed in
$workingSessionIds forever - the sidebar running arc and agents-panel
'running' chrome lied for hours after the turn ended (the stale-flag
half of #53902/#73082; the CSS cost half landed in #91383).
reconcileBusyStatesOnReconnect() downgrades busy/awaitingResponse states
through publishSessionState (watchdogs disarm, stall hints drop, settle/
unread bookkeeping stays consistent), scoped by event-source: the primary
reconnect touches only scope-less runtimes, a secondary (registry)
reconnect touches only its own connection's. needsInput survives - a
blocking prompt is the user's to answer. A genuinely live turn re-asserts
busy on its next post-reconnect event, so the worst case is one arc blink.
Regression tests proven by sabotage run (neutered reconcile -> 5/6 fail).
The arc-border running indicator animated background-position (repaint
every frame: ~3,600 main-thread style recalcs/min per arc, ~4,100ms/min
of renderer task time measured over 60s of true idle) and progress-slide
animated left (forced layout every frame). Both now travel via transform
on the compositor: same visuals, 3,617 -> 79 style recalcs/min and
4,108 -> 153 ms/min task time in the same harness (-96%).
Follow-ups to the salvaged WSL-bridge gating (#66447):
- wsl-path-bridge.ts: discard wsl.exe stderr so the 'WSL is not installed'
banner can never leak into an attached console on WSL-less machines (#80184).
- scripts/desktop-update/windows.ps1: add an explorer.exe-mediated detached
relaunch rung between the WMI attempt and the tethered Start-Process
fallback. When Win32_Process.Create fails (observed ReturnValue 8), the
Desktop no longer re-attaches to the hand-off console, so its stdout stops
flooding the window and the console can close.