updateGroupChat's inline durable-map builder (the local-mutation persist
path) skips tombstoned rooms and carries roomId — but durableGroupChatRooms,
the SEPARATE builder persistGroupChatRooms uses for the remote-merge path
(every pullGroupChatServerState / gateway-swap sync), has neither.
Two independent gaps in the same function:
1. Tombstone resurrection. Disband sets a runtime-only tombstone
({tombstone: true, log: [], ...}) while a drive may still be mid-turn,
with no roomId. mergeRemoteGroupChatSnapshotIntoRooms spreads
...existing before its explicit field overrides (none of which touch
tombstone), so if a remote gateway hasn't received the delete yet
(plausible now that sync fans out to every reachable default-profile
gateway with independent per-connection backoff) and still has a live
copy of the room, the tombstone flag survives into the merged room.
That merged map is handed straight to persistGroupChatRooms, which
wrote it to storage because durableGroupChatRooms had no tombstone
check. On the next cold hydrate the persisted tombstone reads back as
an empty, non-tombstoned room, resurrecting the original bug
(recreating a room under the same name silently becomes "<name> 2")
through a path the earlier tombstone fix didn't cover.
2. roomId loss. mergeRemoteGroupChatSnapshotIntoRooms correctly carries
roomId into the merged in-memory room, but durableGroupChatRooms never
included it in the persisted snapshot. Every room merged in via the
remote-sync path therefore loses its immutable room identity on the
next cold hydrate (comes back with roomId: null) and falls back to
legacy name-keyed identity — breaking id-based rename/merge resolution
and member-session titling ("Group: <roomId>").
Fix: durableGroupChatRooms now mirrors updateGroupChat's inline map
exactly — skip tombstones, carry roomId.
Tests: durableGroupChatRooms unit tests for both gaps, plus an
end-to-end reachability test (tombstone -> mergeRemoteGroupChatSnapshot-
IntoRooms -> persistGroupChatRooms -> storage) proving the merge really
does forward the tombstone and the fix really does keep it out of
storage. Mutation-verified against pre-fix code (all 3 new tests fail).
Full hermes-bots plugin test suite (60+ files) green.
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the deleted stylesheet-text assertions with tests that run the
thing they claim to cover.
e2e/glyph-spinner.spec.ts drives a real browser, where the CSS actually
executes: the strip's animation resolves to steps(N) for N frames, runs
infinitely, and travels a resolved length rather than a percentage (a
percentage translate is layout-dependent and Chromium refuses to
composite it). Both pause gates are covered — the per-spinner
`data-paused` attribute and the global renderer-pause attribute that
window blur / minimize / document-hidden arm — along with the layer
promotion being scoped to running spinners. A sampling test confirms the
transform visits a bounded number of distinct values across one cycle
(steps, not a linear sweep) and that nothing mutates the DOM while it
animates, which is the property the whole change exists to deliver.
status-invalidation-scope.test.tsx pins the scoping itself as a render
count. `useTapbackDoubleClick` is called by AssistantMessageBody and by
nothing else in the tree, which makes it an exact render counter for the
message root without exporting internals. A settle and a delta flush must
both leave that count untouched while the leaves update. Verified by
mutation: reinstating a root-level status subscription fails the settle
test (2 renders where 1 is required).
It also pins node identity across the settle transition, so the
inter-agent collapse cannot go back to swapping element types at the
message-root position and remounting the row under the scroll anchor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Review follow-ups on the compositor spinner and the invalidation scoping.
Spinner CSS:
- Clip each frame to its own box. Braille renders from a system fallback
face (JetBrains Mono has no U+2800 block), whose metrics are not
guaranteed to fit the 1em frame, so neighbouring ink could bleed into
the viewport.
- Name descendants explicitly in the selection guard. The competing
`[data-selectable-text='true'] *` rule has the same (0,1,0)
specificity, so relying on inheritance made the winner depend on
stylesheet order.
- Scope the compositor promotion to spinners that are actually running.
A permanently promoted layer per parked spinner is pure memory at
fan-out breadth, where many sit mounted and paused at once.
- Give every var() the braille default as its fallback, so a missing
custom property degrades to a working spinner rather than an invalid
declaration.
Spinner component: replace the bare `as CSSProperties` cast on the inline
style with an exported GlyphSpinnerVars contract, so a typo in a custom
property name is a compile error rather than a silently dead declaration.
Assistant message:
- Render the inter-agent collapse as a CHILD of the normal body instead
of a competing root. The settled case previously returned a different
element type than the running case, so settling unmounted the whole row
and mounted a fresh one — discarding the DOM the scroll anchor held.
One component, one root, children vary; the truth table is unchanged,
including the collapsed row carrying no tapback listener.
- Collapse AssistantStatusSlot's separate subscriptions into one selector
returning a stable string. The inputs always move together on a status
flip, so reading them separately just multiplied the wake-ups.
- Give StreamingMarker a stable `data-slot` and assert on that rather
than on `span.hidden`.
Repro script: count settled rows by subtracting streaming markers from
message roots instead of `:not(:has(...))`. The selector walked every
row's subtree on each evaluation, inside the very latency window the
probe measures.
Comments: drop the stale translateY(-100%) description, name both pause
triggers, replace hard-coded line-number citations with selector/symbol
ones, note that only the primary window arms the renderer-pause
attribute, and move the forensic trace numbers out of source comments
into the PR.
Delete the three tests that asserted on stylesheet TEXT. AGENTS.md bans
reading source in tests outright, and they demonstrated exactly why: a
var()-fallback edit that changed no rendered pixel broke one of them.
Replacements that exercise the CSS in a real browser follow.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Three items from the adversarial review of the compositor-only spinner.
1. COMPOSITOR PROMOTION. The keyframes travelled translateY(-100%), which
resolves against the strip's own box and so makes the animation
layout-dependent: instrumentation recorded a non-zero compositeFailed on
184/184 records (131072 / 131104) while a sibling transform animation using an
absolute length composited clean. The travel is now
calc(frames * -1 * frame-height), an absolute length for the same distance, and
the strip gets will-change: transform. steps(var(--glyph-spinner-frames)) and
the 1em frame metric are unchanged.
The frame height is now a custom property on .glyph-spinner, used by the clip
viewport, each frame box and the keyframe travel, so those three cannot drift.
On the em-resolution question: the keyframes apply to .glyph-spinner__strip and
nothing below .glyph-spinner declares a font-size, so the strip's em and the
frame's em are the same length -- the property makes that a single declaration
rather than a coincidence to re-verify.
2. SELECTION. .glyph-spinner takes user-select: none (plus -webkit-). These sit
inside [data-selectable-text] subtrees, where the strip contributed all N
glyphs to a transcript copy against the old implementation's one. None is right
for a decorative aria-hidden element.
3. SAME-CLASS SWEEP. chat-swap-overlay.tsx ran its own 80ms setInterval +
setState braille ticker -- the exact mechanism this fix removes. Its setFrame
drove only the glyph (setLabel is independent), and its frame set and cadence
are exactly the `braille` variant, so it now renders GlyphSpinner.
`justify-start` (tailwind-merge lets the caller win) keeps the glyph
left-aligned in its w-3 box as the bare span was. GlyphSpinner gains a `paused`
prop for it: the overlay stays mounted through its fade-out, and the old
cleared-interval behaviour was to stop animating once the swap was done.
Tests: guards for the two invisible-in-jsdom regressions -- keyframes must not
return to a percentage translate, and the selection guard must stay -- plus the
paused prop, and a new chat-swap-overlay.test.tsx pinning that no timer comes
back, the label still survives the fade-out, and the glyph freezes with it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Scheduler attribution on the incident trace (FINDINGS.md round-5-Opus receipt)
puts 133 of 138 wide document-scale recalcs on GlyphSpinner's ticker, and 0 of
254 cheap ones. Replacing glyph.textContent every interval is a structural
text-node mutation, so each tick scheduled a style recalculation that resolved
against the whole document -- with N spinners mounted in a streaming
transcript, that is the incident.
Every frame is now in the DOM from mount as a vertical strip, scrolled by a
transform translateY keyframes animation. Transform animations run on the
compositor: no JS timer, no text mutation, no per-frame style recalc, layout or
schedule.
The strip is N frames tall and each frame is exactly 1em, so translating -100%
travels N frames; steps(N) (jump-end) samples that at 0, 1/N .. (N-1)/N, i.e.
it parks on frame 0..N-1 for one interval each and wraps -- the same sequence
and cadence the setInterval produced. Frame count and duration (N x interval)
arrive as inline custom properties, so all 18 spinner names / 16 distinct
frame-interval shapes share one keyframes rule with no generated or colliding
per-variant CSS.
Sizing, colour and alignment are unchanged: the outer cell keeps its exact
classes, and the 1em clipping viewport is centred by the same items-center that
used to centre the single glyph -- so consumer classNames that set a box
(size-3, size-3.5) or a font-size still land the way they did.
Gating semantics preserved, per the original "N mounted tabs each ticking burns
CPU for pixels nobody can see":
- kept-alive hidden tab -> data-paused -> animation-play-state: paused. Kept
explicit rather than relying on the pane's content-visibility:hidden, since
that containment has a runtime kill switch and older pane layers only set
visibility:hidden, which does not stop an animation.
- window blur / minimize / document hidden -> the strip joins the existing
:root[data-renderer-animations-paused] allowlist in styles.css, driven by
main.tsx's installRendererAnimationPauseState(). That is the mechanism every
other continuous decorative animation here already uses, so the per-spinner
createRendererLoopPauseController goes away.
- reduced motion is now honoured, which the ticker never did: the blanket
@media (prefers-reduced-motion: reduce) rule freezes the animation. This
also makes E2E screenshots deterministic, which that rule exists for.
The frames are marked aria-hidden. role="status" is a live region and the old
implementation rewrote its text ~12x/second, which announced a new glyph on
every tick.
Tests rewritten, deliberately: all five previous cases asserted the ticker
itself (vi.getTimerCount(), per-tick textContent), which no longer exists. The
replacements pin what jsdom can see -- frame order, the custom properties
feeding steps()/duration per variant, zero timers ever created, the hidden-tab
gate, aria-hidden, and that the strip is still named in the global pause rule.
The blur/minimize/document-hidden contracts now belong to that global mechanism
and are covered by lib/renderer-loop-pause.test.ts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The last status-dependent read at the message root, and the most expensive
one: the completedText selector flipped between '' while running and a full
messageContentText(content) join once settled, so every running <-> settled
transition re-ran the join for the whole message AND re-rendered the root. At
stream breadth N that is N joins plus N root re-renders per flip.
completedText and the previewTargets memo it feeds now live in a new
AssistantPreviewEmbeds leaf, mounted at the same position inside
[data-slot='aui_assistant-message-content']. Verified before moving that
previewTargets fed nothing else at the root -- its only consumer was its own
render block. The leaf renders the same wrapper div with the same classes, or
null when there are no targets, so the DOM is byte-identical; a component
boundary adds no node, so unlike StreamingMarker this needed no placement care
around the :first-child/:last-child rules.
The '' branch is preserved deliberately: it is the streaming-side optimization
that keeps the selector referentially stable so per-token flushes skip the
regex scan.
AssistantMessageBody now holds no status-dependent subscription at all -- what
remains is messageId, hasVisibleText, isInterim and turnDurationS, none of
which move on a pending flip.
Adds preview-embeds.test.tsx. The embed had no coverage, and the two cases are
written as a matched pair on the same selector -- present once settled, absent
while running -- so neither can pass vacuously.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Removes the two residual invalidators left by the previous commit.
data-streaming: gone from the message root. The flag is not dead -- it is the
settled-row signal for scripts/run-short-session-hang-repro.mjs -- so it moved
to a permanently-mounted, display:none leaf that is a ROOT-LEVEL sibling, and
the repro now matches on the descendant. Placement is load-bearing three ways:
a node inside [data-slot='aui_assistant-message-content'] would steal
:last-child from the stall indicator and change inter-bubble margins mid-stream
(styles.css:1995-2003); keeping it mounted and toggling only the attribute
keeps the per-flip write on a childless node instead of making it a DOM
structure change; display:none costs no layout or paint while querySelectorAll
and :has() still match it.
Renamed to data-message-streaming rather than reusing data-streaming: shiki
puts that exact attribute on deferred code cards, which are descendants of the
message root, so a descendant-matching selector sharing the name would report
any message holding a deferred code card as still streaming.
root isRunning: gone from the standard path. AssistantMessage now dispatches on
interAgentSender, so the collapse gate's live status subscription lives in
InterAgentAssistantMessage and only the rare inter-agent case pays it. The
enter animation captures its enabled flag once off the runtime, non-reactively,
because use-enter-animation.ts parks the value in a ref behind a useCallback([])
identity and consults it only when the callback ref fires at mount -- a live
subscription fed a value the hook already ignores.
Adds inter-agent-collapse.test.tsx: the collapse gate and the marker contract
both had zero coverage, and nothing in the app reads the marker, so a delete
would otherwise look free and silently regress the repro's response gate.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Prong A of the wide-recalc fix (FINDINGS.md round-4 protocol): read status in
leaf components inside message content, hoist MessagePrimitive.Parts so status
flips cannot re-render the parts subtree. Behavior-identical; invalidation
scope only.
The third primary edit -- dropping the data-streaming root attribute -- is NOT
in this commit. The design doc calls it dead based on a CSS grep, and that grep
is correct (every [data-streaming='true'] rule targets [data-slot='code-card']).
But it has a live non-CSS consumer: scripts/run-short-session-hang-repro.mjs
:928 and :1023 count settled assistant rows via
[data-slot="aui_assistant-message-root"]:not([data-streaming="true"]) and gate
the assistant-response wait on that count growing. Deleting the attribute makes
that selector match every row, so the gate would pass at stream start instead of
completion. Held pending a decision rather than silently weakening the repro.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test harnesses that vi.mock('@/hermes') without setApiRequestProfile make
the session-states transitive graph unloadable; the deferred reconcile
import then rejected unhandled and failed unrelated suites in shard 2.
Catch and skip — the production graph always loads.
Static gateway.ts -> session-states.ts import closed a module cycle that
left $activeGatewayProfile undefined at session-states init (TypeError:
Cannot read properties of undefined (reading 'get')) — the CI red across
all three UI shards. Dynamic import defers the edge past module init;
reconcile semantics unchanged. Proven by re-adding the static edge:
hud/pet suites reproduce the exact CI failure.
A respawned backend re-mints runtime ids, so a pre-reconnect busy state
never receives its terminal busy:false publish and its session stayed in
$workingSessionIds forever - the sidebar running arc and agents-panel
'running' chrome lied for hours after the turn ended (the stale-flag
half of #53902/#73082; the CSS cost half landed in #91383).
reconcileBusyStatesOnReconnect() downgrades busy/awaitingResponse states
through publishSessionState (watchdogs disarm, stall hints drop, settle/
unread bookkeeping stays consistent), scoped by event-source: the primary
reconnect touches only scope-less runtimes, a secondary (registry)
reconnect touches only its own connection's. needsInput survives - a
blocking prompt is the user's to answer. A genuinely live turn re-asserts
busy on its next post-reconnect event, so the worst case is one arc blink.
Regression tests proven by sabotage run (neutered reconcile -> 5/6 fail).
The arc-border running indicator animated background-position (repaint
every frame: ~3,600 main-thread style recalcs/min per arc, ~4,100ms/min
of renderer task time measured over 60s of true idle) and progress-slide
animated left (forced layout every frame). Both now travel via transform
on the compositor: same visuals, 3,617 -> 79 style recalcs/min and
4,108 -> 153 ms/min task time in the same harness (-96%).
Follow-ups to the salvaged WSL-bridge gating (#66447):
- wsl-path-bridge.ts: discard wsl.exe stderr so the 'WSL is not installed'
banner can never leak into an attached console on WSL-less machines (#80184).
- scripts/desktop-update/windows.ps1: add an explorer.exe-mediated detached
relaunch rung between the WMI attempt and the tethered Start-Process
fallback. When Win32_Process.Create fails (observed ReturnValue 8), the
Desktop no longer re-attaches to the hand-off console, so its stdout stops
flooding the window and the console can close.
Seed backend-mode state before creating the first window and update it after runtime resolution. Keep remote reconnects gated until a local backend is confirmed.
Cover Windows child-process suppression and bridge state transitions.
The Pinned section falls back to the server `pinned` flag for rows the
local set doesn't hold, so a backend pin stays reachable when
localStorage is cold (#85969). But an unpin leaves the local set the
instant the user clicks, while the loaded row keeps reporting
`pinned: true` until a page issued after the PATCH lands — so the
fallback read the user's own unpin as a foreign pin, parked the session
at the bottom of Pinned, and only released it a refresh cycle later.
session-pin-sync already knows which rows its own in-flight writes
contradict; publish that fence and have the fallback skip them.
* fix(gateway): persist prompt.submit truncation to the session's own profile DB
`_get_db()` returns the LAUNCH profile's SessionDB handle. App-global
remote mode gives a session its own profile (`session["profile_home"]`)
whose transcript lives in that profile's `state.db`, so a write keyed on
`session_key` that goes through `_get_db()` addresses the wrong database.
In the `prompt.submit` truncate branch that has two consequences. The
edit/resend never sticks — `session.resume` reopens the profile db and
resurrects the undone turns — and when the launch profile happens to hold
a row under the same session id, the truncated transcript is inserted
into a profile the session does not belong to.
It also silently voids the branch's own fail-closed contract. The handler
persists before it rewrites `session["history"]` precisely so that a
failed write refuses the turn and leaves memory and DB aligned; that only
holds if the handle it checks is the one that owns the row.
`_session_db(session)` is the profile-aware resolver that already exists
for this: the profile's `state.db` when `profile_home` is set, otherwise
the shared launch handle. Non-profile sessions are unaffected —
`_session_db` borrows the same shared handle and leaves it open.
`active_only=True` and `archive_dropped=True` are carried through
unchanged; only the handle the call is made against changes.
* fix(gateway): resolve the /undo command against the session's own profile DB
`command.dispatch`'s `/undo` branch opened the launch profile's handle via
`_get_db()`, but every read and write under it is scoped by session id:
`list_recent_user_messages`, `rewind_to_message` and the
`get_messages_as_conversation` reload all key on `session_key`.
For a session with its own profile (`session["profile_home"]`) the rows
live in that profile's `state.db`, so against the launch handle
`list_recent_user_messages` returns nothing and the command fails closed
with `4018 "no user messages to undo"` — for the entire session, on every
invocation, even though the transcript is right there in the profile db.
Route the whole branch through `_session_db(session)`, which yields the db
that owns the session's row and closes a profile handle on exit. Sessions
without a profile keep borrowing the shared launch handle exactly as
before, so this is behaviourally identical for them.
* fix(gateway): read /history and /context from the session's own profile DB
`_format_live_history_output` and `_format_live_context_output` rebuild the
transcript from the database rather than from `session["history"]`, because
the in-memory list is empty for a session this process did not run itself.
Both reads are scoped by session id but were issued against `_get_db()`,
the launch profile's handle.
A session with its own profile (`session["profile_home"]`) keeps its rows
in that profile's `state.db`, so both reads come back empty and the
commands under-report: `/history` renders "No conversation history yet."
and `/context` falls back to the empty in-memory list and reports a
conversation of zero messages. Both swallow their exceptions, so there is
no error either — just a wrong answer about the user's own transcript.
Resolve both through `_session_db(session)`, the profile-aware resolver
used by the rest of the session-scoped paths.
* test(gateway): cover session-scoped transcript ops against a profile DB
Regression coverage for the three session-scoped sites that resolved
against the launch profile's handle instead of the db owning the
session's row. Each test drives the real JSON-RPC entry point with a
session carrying `profile_home`, seeds the transcript into the profile's
own `state.db`, and asserts against both databases.
Per site, with the production change reverted to its pre-fix form:
- `prompt.submit` truncation — `test_truncation_persists_to_the_profile_db`
and `test_truncation_does_not_copy_rows_into_the_launch_profile` fail.
The second seeds a row under the same session id in the launch db so the
foreign write succeeds instead of failing a key check, which is the case
that copies a transcript into a profile it does not belong to.
- `/undo` — `test_undo_rewinds_the_profile_transcript` fails with
`4018 "no user messages to undo"`.
- `/history` and `/context` — `test_history_reads_the_profile_transcript`
and `test_context_reads_the_profile_transcript` fail, reporting an empty
conversation.
`test_undo_still_uses_the_shared_handle_without_a_profile` and
`test_truncation_without_a_profile_uses_the_shared_handle` pin the
unchanged path: with no `profile_home` the resolver must borrow the shared
launch handle and leave it open. Both stay green in every direction, so a
future change cannot satisfy the profile cases by abandoning the shared
one.
* fix(desktop): aim truncations by durable id alone on tail-only transcripts
The cold-open transcript is a newest-first prefetch page
(LATEST_SESSION_MESSAGES_LIMIT = 120) with the resume RPC sent
omit_messages — older rows only arrive via "Show earlier" backfill.
planEdit/planReload/planRestore still counted truncate ordinals over
that windowed list, so every edit/reload/restore in a session longer
than the prefetch page sent a window-relative ordinal alongside the
durable row/message id. The gateway's #82959 cross-check resolved the
durable id to its full-history ordinal, read the offset as drift, and
refused with 4030 — making the Edit affordance permanently dead in
long sessions.
When the transcript may be tail-only (the transcript-tail
bookkeeping's possiblyTruncated), drop the client ordinal and address
the truncation by durable id alone — the same rule runRewindSubmit
already applies to content-resolved row ids (#87059). The ordinal
tripwire stays on whenever the transcript is complete.
Closes#88082
* fix(desktop): drop client rewind ordinal whenever a durable id is present
#88092 gated the drop on tail-only prefetch. After in-place compact the
live scrollback is treated as complete, so Restore still sent a
display-lineage ordinal next to a resolved row id and the gateway
refused with 4030 (#89244). prefix_user_count is structurally 0 on
in-place because get_ancestor_display_prefix is cross-session.
Same choke point: if a durable truncate_before_row_id or a real
truncate_before_message_id is present, omit the client ordinal.
confirm_empty_truncate is still carried from a caller ordinal of 0.
Unknown ids still fail closed at 4018.
Closes#89244
---------
Co-authored-by: briandevans <252620095+briandevans@users.noreply.github.com>
Co-authored-by: zengzheqing <yuntianqing@yahoo.com>
* fix(desktop): stop layering Windows glass windows, and don't make them transparent
A Windows chat window under glass rendered while focused and went dead the
moment it lost it. Two things put it on a compositing path DWM will not draw
acrylic behind, both no-ops that looked free:
`opacity: windowOpacity()` was passed on every window. Under glass on Windows
fade is 0, so the value is always 1 — but Electron's `SetOpacity` calls
`SetLayered()` and `SetLayeredWindowAttributes(..., LWA_ALPHA)` before it looks
at the value, and nothing ever takes `WS_EX_LAYERED` back off. A layered window
composites through the legacy redirection surface, which Windows documents as
mutually exclusive with `UpdateLayeredWindow`. Opacity is now only passed when
the state actually fades, and the runtime path keeps setting it for a window
that is already faded so it can still come back to opaque.
`transparent: true` was set on every glass-capable Windows chat window, on the
premise that DWM materials only reach the client area that way (electron#49443,
which was closed as need-info against an EOL Electron 28). They do not need it:
`IsTranslucent` answers yes off `background_material_` alone, which is what
gives the page its transparent default backing, and `SetBackgroundMaterial`
flips widget translucency live. Its one gate is a frameless window, and
`titleBarStyle: 'hidden'` already satisfies it. What `transparent` did add was
permanent — the widget pinned to kTranslucent for the window's whole life, so
even glass-OFF windows paid a DirectComposition redraw per frame
(electron#39895), plus the documented transparent-window limits, including that
a resizable transparent window is unsupported and breaks (electron#48421).
Both landed latent in #89837 and only surfaced when #90587 turned glass on by
default and dropped the opaque backing that had been hiding them.
* docs(desktop): name the one thing the opacity guard cannot undo
Electron exposes no way back off WS_EX_LAYERED, so a Windows window that has
been faded once keeps the layered compositing path until it is recreated. Not
opening the door on the default path is the whole of the fix; say so where the
guard lives rather than leaving a reviewer to work out the gap.
Resolves the room-lifecycle class on top of the salvaged #89369 projection:
- v3 projection keys rooms by immutable roomId (id:<roomId>) with
name:<name> fallback for legacy rooms; v1/v2 envelopes are normalized
on read so mixed-version fleets share one merge path
- rename is now a same-key field update — no distributed delete+create,
no old-name resurrection from lagging gateways
- id tombstones are FINAL (ids are never reused), so a gateway that was
offline during a disband can never resurrect the room, regardless of
the revision its stale copy carries; same-name recreation is unaffected
because it mints a fresh roomId
- the projection fans out to EVERY reachable default-profile gateway
(per-gateway job queues, CAS revision streams, backoff and retry caps),
so rooms survive any single gateway dying and surface on gateway-only
clients without waiting for a Desktop to foreground that gateway
- cold hydrate follows a remote rename via roomId instead of duplicating
the room under both names
New tests: id-keyed rename continuity, final id-tombstones vs lagging
high-revision copies, rename-job shape (changed+deleted same key),
cold-hydrate re-keying, multi-gateway fan-out. Sabotage-verified: each
new test fails against the pre-class behavior.
Two follow-ups to the salvaged #89369 base against current main:
- compact projection entries no longer overwrite the local rich copy
(attachments survive; watermark accounting stays stable)
- synthetic legacy-N thread ids collapse to one bucket in the entry key,
so id-less entries don't duplicate after a pull and manufacture
phantom member turns into busy sessions
Wire babel-plugin-react-compiler through @vitejs/plugin-react v6's
reactCompilerPreset + @rolldown/plugin-babel in the web and desktop
vite configs, scoped to modules that can actually contain components
or hooks (JSX syntax or a react-ish import — the preset's default
filter babel-parsed every TS module). Both vitest configs run compiled
components, so rules-of-react violations fail in CI.
Also fixes the latent bug the compiler exposed: usePluginI18n kept a
stable translator identity over a mutating locale registry, so
memoized consumers (React.memo today, compiled components tomorrow)
served stale strings after a late bundle registration. The registry
version now keys the translator identity — correct with or without
the compiler.
useRoster() keys its query on [...ROSTER_KEY, connectionId] so every
connection the window has been on gets its own cache entry, but the two
imperative roster readers — the @mention completions and the mention
middleware — still called getQueryData(ROSTER_KEY). TanStack Query's
getQueryData is an exact-key match, so the bare key matched nothing and
both readers saw an empty roster: the composer offered no agent handles,
and remote @name-device mentions (e.g. @default-vera for a same-named
profile on an SSH connection) fell through to the local profiles.list
fallback, which only knows bare local names — the mention passed through
untouched to the local agent instead of being routed over Connections.
Add a cachedRoster() helper: honor the legacy bare-key write first, then
scan the key family with getQueriesData, preferring the active
connection's entry and falling back to any other entry (a stale roster
beats none). Never throws, matching the readers' existing posture.
Regression tests model the real cache shape — entries under the suffixed
key, exact-key reads missing them — so the key-shape mismatch is pinned;
the previous tests stubbed getQueryData to ignore keys, which is exactly
why this slipped through.
useRoster caches under [...ROSTER_KEY, activeConnectionId] — a 3-element
key — but the composer @ autocomplete provider and the mention-routing
middleware both read getQueryData(ROSTER_KEY) with the bare 2-element
key. An exact-match lookup against a 3-element cache key returns
undefined forever, so:
- @ autocomplete offered ZERO bot rows (local and remote alike)
- remote-bot mentions never matched the union roster and fell back to a
local-only profiles.list, so @name-device mentions did not route
That is issue #89303: the remote-source Bot row tells the user to
`@name-device` the agent, and the completion surface never offers it.
Fix: cachedUnionRoster() — read the live connection's cache entry first
(getQueryData with the full 3-element key), fall back to a prefix scan
(getQueriesData) keeping the freshest snapshot. Never throws; callers
keep their existing cold-cache fallbacks.
The plugin test mocked getQueryData as key-agnostic, which is why this
never failed in CI: the mock now reproduces the real QueryClient v5
exact-key semantics (a bare ROSTER_KEY lookup misses the 3-element
cache), and the tests fail against the old code (verified: 2 failures
before the plugin fix, 9/9 pass after).
Fixes#89303
Keep-alive workspace panes (#89788) stay mounted while hidden, so returning
to an already-open room never remounted GroupChatWorkspace and the
mount-time bottom anchor didn't rerun — reopening group A after visiting
group B left A at the old scroll position. GroupChatMainView now subscribes
to the pane's visibility via feature-detected host.paneVisibility (always-
visible atom fallback for older SDKs) and the workspace scrolls the bottom
sentinel into view on the hidden→visible edge. Repro and fix direction from
the live-audit review comment on #90526.
Also updates the composer shape test to slice only GroupMentionInput
(GroupClarifyCard, merged since the branch was cut, legitimately uses
Input) and teaches the legacy-SDK react proxy about useMemo.
Composer was a single-line Input (Enter always submitted, newlines
impossible) — now the SDK Textarea with Enter=send / Shift+Enter=newline
and popover-first key handling. Room log had no scroll anchoring (opened
at position 0) — bottom sentinel + near-bottom-guarded anchor effect.
Stranded replies were only harvested inside an active turn loop (stuck
until the user's next send) — the settle path now runs a bounded
background harvest that yields to a live loop.
Found in live GUI E2E of the roomId salvage: disbanding a room while a
drive is mid-turn leaves an epoch-bump tombstone under the room's key.
Two defects surfaced: (1) updateGroupChat persists the WHOLE atom map, so
the next unrelated room write persisted the tombstone as an empty room;
(2) create/rename collision sets counted the tombstone's key, so
recreating the group under the same display name silently became
'<name> 2'. Tombstones are now flagged, excluded from the durable map,
from create/rename collision sets (liveGroupChatNames), and from roster
rows. Regression coverage in the existing tombstone test.
New rooms title member sessions 'Group: <roomId>'; the dialog copy now
describes the kept sessions generically instead of quoting a title that is
wrong for post-roomId rooms.
The visual half of asking someone to connect a thing: the mark, the card,
the settled scaffold line it collapses to, and the trust badge.
It renders a subject, a state, and an outcome, and calls back. What a
connector IS, how one connects, and where the strings come from all stay
with the caller — the prop types are structural, so a richer type satisfies
them without this layer importing it. That boundary is the point: any
surface that has to ask "connect this?" should look identical without
sharing a data layer.
MarkdownLinkText comes along because setup steps need inline links.
LinkifiedText can't serve them — it finds bare URLs and guesses a label,
and here the label carries the meaning while the URL is a console page
whose own title is useless or, behind a login wall, actively wrong.
A name with no bundled brand glyph had nothing to render, and the renderer
could not go find one — CORS makes another origin's markup unreadable.
So resolution moves to the main process: read the page, collect every icon
it declares in its head and web manifest, rank them, sniff the bytes to
catch a challenge page served as image/png, and cache the answer per host
on disk. Only ever the site's own marks — a public icon service would
answer for the hosts behind a bot wall, but asking one means telling a
third party which tools someone is wiring up, so an unreadable site keeps
its monogram instead.
Messaging platforms, the MCP servers tab, and the connector surfaces each
drew their own square brand chip, and the three copies had already drifted
on radius, glyph size, and what an unknown name falls back to.
They now share one AvatarChip. Size and radius stay the caller's call;
everything else is fixed in one place, so a Slack row and a Linear card
wear the same mark. brandFor also stops caring how a name is spelled —
catalog slug, registry id, or display name all reach the same glyph.
The SDK could contribute a theme through `THEMES_AREA` but never select one,
and `useTheme().setTheme` needs a component to hang the hook on. A plugin that
repaints on an event — a gateway coming up, a socket message — had nowhere to
call, so shipping one meant patching the app and re-patching it after every
upgrade.
`requestTheme(name)` writes to the same one-shot channel the backend `/skin`
sync already uses, so the ThemeProvider drains it through `setTheme` and an
imperative switch normalizes and persists per profile exactly like a manual
pick — one policy, one owner.
An unresolvable name is refused rather than coerced. `setTheme` falls back to
the default skin, which is right for a person picking off a list and wrong for
code reacting to an event: a gateway naming a theme the user never installed
would silently reset an appearance the caller never meant to touch. The
returned boolean doubles as the availability check.