Commit Graph

3814 Commits

Author SHA1 Message Date
hermes-seaeye[bot] 77001a6be7 fmt(js): npm run fix on merge (#95924)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 23:21:02 +00:00
Teknium db127f7502 fix(desktop): session rows are identified by (profile, id) — twins in two profiles stop collapsing and mis-routing (#92454)
Two profiles can hold sessions with the SAME stored id (restored
backups, copied state.dbs, cross-profile imports). mergeSessionPage
keyed rows by bare id, so the twins collapsed into one sidebar row
whose title/activity carry stitched one profile's content onto the
other's route — clicking a row previewing profile A resumed profile B
and wedged on an eternal 'Waking up…' when the mismatched resume never
completed (live-reproduced on main).

- mergeSessionPage: identity + lineage + dedupe keys are now
  (profile, id); a kept twin in another profile survives the incoming
  page dedupe. Local profile normalize (importing @/store/profile would
  be circular).
- Sidebar clicks carry the ROW as the identity: onResumeSession passes
  the clicked SessionInfo, and wiring pins the row's own
  (connection, profile) as the resume owner via requestSessionResume
  before navigating. Untagged rows keep the id-only path.
2026-08-26 15:53:41 -07:00
Hermes Agent 3ee0c62224 fix(desktop): SSH orphan reaper fails CLOSED on remote lockfile schema/ownership skew
A backend.lock.json that exists but doesn't match what this build writes
(unknown/future schemaVersion, truncated JSON, missing or foreign
ownershipId, malformed shape) was previously indistinguishable from 'no
lockfile': connect() would spawn a fresh backend on top of it and
overwrite the record, and cleanup paths could drop foreign state —
disarming the #78872 ownership guard exactly when another (e.g. forked)
desktop build shares the remote. readLockfile now returns a skew
sentinel for existing-but-foreign lockfiles; connect() refuses with a
'remote-lockfile-skew' error and a skew warning instead of
reaping/overwriting, disconnect() and cleanupStale() skip entirely.

Refs #95532
2026-08-26 15:51:54 -07:00
Finn763 c9d7b22e05 fix(desktop): escape reloadUrl in error page inline script (script-tag breakout)
JSON.stringify does not escape '<' — a reloadUrl containing
'</script><script>…' would terminate the inline <script> element of the
data: error page and let an attacker-controlled URL inject markup/script.
Escape <, >, & (and U+2028/U+2029) as \uXXXX sequences after stringify;
add regression test.
2026-08-26 15:51:22 -07:00
Finn763 4bc84d0499 fix(desktop): replace white screen after update with visible error + auto-reload (#95575)
A torn renderer bundle (update replaced the app while its files were
locked, e.g. antivirus or a still-running instance) loads index.html
fine and then dies on the first lazy import — a white screen with only
a desktop.log line. A main-frame load failure (missing index.html,
blocked file) was likewise log-only.

- resolveRendererIndex() already detects torn bundles; the primary
  window now refuses to load one and shows a visible repair page
  (error code, missing assets, 'hermes desktop --force-build', Reload)
  instead of a blank window.
- did-fail-load on the main frame now gets bounded auto-reload through
  the shared rolling reload budget (transient failures self-heal) and,
  once the budget is exhausted, surfaces the visible error page.
  ERR_ABORTED and sub-frame failures stay log-only, and helper windows
  (OAuth/portal) keep their log-only policy (opt-in via
  reloadOnFailedLoad).

Regression tests cover the policy decisions (reload / abort /
budget-exhausted surface), budget sharing with render-process-gone,
and the error page content + data: URL loading.
2026-08-26 15:51:22 -07:00
hermes-seaeye[bot] d9e4b234a3 fmt(js): npm run fix on merge (#95861)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 21:50:14 +00:00
hermes-seaeye[bot] 74cb4cb80c fmt(js): npm run fix on merge (#95858)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 21:44:19 +00:00
ethernet a6d6060d61 desktop: remove js vestiges 2026-08-26 14:38:57 -07:00
Brooklyn Nicholson a2c3b54939 fix(desktop): put attachment close and code copy back on hover-reveal
#95611 painted those two always-on. With hover un-gated, restore the
corner reveal so they match the rest of the chrome.
2026-08-26 16:38:46 -05:00
Brooklyn Nicholson 2f87fa66ad fix(desktop): don't gate hover-reveal on the hover media query
Tailwind v4 wraps hover:/group-hover: in @media (hover: hover). Windows
hosts with a digitizer often answer false even with a mouse, so those
controls stay opacity-0 — clickable, invisible. Trust :hover itself.

Co-authored-by: xxxigm <54813621+xxxigm@users.noreply.github.com>
2026-08-26 16:38:46 -05:00
Brooklyn Nicholson be7eefec8d fix(desktop): restore scroll on the capped thinking preview
The live thinking body pins to the newest tokens and keeps max-h-40
after the turn settles so the transcript doesn't jump. overflow-hidden
let that pin work in JS but clipped the rest of the thought. overflow-auto
makes the cap a real scroller; overscroll-contain keeps the wheel from
chaining into the transcript.

Supersedes #73757.

Co-authored-by: Dan Latimer <latdani@gmail.com>
2026-08-26 16:38:40 -05:00
Brooklyn Nicholson 19f9d1badb fix(desktop): fail closed when onboarding hits a model guard
The OAuth complete-with-model path is headless: a confirm toast would
hang with nobody to click it. Skip the prompt and surface the backend
message instead.

Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
2026-08-26 16:38:22 -05:00
Brooklyn Nicholson 8fe4816edd fix(desktop): confirm guarded Settings model applies
Settings → Model → Apply treated confirm_required as a red error, so
contributor-tier models like muse-spark-1.2-contributor could never be
saved. Prompt and retry with confirm_expensive_model, matching the
in-session picker handshake.

Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
2026-08-26 16:38:22 -05:00
Trevor Nash-Keller 15b673d178 fix(desktop): restore the typecheck by completing the typing-sync harness
`main` is red. 68518c1f9b added a dedicated `renderTypingSync` harness for the
typing-aware deferral tests, but its params object omits `updateSessionState`,
which `BackgroundSyncParams` requires:

    src/app/contrib/hooks/use-background-sync.test.ts(660,25): error TS2345:
    Property 'updateSessionState' is missing in type '{ ... }' but required
    in type 'BackgroundSyncParams'.

`npm run typecheck` exits 2, so `apps/desktop :: check:lint` fails and the
JS & TS checks job goes red on every open PR regardless of its contents.

Vitest did not catch this because it transpiles without typechecking, so the
new suite passes while `tsc -p .` fails.

Supply the missing prop. The updater runs against a throwaway state — this
harness never exercises the transcript path — and it is added to the existing
`stable` object rather than inline, because the harness's own comment requires
every param to keep a stable identity across the tick-driven re-renders; a
fresh `vi.fn()` per render would re-run the connect-reseed effect and
re-subscribe the throttle, polluting the very counts the tests observe.

Verified: all three tsc projects clean (`tsc -p .`, `tsconfig.electron.json`,
`tsconfig.e2e.json`), use-background-sync suites 26/26, eslint clean.
2026-08-26 16:34:23 -05:00
Brooklyn Nicholson 68518c1f9b fix(desktop): hold the sessions list refresh for the whole typing burst
The starvation cap in the original patch re-ran the heavy pass at the
same ~10s mark the freeze is measured at. Hold until the keyboard is
quiet, then land one coalesced pass.
2026-08-26 14:40:43 -05:00
Bruno Bza 01e9b9abb1 fix(desktop): defer heavy sessions.changed list refresh while typing 2026-08-26 14:40:43 -05:00
Leon Phull f0187332d1 fix(desktop): decode-probe app icon candidates instead of existence-only (#94806)
new BrowserWindow({ icon }) and app.dock.setIcon() decode the icon file
synchronously on the main process and throw on undecodable bytes. The
icon ladder was resolved with statSync().isFile(), which only proves a
file exists — a truncated or zero-byte PNG inside a packaged app.asar
(interrupted electron-builder run, partial copy) killed the main process
inside createWindow(): the window never appeared, running turns lost
their renderer, and the desktop log showed 'Uncaught exception: Error:
Failed to load image from path .../app.asar/public/apple-touch-icon.png
at createWindow'.

Resolution now runs through a decoding probe (nativeImage.createFromPath
must yield a non-empty image); a candidate that exists but does not
decode is skipped like a missing one, so the app falls through to the
next rung or starts with the platform default icon instead of dying.
The ladder and probe live in a pure module (electron/app-icon.ts) so
precedence is unit-testable without a running Electron app; window
factories re-resolve per call exactly as before.

Regression tests cover: skip-first-undecodable, all-fail -> undefined,
first-pass wins, missing/empty/directory rejection, and the unchanged
mac/Windows precedence ladder.

Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-08-26 18:44:46 +00:00
xxxigm b519ce29ad fix(desktop): keep attachment close and code copy icons visible (#95611)
* fix(desktop): keep attachment close and code copy icons visible

Hover-only opacity-0 hid the composer remove control and code-block copy button, so they stayed clickable but invisible on Windows and other no-hover surfaces.

* test(desktop): pin attachment close and code copy visibility at rest

The remove chip and code-block copy control must stay in the tree without a hover class, so Windows and no-hover surfaces cannot hide them again.
2026-08-26 13:38:19 -05:00
Gille d0fdbfd655 fix(desktop): show repo-root-only sessions in project drill-in (#94552) 2026-08-26 13:37:43 -05:00
Teknium 1a5547c5c5 fix(desktop): guard setPrimaryGatewayConnectionId against non-primary active scopes
Hardening follow-up to #95396 (#95628): ignore primary connection-id writes
while the active key is a composite secondary scope, so future
presentation-layer writes cannot relabel the primary socket and poison
new-session routing. Regression test is sabotage-proven (fails with the
guard reverted).
2026-08-26 09:52:27 -07:00
David Bartoš 10746a53ba fix(desktop): preserve primary gateway identity across source switches 2026-08-26 09:52:27 -07:00
Thawatchai 7eb11cabfc fix(desktop): route clarify responses through session owner 2026-08-26 09:51:55 -07:00
Luke Roberts 6fdaef6a9f fix(desktop): resolve session owners from the cron and messaging sidebar slices
The owner ladder's row rung (tile route -> hint -> session row) only searched
$sessions (recents). Cron- and messaging-sourced sessions are fetched as their
own sidebar slices ($cronSessions / $messagingSessions), so a scheduler-minted
cron session had no tile, no hint, and no row the rung could see. On a
registry-topology install every session-scoped RPC for such a session then
failed closed:

  Session owner could not be resolved for "xxxx" (approval.respond): no owner
  route, hint, connection-tagged row or profile probe named the backend ...

which made command approvals raised inside a cron chat impossible to answer
(the approval bar errors on every choice), even on a single-local-backend
install whose cron row - with its profile stamp - was already loaded for the
sidebar's cron section. The approval-bar path (requestForOwnedSession) has no
async REST-probe rung, so nothing downstream recovered.

Fix: ownerLookupSessionRows() returns the union of the three source-scoped
slices for OWNER lookups, keeping $sessions' array identity in the common
recents-only case so reference-keyed memo caches still hit. All four row-rung
call sites (session-rpc-dispatcher, knownOwnerForSession, tile delegate, tile
actions) now search it.

Repro: any cron session that raises approval.request (open the cron chat,
reply with a prompt that triggers a gated command, click Run).
2026-08-26 09:51:55 -07:00
Teknium 778feee37a chore: eslint --fix on salvaged e2e spec 2026-08-26 09:47:23 -07:00
686f6c61 bdbb41b62e fix: adapt salvaged staleness-probe tests to the kickoff option-object signature
createCanonicalChat gained { kickoff } on main (#95326) after #91227 was
filed with a bare positional openingStillCurrent; the salvage merges both
into one options object.
2026-08-26 09:47:23 -07:00
686f6c61 0ecbf1ce91 fix(desktop): snapshot Close All pane ids before persist-close
closeSessionTile can rewrite the layout tree. Iterating the live
group array then skips tiles. Copy the list first.
2026-08-26 09:47:23 -07:00
686f6c61 f1755cc1c5 fix(desktop): persist Bot Mode Close All session tiles
Close All only dismissed layout-tree panes. Bot tiles live in the
shared __bots_workspace__ bucket, so clicking a bot or swapping
profiles rehydrated the closed tabs. Persist-close those tiles
before dismissing the rest of the strip.
2026-08-26 09:47:23 -07:00
funky-xamarin 9ea7a37cc1 fix(desktop): restore live group after bot open failure 2026-08-26 09:47:23 -07:00
funky-xamarin fbd149035b test(desktop): verify group-to-local-bot handoff in Electron 2026-08-26 09:47:23 -07:00
funky-xamarin 35ee27b1c6 test(desktop): add RED group-to-local-bot E2E 2026-08-26 09:47:23 -07:00
funky-xamarin 7c255dbbbf test(desktop): focus group-to-local-bot handoff on close-callback RED
Rewrite the handoff proof to ≤150 lines on the hermes-bots VM seam so
base fails because the registered group closer is not invoked, not a
missing helper marker. Cover BotRow/Active Now local close-before-open,
remote no-dismiss, and no-group/old-host safety.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin 2fa39fac8c test(desktop): make group-to-local-bot handoff proof behavioral
Exercise real BotRow and Active Now open handlers so close-before-open
and remote stay-put are load-bearing, not source-regex false greens.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 09:47:23 -07:00
funky-xamarin f72786f764 fix(desktop): close group main tab when opening a local bot chat
Local BotRow and Active Now opens now share a handoff that retires the
selected group workspace before canonical chat open, so the bot chat can
own the main surface after leaving a group tab.
2026-08-26 09:47:23 -07:00
Teknium 9aa7530f7b style(desktop): sort oauth-partition import (lint) 2026-08-26 08:40:07 -07:00
Teknium 697087f2eb fix(desktop): SIGKILL-escalate the owned SSH backend when it survives the graceful quit wait (#91668 remainder)
The #95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
2026-08-26 08:40:07 -07:00
Teknium 8497edb2ac test(desktop): pin the activation-epoch guard against mid-handshake profile switches (#92434 close-candidate)
Reproduces the reported Bot↔Default switch shape at the gateway.ts
activation seam: a switch-back that lands while the outgoing switch's
WS handshake is still pending keeps the route, the late-completing dial
neither steals the foreground nor breaks its socket, and re-activating
the bot works without an app restart. The guard (activation epochs +
open-socket-publish, landed via #89622/#92265/#81094) already prevents
the reported permanent break; this pins it so it cannot regress.
2026-08-26 08:40:07 -07:00
Teknium 1e9a12a71f fix(desktop): stop spawning loopback serve children when the registry primary is remote (#91564, #90316)
'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
2026-08-26 08:40:07 -07:00
Teknium 62e2d6e1e4 fix(desktop): quarantine malformed connections.json entries per-entry instead of dropping or nuking the registry (#94246)
- normalizeRegistry now preserves every malformed entry (unknown kind,
  url-less remote/cloud, host-less ssh, mangled non-object items, and
  any entry whose normalization throws) under a capped 'quarantined'
  key that survives write cycles — healthy entries keep loading and
  user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
  connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
  a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
  surfaces reason+label summaries (never raw entries/token envelopes).
2026-08-26 08:40:07 -07:00
Teknium 31250da505 fix(desktop): key cookie-auth session partitions on connection identity, not auth mode (#92183)
Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
2026-08-26 08:40:07 -07:00
Zeus-Deus fd565c80e9 feat(desktop): fleet profile rail — every registered gateway's agents on one strip
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.

- Every registered gateway's profiles sit on the one strip, in registry order
  (This device first, then by label), each group headed by that gateway's
  kind glyph. The active gateway's squares are unchanged; the others are
  "at rest" (dimmed) with tooltips/accessible names qualified by machine
  (`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
  the statusbar switcher, landing on that exact (gateway, profile):
  `selectConnection(id, { profile })`. The spinner sits on the clicked
  square; the previous source stays painted until the target answers.
  Groups keep their slots whichever gateway is active, so a square never
  moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
  SOUL.md / Delete, executed on the owning gateway (renameProfile,
  getProfileSoul and updateProfileSoul accept the same scope deleteProfile
  already had); the delete confirmation names the machine. The legacy
  per-profile "Connect to a remote host…" item is hidden on multi-gateway
  setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
  two registrations of one backend collapse to one group; past thirteen
  squares across the fleet the strip condenses into a menu sectioned by
  gateway. Roster is fetched on mount / focus / registry change only — no
  periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.

Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.

Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.

Docs: multi-connection-desktop.md describes the fleet rail.

Refs #89304, #92384, #91047, #94724

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 08:39:45 -07:00
Teknium dd0aae4173 fix(desktop): paint stored Bot Chat history immediately instead of stranding the wake on an unsatisfiable profile gate
Fixes #89843. On a shared-remote connection every profile is served
through the primary socket, so waitForFocusedSessionHydration's
profileMatches gate could never become true — a bot chat whose stored
transcript painted within seconds still burned the whole 20s hydration
budget and then stranded the pane with 'Timed out loading <bot>'s
session history'.

The wake now resolves paint-first: once the stored transcript is painted
on exactly the target session, the content is its own proof — the pane
opens immediately and a subtle 'Syncing…' badge (new $hydrationSyncProfile
atom + ChatSyncBadge) shows until the profile gate catches up in the
background. Fail-closed everywhere content is not its own proof: a
superseded/conflicting concurrent wake still rejects, and an
expected-empty chat still waits for the full runtime gate.
2026-08-26 08:39:45 -07:00
etzelvon dbbed456c1 style(desktop): satisfy curly + padding-line lint for guarded switch
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
2026-08-26 08:39:45 -07:00
etzelvon a72d6d6338 fix(desktop): address review — staleness guard, neutral confirm fallback, clarify pending return
- staleness guard in applyConfirmedSwitch: bail and dismiss if
  current model/session no longer matches snapshot this warning was
  created for, preventing stale Confirm from clobbering newer pick
  (Enough1122 review #92492)
- neutral fallback for missing confirm_message: 'Confirm this model
  switch?' instead of modelSwitchFailed
- document selectModel boolean: false means not-applied (pending
  confirmation or failed) — pending already shows warning, not an error
- test mock: add dismissNotification mock for guarded flow

Addresses https://github.com/NousResearch/hermes-agent/pull/92492#issuecomment-5387151656
2026-08-26 08:39:45 -07:00
etzelvon 7450266a3a fix(desktop): confirm guarded model switches instead of snapping back
Desktop composer ignored gateway's confirm_required response for
contributor / expensive models. It painted the target optimistically
then invalidated model-options and refetched the still-active session,
so muse-spark-1.2-contributor appeared to instantly snap back to
gpt-5.6-sol with no explanation.

Now handle the confirmation protocol: on confirm_required rollback the
optimistic state, surface the backend's confirm_message as a warning
notification with a Confirm action, and on Confirm retry config.set with
confirm_expensive_model:true. Preserves data-training consent, keeps
the fix session-scoped and test-covered.
2026-08-26 08:39:45 -07:00
konsisumer b3a2065ff3 fix(installer): target Windows updater shim kills 2026-08-26 07:59:01 -07:00
hermes-seaeye[bot] 277899f10c fmt(js): npm run fix on merge (#95614)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 14:35:33 +00:00
beplee ec8ca8f2cb fix(desktop): re-bind open pane to rebuilt runtime after model switch
A mid-conversation model/provider switch rebuilds the agent runtime. The
rebuilt runtime emits session.info (and all later events) under a NEW
explicit session_id while the pane still holds the dead one as its
active id — isActiveEvent is false for the same conversation from that
moment on, so view-scoped updates stop and the chat freezes until a
full resume (#93942 scenario B; backend even logs 'client should resume
the stored session', but the client never does).

Fix: when a session.info event lineage-matches the selected conversation
(sessionMatchesStoredId over stored_session_id) but carries a different
runtime id, adopt the new runtime id as the active session id — keeping
the durable selection untouched — so every subsequent isActiveEvent gate
keeps matching without a resume. Guarded: the old runtime must show no
live turn (not busy/awaiting/streaming) or the adoption is refused, so
an overlapping manual switch can never split one conversation across
two panes.

The existing compression-rotation path does not cover this case: it
fires when the SAME runtime's stored id rotates, while a rebuild
produces a NEW runtime with a NEW stored id.

Together with #94255 (tile reconcile on sessions.changed), closes
#93942.

Regression tests verified failing pre-fix on 41447a6d70.
2026-08-26 07:28:10 -07:00
beplee 7cfed25370 test(desktop): behavior tests for tile reconcile + signature pruning (#94255 review)
Enough1122 review points on #94255, all addressed:

1. Source-grep Python tests replaced with real vitest behavior tests:
   - a tile whose stored transcript gained a background delivery IS
     reconciled (updater invoked, correct stored id fetched)
   - an unchanged transcript is skipped entirely (no updater call —
     the signature gate proven, not asserted by regex)
   The structural smoke test in test_bots_chat_live_append.py stays as
   a cheap drift alarm; the contract now lives here.

2. Shared-sequence latest-wins semantics documented in the reconcile
   docstring (review point 2).

3. Signature map pruning: a tile closed/superseded mid-read now deletes
   its signature entry instead of leaking one map slot per ever-opened
   tile for the app lifetime (review point 3).

4. Test harness: typed updateSessionState mock via Parameters<> instead
   of the {}-as-state cast; no more silent spread corruption.

No production behavior change beyond pruning: 18/18 hook suite green,
typecheck clean.
2026-08-26 07:28:10 -07:00
beplee 3669fa3095 fix(desktop): satisfy lint — read busy atom directly in tile reconcile
CI caught two lint issues in the tile-reconcile path:
- no-restricted-syntax: don't mirror the $busy atom into a ref via
  useEffect (stale-read hazard); reconcileTileTranscripts now receives a
  live getter view so the loop reads the current value at tick time.
- react-hooks/exhaustive-deps: add updateSessionState to the
  sessions.changed effect deps (stable useCallback from
  useSessionStateCache, so no extra re-subscription).

No behavioral change: regression tests 2/2, adjacent hook suite 58/58,
typecheck clean.
2026-08-26 07:28:10 -07:00
beplee db8ff4eb75 fix(desktop): reconcile workspace-tile transcripts on sessions.changed
Bot canonical chats open as workspace tiles (workspaceMode: 'bots') and
are deliberately hidden from $sessions/$messagingSessions, so the
sessions.changed transcript refresh skipped them twice over: it covers
only the main pane's selection, and its resolveSession() bails on hidden
sessions. A background delivery (bot-to-bot DM via bot_relay.deliver, a
cron run's output, another machine) therefore never reached an open bot
chat — the roster updated but the pane stayed stale until remount
(#93942 scenario A).

Fix: the sessions.changed tick now also reconciles every visible
workspace tile through a dedicated signature-gated path. Each tile
carries its own stored↔runtime id pair so no resolution step is needed;
per-tile signatures make no-change ticks free; busy tiles are skipped
(their own stream owns the view); closed/superseded tiles discard their
in-flight read.

Slice 1 of 2 for #93942 (scenario A only). Scenario B (stream re-key
after mid-conversation model switch) follows separately.

Fixes part of #93942
2026-08-26 07:28:10 -07:00