The Bots editor's model write (profiles.configure) was the one switch
surface that bypassed the data-policy / expensive-model selection guard:
a guarded pick (e.g. muse-spark contributor tier) was applied silently,
with no confirm flow anywhere — the #95293 remainder after the core
picker's confirm handshake landed in use-model-controls.
Gateway: profiles.configure now answers confirm_required +
confirm_message for a guarded model (same handshake as config.set
model) and writes NOTHING until the client resends with
confirm_expensive_model: true. Other sections still apply; the pending
model section is not reported as failed.
Desktop: the confirm flow is extracted out of use-model-controls into
one shared applier (lib/guarded-model-switch.ts, exported through the
plugin SDK) — warning toast, staleness-guarded Confirm, single
confirmed resend, never a retry loop. The core picker and the Bots
editor now consume the SAME handler; the Bots editor's Confirm resends
only the model section with confirm_expensive_model: true.
Fixes#95293 (Bots surface remainder).
Refresh Models only updated the catalog cache, so the composer kept showing a model that was no longer in the new group list.
Co-authored-by: Cursor <cursoragent@cursor.com>
The Bots model picker's catalog read rode the bot's own socket with no
deadline: a wedged dial left the query pending forever, so the picker
spun indefinitely. On top of that every fetch forced refresh:true,
bypassing the staleTime cache, so each Bots view remount (tab re-front,
dialog reopen, pane visibility flip) knocked the picker back into its
loading state and wiped the staged provider/model pick mid-edit.
Bound every attempt to 20s (rejection falls through to the picker's
existing free-text fallback), drop the forced refresh so the read
participates in the cache like every other surface's catalog, and pin
the contract with a red->green regression suite.
Slim renderer UI for the managed SSH remote update engine (#95942),
adapted from #93042's renderer unit with the deferred canary/rollout
scope stripped. Adds a per-connection store (idle/updating/terminal
states, receipt, managed-update-in-progress busy envelope) and a
'Managed updates' section on the Gateways settings page with an Update
button, progress line, and correlated receipt per registered
Desktop-managed SSH connection. Fails closed when the Electron main
lacks connections.updateManaged.
The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).
Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
Group rooms persist source-qualified members. After Desktop switches to
the built-in This-device source, a dead loopback row still listed next
to the live profile and looked like a second agent. Collapse only the
sidebar tiles; $lastRoster, group seats, and mentions keep every
(connectionId, profile) identity.
main's hermes:connections:update-all handler grew (renderer-side exclusions +
the managed-SSH dispatch branch) after #95606 branched, pushing the claim-
guarded dial past the 2,000-char slice.
(connectionId, profile) scope, but it was only wired into hermes:connection,
hermes:connection:for, and the power-resume pool rebuild. Five other real
call sites still invoke ensureRegistryBackend()/ensureBackend() directly:
the media-protocol connection resolver, the terminal-pane backend resolver,
the ~5s roster-enumeration probe, the connections update-all dispatch, and
dispatchRegistryApiRequest (every registry-scoped hermes:api REST call).
ensureRegistryBackend() has a genuine await-then-check race
(reuseMatchingPrimarySshBackend before the pool entry check-and-set), so any
of these five racing a guarded dial for the same scope can each bootstrap
their own SSH tunnel / remote dashboard for the same connection — exactly
the symptom class #90812 was written to prevent, just reached through an
unguarded path instead of two renderer windows.
Left the ensureRegistryBackend() self-call inside its own dispatch-time
health-probe reconnect branch (registryDispatchRevalidation) unguarded —
that recursive path has its own coordination semantics and touching it
without modeling reentrancy against the newly-guarded entry points here is
a separate, riskier change better done on its own.
A wake-path liveness-probe timeout force-closed the primary renderer
socket even when the backend was merely busy mid-tool-call; the gateway
then saw its client vanish, ws_orphan_reap expired, and the running turn
died as a bare "Operation interrupted." placeholder.
While any session still reports working, the first inconclusive probe
timeout now defers the teardown behind one bounded re-probe; only an
exhausted consecutive-failure streak (or no in-flight work at all)
rebuilds the transport. The streak resets on a successful ping, a
healthy -32601 answer, a clean open, a gateway switch, and unmount.
The registry's active-connection-invalidated fallback re-dial was the one
getConnection() await in this file the #93454 bound-every-IPC-round-trip
sweep never reached. A wedged main-process round-trip during an eviction
fallback (idle reap, connection removal, profile delete) left $connection
latched on a promise that never settles instead of rejecting into the
existing catch/publish(null) path.
The #92434 mid-handshake pin (added on main after #95343 branched) latches
gatewayMocks.connect on a never-resolving mockImplementation; vi.clearAllMocks()
clears calls, not implementations, so the salvaged wedged-probe test inherited
a dial that never completes and timed out.
20s withTimeout() on the boot() and soft-switch paths (use-gateway-boot.ts),
since these are IPC round-trips into the main process with no timeout of
their own — a wedged main-process round-trip hangs the awaiting caller
forever instead of surfacing a failure.
Every other production call site of the same IPC pair was still unbounded:
- store/gateway.ts's openSecondary() and sharedPrimaryRoute() — the actual
connection-establishment underneath requestGatewayForProfile/Agent,
ensureGatewayForProfile/Agent, and every other exported routing entry
point that opens a non-primary profile's socket.
- use-gateway-request.ts's on-demand reconnect (the primary gateway's
"not connected" retry path hit by every RPC).
- voice-playback.ts's resolveSpeakStreamUrl().
- api/plugins.ts's activeConnection() (pluginSocket's connect()).
Extracted RECONNECT_ATTEMPT_TIMEOUT_MS into the shared lib/with-timeout.ts
(previously local to use-gateway-boot.ts) so every call site uses the same
budget instead of duplicating the constant.
Regression tests mirror the existing use-gateway-boot.test.tsx hang-repro
pattern: wedge getConnection()/getConnectionFor() with a never-resolving
promise, advance fake timers past the 20s bound, assert the caller settles
instead of hanging. Mutation-verified: reverted the production fix (kept
tests) and confirmed the 6 new tests fail — 4 by genuinely timing out at the
vitest level, 2 by TypeError on the not-yet-exported activeConnection —
restored the fix and confirmed all 70 tests across the gateway/voice/boot/
plugins suites pass, with tsc -p . --noEmit clean throughout.
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.
- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
Review follow-ups on the salvage (#94914): the delete path derives its
scope from the removed row while saves derive from ownerRoute — a shape
drift orphaned the entry until LRU eviction. Deletes now sweep every
scope of the stored id (safe: deleted ids are never reused), while the
failed-resume path keeps the exact-scope drop (a same-id twin in another
profile must keep its tail). purgeLegacyV1 latches only after a
completed sweep so a mid-sweep throw retries next touch.
Stored session ids are only unique within one profile's state.db, and
localStorage survives profile switches in the same window. The durable
transcript-tail cache keyed entries by bare stored id, so a tail cached
while working in profile A was painted against profile B's backend after
a switch; that backend never held the session and retried it on every
wake ('session not found'), matching the desktop.log pattern of repeated
session.reclaimed (ws_orphan_reap) events for ids that live in another
profile's state.db (#94828).
Scope every entry by its owner — the same {connectionId, profile} shape
the REST layer already threads through getLatestSessionMessages and the
in-memory twin stores in TranscriptTailState:
- key entries under v2:[connectionId, profile, storedId]; loads only
return a tail saved under the SAME scope
- thread the session's owning scope through all call sites
(resumeSession's warm-path save + cold-paint load/rollback drop,
final save, removeSession's delete drop)
- sweep pre-scoping v1 entries once per window: a bare-id key cannot be
attributed to an owner, so it must never paint again
- dropTranscriptTail with a scope leaves other backends' same-id tails
intact
Re-implements the #93042 main.ts wiring against current main (post-#94724
drift), making the extracted managed-ssh-update engine reachable:
- ManagedConnectionUpdateGate instance + owner-only recovery journal at
DESKTOP_MANAGED_SSH_RECOVERY_PATH (read/write/persist/mark/clear with
strict record validation).
- IPC: hermes:connections:update-managed (requestManagedSshUpdate with
correlation-id claim + in-flight dedupe); update-all's ssh rows now route
through the transactional drain/update/restore lifecycle instead of
POSTing the remote backend updater.
- Gate enforcement at every dial/mutate seam: bootstrapSshConnectionInner
(pre-dial + publication fence with exact-serve rollback via
rollbackSshBootstrapResult), resolveRemoteBackend, ensureRegistryBackend,
saveRegistryConnection dial-field edits, connections:remove, and
primary-routing mutations (set-primary, set-launch-mode,
connection-config save/apply, profile:set) via
assertCanMutateManagedPrimaryRouting.
- Scope capture/drain/restore drivers: captureManagedSshScopes (pool +
primary discovery, bootstrap fence join), drainManagedSshScope (exact
identity-re-proof termination, no-kill forward recovery),
ensureManagedSshBackend(AtKey)/restoreManagedPrimarySshBackend restores,
openManagedSshUpdateTransport for serve-less connections.
- Startup recovery (resumeManagedSshRecoveries before createWindow) and
before-quit join of in-flight update/recovery operations BEFORE the SSH
coordinator is sealed, so restore dials are not refused during quit.
- Extended sshConnections state (spawnNonce/creationTime(Ns)/startedAt/
hermesPath/hermesHome/pythonPath/remoteProfile/registryConnectionId/
primaryRegistryScope) so drain can prove the exact serve it owns;
bootstrap coordinator entries carry metadata for the update fence;
persistSshConnectionToken mirrors tokens per managedSshTokenPersistencePlan.
- preload/global.d.ts: connections.updateManaged +
DesktopManagedConnectionUpdateResult/Receipt types.
Renderer UI (fleet-updates store, about-settings, system.ts ProfileScope
plumbing, i18n) intentionally NOT wired — it belongs to the deferred fleet
rollout UI and follows separately.
Wiring re-implemented against current main; design from #93042 by @andrexibiza
tsc -p apps/desktop clean; electron project 1924/1924 passed (133 files,
incl. 131/131 across the three engine suites); eslint clean on touched files.
The Desktop can drive a full update of a REMOTE SSH instance: claim the
connection (ManagedConnectionUpdateGate pauses dials/mutations while the
update owns it), terminate the owned backend with an identity re-proof
at the signal boundary (terminateOwnedDashboardForUpdate — argv+creation
time re-read in the same remote shell that signals), run the updater
under an update-in-progress marker/mutex on the remote install root,
bootstrap the new backend, and fence its publication so a rollback
cannot leak a half-published serve (fenceManagedSshBootstrapPublication
+ waitForManagedSshBootstrapFence barrier).
Extraction notes (campaign #91277 Phase 4; rollout/canary engine stays
behind per sequencing):
- managed-ssh-update.ts + test: clean cherry-pick (new files).
- windows-remote-lifecycle.ts + test: applied as-is (zero main drift).
- remote-lifecycle.ts: PR hunks stitched AROUND current main's #95532
skew guards and #91668 SIGKILL-escalating cleanupStale, which are
PRESERVED verbatim — the PR's python identity-re-proof termination is
wired for the managed-update path only; connect's stale replacement
keeps main's proven kill. One PR test assertion re-pinned accordingly.
- One PR test fixed: floating coordinator.start() promise whose
rejection IS the contract under test now has an explicit handler
(vitest flagged it as an unhandled rejection).
tsc clean; 131/131 across the three touched electron suites. main.ts
wiring (IPC + deps bag) follows as a separate commit.
Two profiles can hold sessions with the SAME stored id (restored
backups, copied state.dbs, cross-profile imports). mergeSessionPage
keyed rows by bare id, so the twins collapsed into one sidebar row
whose title/activity carry stitched one profile's content onto the
other's route — clicking a row previewing profile A resumed profile B
and wedged on an eternal 'Waking up…' when the mismatched resume never
completed (live-reproduced on main).
- mergeSessionPage: identity + lineage + dedupe keys are now
(profile, id); a kept twin in another profile survives the incoming
page dedupe. Local profile normalize (importing @/store/profile would
be circular).
- Sidebar clicks carry the ROW as the identity: onResumeSession passes
the clicked SessionInfo, and wiring pins the row's own
(connection, profile) as the resume owner via requestSessionResume
before navigating. Untagged rows keep the id-only path.
A backend.lock.json that exists but doesn't match what this build writes
(unknown/future schemaVersion, truncated JSON, missing or foreign
ownershipId, malformed shape) was previously indistinguishable from 'no
lockfile': connect() would spawn a fresh backend on top of it and
overwrite the record, and cleanup paths could drop foreign state —
disarming the #78872 ownership guard exactly when another (e.g. forked)
desktop build shares the remote. readLockfile now returns a skew
sentinel for existing-but-foreign lockfiles; connect() refuses with a
'remote-lockfile-skew' error and a skew warning instead of
reaping/overwriting, disconnect() and cleanupStale() skip entirely.
Refs #95532
JSON.stringify does not escape '<' — a reloadUrl containing
'</script><script>…' would terminate the inline <script> element of the
data: error page and let an attacker-controlled URL inject markup/script.
Escape <, >, & (and U+2028/U+2029) as \uXXXX sequences after stringify;
add regression test.
A torn renderer bundle (update replaced the app while its files were
locked, e.g. antivirus or a still-running instance) loads index.html
fine and then dies on the first lazy import — a white screen with only
a desktop.log line. A main-frame load failure (missing index.html,
blocked file) was likewise log-only.
- resolveRendererIndex() already detects torn bundles; the primary
window now refuses to load one and shows a visible repair page
(error code, missing assets, 'hermes desktop --force-build', Reload)
instead of a blank window.
- did-fail-load on the main frame now gets bounded auto-reload through
the shared rolling reload budget (transient failures self-heal) and,
once the budget is exhausted, surfaces the visible error page.
ERR_ABORTED and sub-frame failures stay log-only, and helper windows
(OAuth/portal) keep their log-only policy (opt-in via
reloadOnFailedLoad).
Regression tests cover the policy decisions (reload / abort /
budget-exhausted surface), budget sharing with render-process-gone,
and the error page content + data: URL loading.
Tailwind v4 wraps hover:/group-hover: in @media (hover: hover). Windows
hosts with a digitizer often answer false even with a mouse, so those
controls stay opacity-0 — clickable, invisible. Trust :hover itself.
Co-authored-by: xxxigm <54813621+xxxigm@users.noreply.github.com>
The live thinking body pins to the newest tokens and keeps max-h-40
after the turn settles so the transcript doesn't jump. overflow-hidden
let that pin work in JS but clipped the rest of the thought. overflow-auto
makes the cap a real scroller; overscroll-contain keeps the wheel from
chaining into the transcript.
Supersedes #73757.
Co-authored-by: Dan Latimer <latdani@gmail.com>
The OAuth complete-with-model path is headless: a confirm toast would
hang with nobody to click it. Skip the prompt and surface the backend
message instead.
Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
Settings → Model → Apply treated confirm_required as a red error, so
contributor-tier models like muse-spark-1.2-contributor could never be
saved. Prompt and retry with confirm_expensive_model, matching the
in-session picker handshake.
Co-authored-by: Silvio S. <silviomanuel297@gmail.com>
`main` is red. 68518c1f9b added a dedicated `renderTypingSync` harness for the
typing-aware deferral tests, but its params object omits `updateSessionState`,
which `BackgroundSyncParams` requires:
src/app/contrib/hooks/use-background-sync.test.ts(660,25): error TS2345:
Property 'updateSessionState' is missing in type '{ ... }' but required
in type 'BackgroundSyncParams'.
`npm run typecheck` exits 2, so `apps/desktop :: check:lint` fails and the
JS & TS checks job goes red on every open PR regardless of its contents.
Vitest did not catch this because it transpiles without typechecking, so the
new suite passes while `tsc -p .` fails.
Supply the missing prop. The updater runs against a throwaway state — this
harness never exercises the transcript path — and it is added to the existing
`stable` object rather than inline, because the harness's own comment requires
every param to keep a stable identity across the tick-driven re-renders; a
fresh `vi.fn()` per render would re-run the connect-reseed effect and
re-subscribe the throttle, polluting the very counts the tests observe.
Verified: all three tsc projects clean (`tsc -p .`, `tsconfig.electron.json`,
`tsconfig.e2e.json`), use-background-sync suites 26/26, eslint clean.
The starvation cap in the original patch re-ran the heavy pass at the
same ~10s mark the freeze is measured at. Hold until the keyboard is
quiet, then land one coalesced pass.
new BrowserWindow({ icon }) and app.dock.setIcon() decode the icon file
synchronously on the main process and throw on undecodable bytes. The
icon ladder was resolved with statSync().isFile(), which only proves a
file exists — a truncated or zero-byte PNG inside a packaged app.asar
(interrupted electron-builder run, partial copy) killed the main process
inside createWindow(): the window never appeared, running turns lost
their renderer, and the desktop log showed 'Uncaught exception: Error:
Failed to load image from path .../app.asar/public/apple-touch-icon.png
at createWindow'.
Resolution now runs through a decoding probe (nativeImage.createFromPath
must yield a non-empty image); a candidate that exists but does not
decode is skipped like a missing one, so the app falls through to the
next rung or starts with the platform default icon instead of dying.
The ladder and probe live in a pure module (electron/app-icon.ts) so
precedence is unit-testable without a running Electron app; window
factories re-resolve per call exactly as before.
Regression tests cover: skip-first-undecodable, all-fail -> undefined,
first-pass wins, missing/empty/directory rejection, and the unchanged
mac/Windows precedence ladder.
Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
* fix(desktop): keep attachment close and code copy icons visible
Hover-only opacity-0 hid the composer remove control and code-block copy button, so they stayed clickable but invisible on Windows and other no-hover surfaces.
* test(desktop): pin attachment close and code copy visibility at rest
The remove chip and code-block copy control must stay in the tree without a hover class, so Windows and no-hover surfaces cannot hide them again.
Hardening follow-up to #95396 (#95628): ignore primary connection-id writes
while the active key is a composite secondary scope, so future
presentation-layer writes cannot relabel the primary socket and poison
new-session routing. Regression test is sabotage-proven (fails with the
guard reverted).
The owner ladder's row rung (tile route -> hint -> session row) only searched
$sessions (recents). Cron- and messaging-sourced sessions are fetched as their
own sidebar slices ($cronSessions / $messagingSessions), so a scheduler-minted
cron session had no tile, no hint, and no row the rung could see. On a
registry-topology install every session-scoped RPC for such a session then
failed closed:
Session owner could not be resolved for "xxxx" (approval.respond): no owner
route, hint, connection-tagged row or profile probe named the backend ...
which made command approvals raised inside a cron chat impossible to answer
(the approval bar errors on every choice), even on a single-local-backend
install whose cron row - with its profile stamp - was already loaded for the
sidebar's cron section. The approval-bar path (requestForOwnedSession) has no
async REST-probe rung, so nothing downstream recovered.
Fix: ownerLookupSessionRows() returns the union of the three source-scoped
slices for OWNER lookups, keeping $sessions' array identity in the common
recents-only case so reference-keyed memo caches still hit. All four row-rung
call sites (session-rpc-dispatcher, knownOwnerForSession, tile delegate, tile
actions) now search it.
Repro: any cron session that raises approval.request (open the cron chat,
reply with a prompt that triggers a gated command, click Run).
createCanonicalChat gained { kickoff } on main (#95326) after #91227 was
filed with a bare positional openingStillCurrent; the salvage merges both
into one options object.