Exports createBudgetedLoop (+types) through the plugin SDK and migrates the
Bots plugin's face-animation clock onto it. The hand-rolled clock only
checked document.hidden; via the shared loop it now also pauses while the
window is minimized or unfocused, matching every other desktop render loop.
- sdk/index.ts: export createBudgetedLoop/BudgetedLoop/BudgetedLoopOptions
from @/lib/budgeted-loop with plugin-facing guidance
- hermes-bots/plugin.js: startFaceClock delegates scheduling to the SDK loop
when present (fps 15, idleWhen = no visible faces); paint body, IO-based
visibility tracking, and the 1Hz rescan are shared; the hand-rolled rAF
path remains as a feature-detected fallback for older desktops, per the
plugin's established SkillsView/McpTab pattern
- tests: new SDK-path case (fps/idleWhen wiring, visibility wake, re-entry
wake, dispose-on-stop) driven through an injected fake loop; verified to
fail with the SDK branch removed; existing fallback-path tests unchanged
CI's check:lint caught three real issues in the directive surface:
- TranscriptDirectiveLeaf called the contribution's render() inline in JSX
— the exact pattern no-restricted-syntax bans because the callback's hooks
land in the host and a plugin reload changes the host hook count
(React #310). The callback is now memoized and mounted via ContribRender.
- The inline frame mirrored its height state into a ref from the message
handler (the stale-read pattern no-restricted-syntax flags). Functional
setState reads current state directly; the shadow refs are gone.
- perfectionist import/export ordering in the frame and the SDK index.
The first inline frame was a full-width bordered box at a fixed height:
webpage-in-a-rectangle, not a widget. Now the frame disappears into the
message flow:
- Content-driven size. The injected measurer reports height (live) and
intrinsic width (adopted once, so %-width children can't feedback-loop the
frame toward zero). A sparkline shrink-wraps and sits flush left like an
inline image; a full-bleed page measures the whole viewport and stays
column-wide. The height attribute is now only a starting value.
- Theme bridge. A style prelude injects first with the app's resolved theme
tokens under stable names (--foreground, --muted-foreground, --accent,
--border, --card), the app font, zero body margin/padding, and a
transparent background — reference HTML written against those vars renders
native in any theme. Page styles override the prelude, so a page that
brings its own design keeps it.
- No chrome. Border, rounded box, and the rail-opener card under the frame
are gone; the fallback paths (non-HTML, remote gateway, unreadable file)
keep the classic card. The wheel gate went with the border — frames size
to content, so there is nothing to scroll inside, and widgets are fully
interactive.
- The desktop platform hint now teaches the default: an inline widget is
transparent, token-colored, flush left, no page chrome — only a standalone
page brings its own background. "Make me an inline sparkline" gets native
styling without the user spelling it out.
Some providers re-send the previous assistant text verbatim when a turn
continues past a tool call (a tool_calls row, then a stop row with identical
prose — both persisted). The turn merge folds both rows into one bubble, so
every paragraph in the reply rendered twice; inline ::preview frames made it
obvious. Repeated text parts now dedupe in the same pass as generated-image
echoes — the last occurrence wins.
The frame was a hardcoded 280px unless the model guessed a height attribute
— tall pages clipped (the flip-clock demo cut its last digit), short ones
floated in dead space. The opaque-origin sandbox means the parent can't
measure the document, but we own the srcdoc string: a tiny injected script
observes the document with ResizeObserver and posts its scrollHeight up via
postMessage, and the frame tracks it live within the 120-1200 clamp.
Reports are validated before they can move layout — per-mount random token
(two previews in one transcript, or a hostile page inventing messages,
can't move each other's frames), finite-number check, clamp. A 4px
tolerance stops vh-sized pages (which measure exactly what they're given)
from oscillating; an explicit height attribute still opts out of
auto-sizing entirely.
The first cut of the core ::preview consumer rendered the classic
preview-attachment card — a button into the right rail we already had, which
made the directive indistinguishable from an ordinary preview link. Now the
directive shows the thing itself: the workspace HTML file renders in a
sandboxed srcdoc iframe inline in the assistant message (opaque origin,
allow-scripts only — no reach into the app, its storage, or the bridge),
with an optional height attribute clamped to 120-1200px and the classic
card kept below as the rail escape hatch.
The frame waits for turn settle before reading the file (mid-stream it is
often mid-write), resolves relative paths against the session's own cwd,
and falls back to the plain card for non-HTML targets and remote gateways
(no local file door there).
The transcript becomes a contribution area (transcript.directives). A plugin
registers a named directive and the model addresses it by emitting
::name{key="value"} as its own paragraph; that leaf renders as the plugin's
component, wrapped in the contribution error boundary. Unclaimed or malformed
directives stay plain prose, so nothing changes for text that merely looks
like a directive (std::vector) or for users with the plugin disabled.
Core ships ::preview{file="..."} as the reference consumer (the existing
preview-attachment card), the desktop platform hint teaches the model the
syntax, and the SDK exports the area + types so runtime plugin.js files get
the surface through the normal plugins API.
Desktop shipped the same bug class four times in one week: a decorative
animation loop that never sleeps (bots face clock #88543, pixel egg #88406,
diffusion placeholder #88564, plus the #77651 hidden-renderer wave). Each fix
hand-rolled the same four behaviors. This extracts them into one helper in
src/lib/budgeted-loop.ts:
- fps budget (default 15) on top of rAF
- observability pause via the existing createRendererLoopPauseController
- idle dormancy: idleWhen() true after a draw parks the loop with zero
pending work until wake() — the piece every hand-rolled loop forgot
- teardown: dispose() cancels the frame, disposes the controller, and makes
wake() a no-op
DiffusionCanvas migrates onto it as the first consumer (net -40 lines at the
call site); its existing scheduling/budget/instance-cap tests pass unchanged
against the migrated implementation. Helper suite covers budget, pause,
park/wake, dormancy-survives-focus-churn, and dispose idempotency; sabotage
run (budget+dormancy stripped) fails 4/5.
Follow-up to #88523/#88542 from a community report (main agent listed twice,
both rows unnamed @default handles). Two distinct bugs, both specific to
remote-gateway-primary desktops:
1. Phantom "This device" default (electron): the roster enumeration dialed
ensureRegistryBackend for the registry's local entry unconditionally. On a
remote-primary desktop that forces resolveRegistryLocalRoute into the
forced-local branch — SPAWNING a local backend the user never asked for.
That backend enumerates a `default` profile, so a second default agent
appears AND the duplicate-handle rule forces -device suffixes onto the
real one. New shouldDeferLocalEnumeration() treats the forced-local route
as connect-on-demand (same courtesy as undialed SSH sources): the local
entry only enumerates when it is the delegate route (local-primary
desktops, byte-identical behavior) or a forced-local child is already
pooled (the user opened one).
2. Main agent renamed to a connection label (plugin): displayName keyed the
"show the connection label for a default row" rule off sourceScoped —
which annotation also sets on ACTIVE-source rows. A remote-gateway user's
main agent rendered as an IP-derived label (or bare handle) instead of
"Hermes"/their title. Key it off remoteSource: only THIN rows from
another source trade the friendly name for their source label.
Tests: shouldDeferLocalEnumeration route matrix (delegate always enumerates;
forced-local defers until a conn:local:: child exists; bare-key remote
descriptor doesn't count), displayName regression (active default stays
Hermes/title; thin remote default still shows its source label).
Follow-up to the cherry-picked fix for #88391: the bare `streamdown` import
broke the plugin's side-load contract (vm-harness tests and the legacy-sdk
tmpdir loader can't resolve bare specifiers, and the runtime plugin door only
maps @hermes/plugin-sdk and react). Export Streamdown from the SDK instead,
feature-detect it in plugin.js like SkillsView/McpTab (plain-text fallback on
older desktops), and fix the indentation at the render site.
Extends the scheduling test suite from #88407 with behavioral coverage for
the salvaged #79327 work: frames inside the 1000/15 budget reschedule
without repainting, instances over MAX_ANIMATED_INSTANCES draw one static
frame and skip the loop, and the counter releases on unmount. Both tests
verified to fail against the pause-controller-only version.
- Throttle frame rate to ~15fps via timestamp check instead of
redrawing on every requestAnimationFrame callback
- Pause animation when document.hidden, resume on visibilitychange
- Cap concurrent animated instances at 2; extras render a single
static frame instead of starting another animation loop
- Fixes#79077
A URL-remote desktop whose PRIMARY profile has a per-profile remote
override lists the gateway's sub-profiles in the Bots pane, but
clicking one fell through the routing table's last case and spawned a
fresh local backend that shared nothing but the name (#88296).
resolveProfileBackendRoute now consults primaryRemoteActive: when the
primary's own backend is remote and the sub-profile has no stored
entry of its own, it routes through the primary gateway with profile
scoping (the same shared-primary flow global remote uses). Profiles
with their own local entries still pool locally.
Follow-up to the salvaged #88219 visibility/15fps work:
- Dormancy: the rAF loop stops scheduling frames when no faces are mounted
or none are visible, instead of running the 1Hz whole-document shadow-root
scan forever. A mounting BotFace or a face scrolling into view wakes it.
- Teardown: register() now hooks ctx.onDispose so disabling the Bots plugin
(or a hot reload) cancels the animation frame, disconnects the
IntersectionObserver, drops cached nodes, and clears window.__hbFaceClock.
Previously the loop ran until app restart even with the plugin disabled.
- Behavioral tests for park/wake/stop via a vm-extracted clock harness,
verified to fail against the pre-fix source.
Layers on the salvaged #88489 (@29206394) and #88341 (@frizikk):
- sdk: host.activeConnectionId() — registry id of the LIVE active gateway.
The salvaged fix classifies against the registry primary; after the user
activates a non-primary source's agent, profiles.list answers from THAT
source and primary-based matching would duplicate the active source's
agents again. Live id wins, primaryConnectionId is the fallback, the
legacy kind==='local' rule covers older desktops.
- plugin: roster/chip/picker list keys are botRowKey(bot) — source-qualified
(connectionId, name) — so same-named agents on two genuine sources can
never collide as duplicate React keys (the render half of the dupe-bots
smear: name-keyed rows + duplicate names = repeated blocks every poll).
Annotated active-source rows keep the plain-name key, so nothing remounts
when a desktop gains the union roster.
- tests: live-id-beats-primary regression, botRowKey stability, source-shape
anchor refresh.
The union agent roster (host.agents) enumerates EVERY registered connection,
including the active gateway that already answered profiles.list. The plugin
merger treated the active gateway's own agents as rows from other sources
because a remote-primary desktop reports them with connectionKind 'remote',
so every bot appeared twice (baseline) and kept growing with each refetch.
Match union agents to the active gateway via the new primaryConnectionId
field on the roster RPC response and annotate the local rows in place
instead of appending phantom copies. Same-named profiles on genuinely
separate sources (This device, other remotes) still get their own tagged
rows, preserving the @name-device disambiguation rule.
Fall back to the legacy connectionKind==='local' rule when
primaryConnectionId is absent (older Electron builds), so single-source
behavior is byte-identical.
Fixes#88344
The no-payload settle gate in gateway-event.ts held session.info
running=false off unconditionally while an optimistically armed turn
(busy/awaitingResponse from restore/edit/submit) had not gone live
backend-side. When the turn never went live at all — a rewind refused
after the optimistic arm, a submit response lost to a gateway bounce, a
terminal error event that never arrived — busy latched forever:
isTargetSessionBusy refused every send, the composer queued each message
('moves to the send area'), and the queue drain (gated on busy→false)
never fired. Only an app restart cleared it (#86795).
Bound the hold to PRE_TURN_LIVE_SETTLE_GRACE_MS (15s) measured from
turnStartedAt; past the window (or with no clock) the gateway's
running=false is authoritative and settles the session. Seed the clock +
reset turnLive in applyRewindOptimistic/applyReloadOptimistic (the
restore/edit/regenerate arm sites), and clear both on every rewind
rollback path in use-prompt-actions and session-tile-actions so a failed
rewind can't leave a stale seed.
Fixes#86795
Addresses review feedback from the hermes-sweeper (salvageability=high,
keep_open): "The new focus/visibility listener behavior lacks a runtime
UI regression test... no ProfileRail test."
Rendering the full ProfileRail component for this would drag in
drag-and-drop, dialogs, hotkeys, and i18n unrelated to what needs testing.
Instead, extracted the focus/visibilitychange wiring into its own
use-profile-rail-refresh-on-active hook, matching this exact directory's
own established convention (use-profile-prewarm.ts is the same shape:
a small side-effect hook pulled out of ProfileRail specifically so it's
unit-testable in isolation).
Added 6 tests covering exactly what the review asked for: refresh on
mount, refresh on window focus, refresh on visibilitychange while
visible, NO refresh on visibilitychange while hidden, listener cleanup
on unmount, and no listener accumulation across repeated mount/unmount
cycles.
Verified the tests have real teeth: simulated the exact bug this PR
originally fixed (dropped the cleanup return, leaving listeners attached
after unmount) and confirmed 4 of 6 tests correctly fail against it --
including "no accumulate listeners" showing 7 calls instead of 1, the
exact leaked-listener signature. Restored the real fix and all 6 pass.
ProfileRail itself is otherwise unchanged in behavior -- this is a pure
extraction (same effect, same dependencies, same cleanup), not a
behavior change. Full sidebar test suite: 93 passed across 12 files (up
from 87 across 11), 0 regressions. Python side unaffected: 158 passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent bugs let a deleted profile reappear / leave orphaned
resources on next launch:
1. hermes_cli/profiles.py's backend-process scanner required argv[0] to
resolve to an executable literally named "hermes". Electron's
pool-backend spawn resolves the hermes console-script shim's path and
execs it via the interpreter directly (python3 /path/to/hermes ...), so
argv[0] reports as "python3" and the scanner never matched the running
backend -- delete removed the profile's files but left its live backend
process running (still bound to a port via uvicorn), which
accumulates across repeated delete/recreate cycles.
2. The desktop sidebar's ProfileRail only refreshed its cached profile
list once, on mount, so a delete/create/rename from another surface
(another window, or the CLI) left a stale ghost entry until something
unrelated triggered a refetch. Note: a delete via this window's own
Manage-Profiles view already refreshes the shared $profiles atom
ProfileRail subscribes to (confirmed by reading refreshProfiles() and
handleConfirmDelete()) -- this fix only covers the cross-window/cross-
process staleness gap, not a duplicate of the already-merged
#57329's Manage-Profiles rail-refresh work.
Fix 1: recognize a python-interpreter argv[0] exec'ing a hermes-named
console-script shim via argv[1]. Fix 2: refresh the profile list on window
focus/visibilitychange, matching the existing pattern used elsewhere in
the sidebar (sidebar/index.tsx, use-background-sync.ts, star-map.tsx,
use-gateway-boot.ts all use the same focus+visibilitychange pattern).
## Related work already on main
PR #57329 (merged) fixed the *headline* symptom from issue #52279
(deleted profile respawns) via a different, non-overlapping mechanism:
routing profile-delete through the primary backend instead of spawning a
fresh pool backend, plus a separate recreation guard in
ensure_hermes_home() (#49435, merged) that makes a backend spawned into a
deleted profile's directory raise FileNotFoundError instead of silently
recreating it.
This PR is NOT a duplicate of that fix. Verified: even with both of those
merged, a backend process that survives because of gap #1 above still
holds a bound port via uvicorn -- it just can no longer resurrect the
profile directory. That's real resource-hygiene, not a symptom already
covered. Gap #2 touches a different file/component (ProfileRail /
profile-switcher.tsx) than #57329's rail-refresh half (which touched the
Manage-Profiles view's own $profiles.ts / index.tsx) and covers a
distinct staleness path (cross-window/cross-process, not same-window
delete-then-refresh).
Tests: tests/hermes_cli/test_profiles.py -- 156 passed (existing +
regression coverage for the argv[0] python-interpreter detection case).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
A user typed their root password into the Desktop SSH host field
(root@IP:PASSWORD form). Three failures compounded:
1. validateSshTarget() only checked for option injection (leading dash),
control chars, and port range — commas in an IP, whitespace ("ssh "
prefix pastes), and non-numeric ":<segment>" leftovers all dialed ssh
with garbage and failed silently five times.
2. normalizeSshConfig() only strips a ":<segment>" when it is numeric, so
a pasted password stayed glued to the hostname all the way into ssh
argv and the desktop.log connect line.
3. redactSecrets() had no pattern for ssh targets, so the password landed
verbatim in desktop.log and then in a PUBLIC debug-share paste.
Changes:
- validateSshTarget(): reject whitespace, commas, non-numeric colon
segments (with a "never put a password in the host field" hint that
does NOT echo the credential), and garbage hostnames; still accepts
bare IPv6 (::1, fe80::1%eth0). Reject whitespace/@ in user.
- redactSecrets(): new pattern masks any non-numeric segment where a
port belongs in user@host:... strings — defense in depth so future
parse gaps can't leak credentials into logs or debug shares.
- normalizeSshConfig(): strip a pasted leading "ssh " prefix.
- Tests for all three, including the exact incident shapes.
The @ popover only completed filesystem references; bot handles worked
when fully typed (mention middleware parses at submit) but were never
offered, so users had to know the exact handle — worse with multi-source
@name-device handles. Fixes#88060 (ported from Hermes-Bot-Mode#43).
- composer contrib: new 'composer.atCompletions' data area
(ComposerAtCompletionSource) — contributed rows merge AHEAD of path
results; a throwing source drops its rows, never the popover
- use-at-completions: merge contributed entries in all three fetch paths
(gateway results, gateway-less, fetch error)
- SDK: export the new area + types for plugins
- bundled Bot Mode plugin: registers 'mention-completions' — roster
handles from the query cache (\u22645s stale), active profile excluded,
'default' offered as @hermes, multi-source @name-device handles via
botHandle, display name + connection label in the row meta, capped at 8
- registered early in register(ctx) so vm harnesses reach it before the
pane/UI registrations that stubs can't fully model
The Bots pane header + is now a dropdown (New Agent / New Group Chat).
New Group Chat opens a checkbox-picker modal: searchable roster list,
member cap at GROUP_CHAT_MAX_MEMBERS, group-name input that defaults to
the selected members' names, and a Create button that assigns the
existing per-bot group meta field - so the room rides the ui_meta sync
path unchanged and the user lands directly in the new room.
The teammate-messaging protocol told bots to send with
`chat -c "Bot Chat" -Q -q` and, on 'No session found', fall back to a
manual two-step (send without -c, then sessions rename) — a dance the
CLI already made unnecessary when --create-if-missing landed (#86794).
Profiles that never went through the Bots-panel birth flow (CLI-created,
pre-Bot-Mode, remote-source) hit that miss on every first contact, and
background sends swallowed the error entirely (the original silent-drop
in Hermes-Bot-Mode#48 / #88059).
- tools/bot_mode_probe.py: protocol command gains --create-if-missing
- bundled plugin: Bot Chat prompt section + @mention handoff note use
the flag; rename-dance instructions deleted
The capability epoch hashes the protocol section, so existing eternal
Bot Chat sessions pick the new instructions up on their next message via
the established once-per-change rebuild — no per-turn cache drift.
Live-verified: fresh profile with zero sessions, protocol command
created 'Bot Chat' and delivered (PONG round-trip); second send resolved
the same session by title (no duplicate); missing-title send WITHOUT the
flag still errors loudly on stderr.
after a stale runtime-session drop (the sleep/wake 404 that resumes the
stored session and retries once), but two call sites still build their
gateway call directly instead of routing through withSessionNotFoundResume:
- session-tile-actions.ts's own cancelRun/steerPrompt/reloadFromMessage —
the tile's OWN UI handlers (wired directly by session-tile.tsx as
onCancel/onSteer/onReload), distinct from use-session-tile-delegate.ts's
interruptSession/submitToSession (used by external callers like
quick-entry-bridge), which #81261 did wrap.
- use-prompt-actions/index.ts's reloadFromMessage (the primary chat's own
"Regenerate") — it builds its prompt.submit call inline instead of going
through the shared send() helper every other action in this file uses,
so it never picked up the recovery wrapper.
After sleep/wake (the exact scenario #81261 targets), clicking Stop,
sending a steering correction, or clicking Regenerate on a tile or the
primary chat surfaces a raw "session not found" error instead of silently
resuming, even though #81261 landed the day before.
Wrap all four call sites in withSessionNotFoundResume, mirroring the
existing pattern each file already uses elsewhere (submitRewind/
syncAttachmentsForSubmit in session-tile-actions.ts, redirectPrompt/send in
index.ts) — resolve the stored session, resume once, retry, and rebind the
live runtime ref via onRecovered.
A tab/tile whose live runtime the backend reclaims (ws_orphan_reap,
idle_timeout, lru_evict) rendered an empty transcript under healthy
chrome, permanently: session.reclaimed dropped the runtime's cached
state but left the tile bound to the dead runtime id, and the tile's
resume effect is gated on !runtimeId so it never refired. Sidebar
re-click could not recover it; only close-tab or an app restart did
(tile persistence strips runtime ids, which is why a remount healed).
The reconnect-path resetTileRuntimeBindings() cannot cover this case:
the WS re-dials immediately while the orphan reaper fires a grace
window later, so the reclaim always lands after that unbind ran.
On session.reclaimed, unbind whichever tile holds the reclaimed
runtime (new unbindTileRuntime, the targeted sibling of
resetTileRuntimeBindings) so the existing resume effect refires
against the intact stored session, and purge the wiring cache's entry
so resumeTile's warm path can't hand the dead runtime straight back.
Live-reproduced both ways on an isolated dev instance (20s reap
grace): pre-fix the tile stays bound to the dead runtime with its
state gone (blank pane); post-fix it sheds the binding and repaints
the transcript within seconds, across two consecutive reap cycles.
Fixes#82620
After sleep/wake the gateway reconnect path called resetTileRuntimeBindings()
to force every tile to re-resume, but it only cleared the tile atoms'
runtimeId. resumeTile()'s warm path then re-bound each tile from the wiring
cache's stored->runtime map - the same dead pre-sleep runtime id, with a
released (empty) cached transcript. Result: every split pane except the
primary repainted as an empty pane with only its header, and prompts
submitted to it recovered into the primary view instead.
- resetTileRuntimeBindings() now also invalidates the delegate's wiring
cache (new optional SessionTileDelegate.invalidateRuntimeBindings, backed
by runtimeIdByStoredSessionIdRef.clear()) so post-reconnect resumes go
cold and bind a live runtime id.
- resumeTile()'s warm path now requires the cached state to carry a
transcript (or be mid-turn): a released/stale empty state goes through
to a real session.resume + transcript hydration instead of repainting
an empty tile.
Regression tests fail without the fix (verified via sabotage run) and pass
with it; tsc + eslint clean.
Wire the continuity flag through every cron-creation surface, not just the
model tool:
- dashboard (web/): checkbox in the cron job editor; form state round-trips
the stored reserved 'self' entry into the toggle and strips it from the
context_from textarea; web_server dashboard validator skips 'self'
(create precedes the job's existence)
- Bot Mode Routines tab (hermes-bots plugin): Continuity checkbox in the
New Cronjob dialog, forwarded through cron.manage
- tui_gateway cron.manage RPC: optional continuity param on action=add
vitest cron-job suite 10/10 (4 new), tsc app project clean, py_compile clean.
feat(desktop): resolve the user's login-shell PATH once at startup and
merge it into process.env before the backend spawns.
GUI launches (Finder/Dock on macOS, desktop launchers on Linux) inherit
a minimal PATH that never runs the user's shell profiles, so the
backend process — and everything it spawns or probes (shutil.which
availability checks like cua-driver, stdio MCP servers, Electron-side
git/gh/hermes resolvers) — cannot see Homebrew-, nvm-, pyenv-, cargo-,
or ~/.local/bin-installed tools. backend-env.ts's static sane-entry
list covers Homebrew//usr/local but not profile-added dirs.
Approach (ported from cline/cline#12429, mirrors VS Code's shell
environment resolution):
- new electron/shell-path.ts: run $SHELL -ilc (fallback -lc for the
macOS system-bash-3.2 swallow) printing $PATH between sentinel
markers so profile banners can't corrupt the capture
- merge login-shell entries first, current-only entries appended,
deduped via backend-env's appendUniquePathEntries
- single-flight, timeout-bounded, failure-hardened: a broken or slow
shell profile never blocks boot; win32 no-op
- warmed at app.whenReady, awaited before backend runtime resolution
12 unit tests + live E2E verified (GUI-minimal PATH enriched with
~/.local/bin, nvm, cargo, go entries on a real shell).
Cmd/Ctrl+Shift+B worktree flows on a remote gateway route through the
backend's /api/git mirror (hermes_cli/web_git.py), but that mirror had
drifted behind the Electron-local git ops the same UI drives locally, so
the flows broke exactly and only on remote connections:
- Convert-a-branch: the picker offers remote-tracking refs, and the
Electron op turns "origin/feature" into a local tracking branch. The
mirror ran `git worktree add <dir> origin/feature` verbatim, which
either fails or detaches HEAD. It now resolves the ref's remote via
git (never assuming "origin"), fetches best-effort, and creates the
worktree with `--track -b <short-name>`.
- branch_list omitted remote-tracking refs entirely and never set the
`isRemote` flag the renderer's HermesGitBranch contract requires —
the convert picker on a remote gateway couldn't reach a teammate's
branch and mislabeled every row's action.
- Branching off an `origin/…` base silently wired the new branch to the
remote upstream; the mirror now passes `--no-track` like the Electron
op does.
Renderer side, replace the silent degradation with a capability gate:
when a remote backend predates the /api/git worktree routes, worktree
creation failed with an opaque "Expected JSON … got HTML" toast. The
route-missing shapes now surface a clear "update the Hermes backend"
message (isGitEndpointMissingError, mirroring the sidebar batch-endpoint
detector); real git errors still pass through untouched.
Sibling audit (documented, no code change needed): repo status / review /
file-diff / git-root / default-cwd already route through desktopGit()'s
REST bridge or /api/fs on remote; repo scan is deliberately a no-op there.
Stale comments claiming "empty/false on a remote backend" in projects.ts
and coding-status.ts updated to describe the backend-routed reality.
Fixes#81724
The white catchlight dots in BotFace were static circles pinned at the
circle-face eye line (cy 16.5), while the animation clock moves the
pupils to the shape-aware eye line (cy 22 for the cloud). On the cloud
avatar the highlights floated above the eyes instead of inside them.
- Tag the catchlights (data-hb-hl-l/r) and move them with the pupils in
paintMathFace, offset upper-left of each pupil center.
- Render the initial eyes/catchlights/shut-lids at the shape-aware eye
line so the first frame matches the animated frames.
Completes PR #74468 (remote gateway headers for Cloudflare Access, #74466)
against the v2 multi-connection registry that landed after the PR was
authored, and closes the review blockers:
- connection-registry: additive optional `headers` field on remote/cloud
entries (normalized through the same forbidden-name filter, secret
envelopes like `token`); inherited on edit, treated as dial material by
connectionDialFieldsChanged, preserved by normalizeRegistry, and carried
through migrateV1ToRegistry. v2 registries without the field load
unchanged — no version bump.
- main.ts registry paths: connectRegistryBackend dials with the entry's
headers (readiness probe, ticket mint, descriptor REST via
getJsonForBackend/fetchJsonForBackend, registry ws-url minting with
rememberRemoteWsHeaders so renderer upgrades get them injected).
- saveRegistryConnection encrypts incoming plaintext header values with the
same safeStorage/allowPlainText seam as tokens; sanitizeRegistryConnection
exposes only header NAMES to the renderer — values never cross IPC.
- Connection tests exercise the leg they validate: both
hermes:connection-config:test and hermes:connections:test now send the
configured headers on the HTTP status call, the ws-ticket mint, AND the
live WebSocket probe (probeGatewayWebSocket grew an injectable `headers`
option passed as the undici WebSocket constructor's second argument).
- Settings → Connections gains an "Extra gateway headers" editor for
remote/cloud entries (name + secret value rows, stored values shown as
saved-but-hidden, clearable), with i18n keys (en + zh; other locales fall
back through defineLocale).
A dropped registered remote connection (SSH or HTTP) never recovered on
its own: the next boot attempt failed with a transient transport error
("Could not verify the existing SSH backend", ERR_CONNECTION_RESET,
mint timeout), the failure was correctly NOT latched, but nothing ever
re-attempted the boot — the renderer's reconnect machinery only arms
after a completed boot. The app parked on "Desktop boot failed" until
the user manually deleted and re-entered the same connection details,
which merely forced the fresh bootstrap an automatic retry would have
performed (issue 82679, feature ask 80430).
Root causes and fixes:
- electron/backend-start-failure.ts: new isRetryableRemoteBootFailure()
predicate — a remote, non-reauth boot failure is transient and may be
retried; local failures and confirmed 401/403 rejections are not
(a missing capability differs from a transient failure).
- electron/main.ts: the boot-failure progress broadcast now carries
`retryable` (rides with `error` through updateBootProgress), and a
failed reuse probe against a cached SSH master tears the stale
master/tunnel down so the next attempt bootstraps fresh — exactly
what manual re-entry did.
- use-gateway-boot.ts: bounded self-heal loop for a failed boot whose
progress is marked retryable — up to 5 re-attempts with the same
full-jitter backoff as the socket reconnect loop (2s base, 15s cap).
Exhausted retries end in the real boot-failure recovery overlay,
never an infinite spinner. Reset on success and on soft switch;
timer cleared on unmount.
- store/boot.ts: resumeDesktopBootForRetry() re-arms the overlay with a
retry status while an automatic retry is in flight.
Secondaries already had full-jitter backoff (store/gateway.ts); this
closes the same class for the PRIMARY/registered-connection path.
Tests: predicate matrix (retryable vs reauth-latch mutually exclusive),
plus renderer hook tests proving a transient SSH failure self-heals on
the next attempt, retries are bounded (6 total dials then the recovery
overlay, no further attempts), and non-retryable failures never enter
the loop. Sabotage-verified (disabling either half fails 4 tests).
Fixes#82679Fixes#80430
Expose the existing profile-aware gateway boot reconnect path through a
single-flight renderer action, and surface a Reconnect button in the
gateway status menu panel whenever the socket is not open. Repeated
clicks share one in-flight reconnect; failures surface through the
existing non-destructive notification UI. Localized copy for all
supported Desktop locales.
Salvaged from PR #80694 (net diff re-applied onto current main; panel
code lives in app/shell/gateway-menu-panel.tsx now).