recordSessionEventScope already captures the exact (connectionId, profile) a
runtime's inbound events proved, but knownOwnerForSession never consulted it:
with no tile/hint/row binding for the runtime id, approval.respond failed
owner resolution (SessionOwnerResolutionError) even though the event source
itself named the owner.
Add a structured owner twin of the scope ledger, written and cleared with it,
consumed as the LAST rung of knownOwnerForSession so durable stored identity
still outranks it and untagged/unknown runtimes keep failing closed.
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).
Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
Refs #86443
assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.
Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.
Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.
The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.
The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
- ambientRequestFor(gateway): one adapter for the six copy-pasted
`<T>(method, params) => gateway.request(method, params ?? {})` lambdas
the pollers built to call requestForOwnedSession.
- resetBackgroundPollingGuard() with no argument (primary reconnect,
use-gateway-boot.ts) now also clears healsByStoredId. The latch and the
heal budget share a lifetime: a respawned backend re-mints every runtime
id, so a stored session that had exhausted its 3 heals must be healable
again. Previously only the latch was cleared there.
- Leaf module comment states the actual import convention.
The approvals loop and broadcast_session_info fan out UNSCOPED session.info
frames for every live session. The event router attributes an unscoped
frame to the active session, and approvalReplaySessionId then re-pulls
approval.pending for it. When the active runtime is already gone (4001),
that made every fan-out tick a fresh dead-id request — the dominant source
of the 1,369 post-restart approval.pending rejections in #100639.
approvalReplaySessionId now takes the frame's explicitness and a gone
predicate and returns null for an unscoped replay onto a latched runtime.
An explicitly scoped frame is the runtime speaking for itself and is never
skipped.
Refs #100639
The salvaged #95647 commits predate runtime-gone.ts and shipped their own
gone-latch (session-rpc-guard.ts) beside the one main already had. Fold
them: one latch, one classifier, one clear seam.
- session-gone-latch.ts is a dependency-free leaf holding the latch, the
4001 classifier (now also rejecting "mentions session not found" tool
strings and unwrapping IPC bridge prefixes), and the rebind seam. It
exists because session-request-router — imported by every store — must
clear the latch after a successful session.resume/activate without
pulling the session/tile stores into its import graph.
- runtime-gone.ts re-exports the leaf and keeps the heal logic.
- A successful rebind now also refunds the stored session's heal budget.
markRuntimeGone caps consecutive heals at 3 per stored id and only a
successful process.list refunded it, so a backend that reaps a detached
runtime a few times left the view stuck on a phantom id with every
poller latched and the socket-reconnect global clear (removed by the
salvaged commit) re-arming the storm. #100639: 1,230 approval.pending
4001s on one runtime id in 42 minutes, zero recovery.
Refs #100639
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
session-tile-owner-route.test.ts asserted against the TEXT of
session-tile.tsx, so it passed on a broken implementation whose call site
merely looked right, and failed on this refactor, which changed nothing the
tile actually does. AGENTS.md bans the pattern outright.
Extracted the ladder as tileOwnerRoute() and replaced the three regex
assertions with six that call it: tile route wins, row falls back, hint
falls back, targetProfile carries through, a bare profile narrows away, an
untagged session stays ambient. 12s of source-matching becomes 1.1s of
behavior.
The coalescing key now carries the owner, so two backends that both expose a
session called `parent` get their own create instead of the second caller
receiving the first's child. Both creates are held open, which is the only
state the key guards — a sequential version passes even with the owner
stripped out.
Co-authored-by: Ahmett101 <ahmet.tunc@gmail.com>
TileChat re-renders per streamed token, and the owner ladder spread three
session arrays before scanning them on each one. Subscribe to the atoms it
actually reads and memoise the lookup on them, so it recomputes when the
tile store or a session list changes rather than per frame.
branchStoredSession looked its parent up in $sessions alone, and
branchCurrentSession did the same. A conversation reachable through a
profile-scoped project tree has no row there, and when it appears in both
places the flat Recents copy is the ownerless one — so the lookup returned
the row that cannot route, and the branch created its child on whichever
backend happened to be active.
cachedSessionRow spans Recents, cron, messaging and the project tree, and
prefers the self-describing candidate. One ladder, used by both branch
entry points and by resolveStoredSession.
Co-authored-by: evan-bradford <evan-bradford@users.noreply.github.com>
Routing the create and stamping the optimistic row still left a
remote-owned branch child flickering into "Couldn't open this session —
Session keeps losing its backend runtime right after resuming". The RPCs
were right and every resume succeeded; the OWNING SOCKET was the
casualty. Three gaps, one cause — nothing durable named the owner:
- openSessionTile persisted ownerRoute only for workspaceMode==='bots',
so the branch tile pinned nothing in the gateway keep-set
(openTileGatewayScopes / foregroundSessionScopes). The pruner closed
the owner socket, the backend orphan-reaped the draft runtime,
session.reclaimed unbound the tile, resume re-armed and succeeded on a
fresh socket the next recompute closed again — until the resume-storm
breaker (#93892) latched the error card at TILE_RESUME_STORM_LIMIT.
Persist the route for sessions-mode tiles whose opener knows the exact
owner, and stop a route-less re-scope from clobbering it.
- forkBranch, unlike both sibling routed creates in the same file, never
called setSessionOwnerHint/holdSessionOwnerUntilForeground — so in the
gap between session.branch returning and the tile publication landing,
no keep-set rung named the owner and a prune could reap the just-minted
draft runtime before the first prompt. Add both, mirroring the
siblings; the hold retires once the tile's own route covers the scope.
- resetTileRuntimeBindings preserved cross-connection runtimes only for
bot tabs, so a flapping sibling connection (an SSH source re-dialing)
dropped the branch tile's healthy binding on every reconnect, re-arming
resume each time — the same storm by another path. Preserve any
owner-routed tile; the owner's own reconnect still rebinds.
Also keep open tiles' rows in sessionsToKeep: a branch child is a draft
the aggregator cannot return until its first turn persists it, so the
next background refresh silently dropped the optimistic "draft: branch
#N" row and the sidebar showed no trace of the branch until first send.
Each fix verified RED by reverting its line. Live acceptance against a
remote-owned parent branched from another connection: draft row visible
immediately and stable across refreshes, tile carries the owner route,
first turn accepted and completed by the owning backend, and no
storm/error card through the full 120s storm window — before the fixes
the card latched at ~20s.
Routing the branch create to the parent's owning connection was only half the
job. The child then landed in the sidebar as a row that lied about who owned
it, so the chat pane spun forever on "draft: branch #1" and never hydrated —
the create was right, the row was wrong.
upsertOptimisticSession stamps the row's profile from $activeGatewayProfile and
omits connection_id entirely when no owner is passed (utils.ts:1318-1342), and
it also skips setSessionOwnerHint. The branch call site passed no owner, so the
child got NEITHER a row tag NOR a hint. resumeSession's owner ladder starts at
`capturedOwner || getSessionOwnerHint(storedSessionId)` and forkBranch calls it
without a capturedOwner, so the missing hint alone was enough to send the
resume to whichever backend happened to be active. Pass the parent's route as
the owner argument, restoring both mechanisms. The two sibling routed creates
in this file already did exactly this.
The tile path had the same defect one rung further out. A branch of a session
that is not the open chat opens a tile instead of resuming, and
SessionTileChrome resolved its owner from the tile route alone. openSessionTile
is called for a branch child with no workspaceScope, and session-states.ts only
persists a tile ownerRoute in bots mode, so that tile had no owner at all and
its model + composer RPCs fell back to the ambient socket. Use the same
tile-route-then-row ladder its sibling in session-tile-actions.ts already uses,
resolved per render so it cannot go stale against the tile store, the
recents/cron/messaging rows, or the hint map, with only the resulting identity
memoised on primitives.
An untagged parent row still reproduces the previous ambient behaviour exactly,
so single-connection users are unaffected.
Verified end to end against two real gateways: a session owned by a remote
connection, branched through the actual sidebar context menu in a running dev
app. The remote gateway served the create (ws closed ... messages=11
detached_sessions=1) and the resulting row polled stable at connection_id =
the remote for the full 8s window. Before the fix the same gesture produced a
row with no connection_id.
Branching a session owned by one connection while another is active created
the child on the wrong backend — or nowhere — while the sidebar still painted
an optimistic row. That row pointed at an id no backend owned, so hydration
retried, exhausted, and armed the stranded-session overlay: "Couldn't load
this session. The connection to this session failed and automatic retries
gave up." Retry re-ran the same mis-route, so it never recovered.
branchStoredSession and branchCurrentSession resolved only the parent's
PROFILE and called ensureGatewayProfile(profile), then dispatched through the
ambient requestGateway. A profile name does not identify a backend once
several connections expose the same name, so both the parent transcript read
and the session.create/session.branch RPC landed on whichever socket happened
to be active. removeSession, twelve lines away, already routed by
(connectionId, profile) via SessionOwnerScope — branch simply never got the
same treatment.
Reuse that existing contract: derive the exact owner from the parent row with
sessionOwnerRouteFromRow, activate it with ensureGatewayAgent, and dispatch
via requestGatewayForAgent. getAllSessionMessages takes the same owner scope
so the transcript read cannot silently come back empty and abort the branch
as "nothing to branch" before any create is attempted. Both arms of forkBranch
(session.branch for the open chat, session.create for a sidebar right-click)
are covered.
An untagged parent row — the single-backend case — keeps the previous
profile-only path exactly, so behaviour is unchanged for users with one
connection.
Tests: three call-site regressions asserting the create rides the owning
(connection, profile) socket, that the transcript read carries the same owner
scope, and that an untagged parent still uses the ambient socket. Plus an
integration test that mocks nothing inside the router — the real
requestGatewayForAgent runs against a fake Electron bridge and transport, so
a regression that re-collapses a registry route onto the ambient socket fails
even if the call-site assertions still pass.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
The everything-flow's client leg read `$updateStatus.get() ?? await
checkUpdates()`, so a cached row always won. That row can be up to a
poll interval (30 minutes) old and is captured before the backend leg
runs, so a cached "already current" skipped the client apply entirely —
the stale-GUI gap the flow exists to close.
Re-check first and fall back to the pre-flow snapshot when the live
check can't answer. `checkUpdates()` resolves with an error status
rather than rejecting and overwrites the atom with it, so the snapshot
is taken before any leg runs.
Co-authored-by: Dhana <227747512+whoisdhana@users.noreply.github.com>
The update entry points chose their target from the connection mode, so
every surface in remote mode acted on the backend — including the ones
showing the client's own status.
The macOS "Check for Updates…" app-menu item is the clearest case: it
sits next to "About Hermes" and is the OS-standard way to update THIS
app, but on a Mac connected to a remote Linux backend it checked the
Linux box. The backend was already current, so the action reported
nothing and did nothing, and the desktop app drifted months behind with
no error and no updater log to explain it. The update-available toast
had the same split: a client check raised it, clicking it opened the
backend's overlay, which has no target switcher and no way back.
Surfaces bound to one target now name it; only genuinely generic
commands still take the connection-mode default, so the command
palette's remote-mode backend target and the everything-flow are
unchanged.
Co-authored-by: BerneYue <14088768+yuexiongHNU@users.noreply.github.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
Co-authored-by: clayduncan <234173110+clayduncan@users.noreply.github.com>
Co-authored-by: Dhana <227747512+whoisdhana@users.noreply.github.com>
When the bundle was swapped under a running process, the About banner sent the
user to the installer — a download and a reinstall for a state that a plain
restart repairs, and the reason reinstalling never helped these reports.
Report bundleSwapPending on hermes:version and give that case its own copy and
a "Restart Hermes" button. It gets its own headline too: reusing "App build out
of date" over a body that says the app is already installed repeats the
contradiction with the Updates card that the banner is supposed to resolve. The
installer link stays for the genuinely-stale-bundle case.
Packaged builds only — a dev `--build-only` rewrites the stamp under a running
`npm start`, and that is a rebuild the developer asked for, not a torn install.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
A user who reopens Hermes while an update is running lands on the boot gate,
which is what it is for. But the updater swaps the packaged bundle on disk
after `hermes update` exits, and its `open` leg only focuses this already-
running process, so nothing ever loads the new build. The parked instance then
passes the gate and boots the new runtime under the old renderer — the "App
build out of date" banner immediately after a fully successful update, over an
Updates card that says "You're on the latest version" and so offers no remedy.
Compare the install stamp this process loaded at boot with the one on disk when
the gate clears. On positive proof of a swap — different commit, or a different
builtAt at the same commit — relaunch instead of starting a backend. Detection
fails quiet like bundle-skew, so a swap that never happened (the Windows
locked-binary case) is unchanged. A one-shot argv flag makes a relaunch loop
impossible and a 15s failsafe falls back to the old behavior.
Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
Co-authored-by: aeonsong <aeonsong@users.noreply.github.com>
detectBundleSkew() trusted `git rev-list --count <stamp>..HEAD -- apps/desktop`
outright, which claims skew in two states where the install is not torn.
Ancestry: `A..HEAD` only measures how far HEAD is ahead of A when A is an
ancestor of HEAD. A ZIP-fallback update rewrites the tree onto a synthetic
root, so the stamp still resolves but is unreachable; the range degenerates to
HEAD's own history and reports a permanent >= 1 while apps/desktop is
byte-identical. Ask `merge-base --is-ancestor` first and go quiet unless it
answers yes.
Scope: the pathspec counted every file under apps/desktop/, so a docs- or
e2e-only commit produced a banner promising missing UI features that do not
exist. Count only the paths that reach the shipped app.
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
Co-authored-by: kokhlo <kokhlo@users.noreply.github.com>
The two #57911 regression tests hand-encoded the remote workspaceCwdKey()
localStorage literal; seed it through the exported setCurrentCwd() like the
sibling isolation test so a key-shape change can't silently strand them.
(/simplify-code reuse finding)
The remote remembered cwd is no longer a bare-new-session default after
the #57911 fix; the comment listed it as priority 2. Follow-up from
salvage self-review of PR #57926.
A bare-new-session (Cmd+N without a project scope) in remote mode used to
inherit the remembered cwd under workspaceCwdKey()'s remote variant — i.e.
the last project the user attached to on that backend. Symptom: Cmd+N landed
on the wrong project's workspace (e.g. a configuration discussion thread
opened in the tradingview directory).
The remote branch in workspaceCwdForNewSession is the entire bug class:
- It overrides the configured-default-pre-attaches rule, so users who
configured an explicit default didn't get it under remote mode either.
- The bare-new-session path has no notion of "remember where I was" —
that's startSessionInWorktree's job. The remembered cwd is only meant
to assist resume/restore (ensureDefaultWorkspaceCwd, which keeps its
remote-keyed sticky seed and is unaffected by this change).
Drop the mode === 'remote' early return and route both modes through
getConfiguredDefaultProjectDir(), matching the proposed fix in the issue.
Test: the existing 'keeps remote workspace memory separate from local and
other remotes' case set the local key then switched connection to remote,
so under the buggy code it returned '' regardless — it never gated the
regression. Add a remote-key repro and a symmetric explicit-default test
so the fix is observable from the suite.
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
_migrated:true files were skipped on later boots, so #100576 installs
stayed stuck on the named profile. Re-evaluate those files only; leave
user-selected pins (no _migrated) alone. If default now wins, write
{profile:null} instead of pinning default.
First-boot migrateActiveProfileIfMissing only listed ~/.hermes/profiles/*
and scored profiles/<name>/state.db. Default's real DB is ~/.hermes/state.db,
so a tiny named profile could be pinned after an update.
Always candidate default, score/pid-check it at HERMES_HOME, and do not
write active-profile.json when the winner is default.
Fixes#100576
The scripted turns execute real terminal commands, and the sidebar
sentinel-wait loop trips the dangerous-command guard: the turn parks
behind a Run/Reject approval card, and the default 'smart' mode fires an
aux LLM approval call at the same mock provider — consuming a
scripted-turn index and never resolving. On the slower CI runner this
stalled the sidebar-dot family (sidebar-states 157/245, tile-unread 166)
until spec timeout; run 33543723331's error-context snapshots show the
approval card blocking each stalled turn. Locally the race usually won
the other way, which is why these passed on dev machines.
Fix: fixtures write 'approvals: mode: "off"' into the mock provider
config by default (specs supplying their own approvals: section own it),
mirroring the auto-title default. Also drop the DOT-DEBUG diagnostics
from tile-unread-bug now that the root cause is identified.
Local: sidebar-states + tile-unread + correction-session-switch all
green in seconds (3-9s vs 90s timeouts); full suite 62 passed /
11 skipped / 1 flaky-passed.
The Desktop E2E lane was disabled Aug 2 – Sep 1; the app and gateway kept
moving, so 16 specs rotted against current main. All failures traced to
spec/harness drift, not product regressions:
- fixtures.ts: title generation now rides the main model (#83636), firing a
background completion at the mock after every turn — it contains the whole
conversation (trigger keywords included), advancing scripted-turn indices
and tripping hold-for-prompt matchers. Disabled by default in the mock
provider config; specs supplying their own `auxiliary:` section own it.
- chat/interim-messages/session-compression/correction-session-switch/
hidden-history-messages: busy-state and transcript assertions updated to
the current composer aria-labels, interim-message semantics, and
verify-on-stop continuation behavior on main.
- bot-mode-closed-chat-stays-closed/group-to-local-bot-handoff: Bot Chat tab
selectors updated for the Bot Mode rework (tabs keyed by
connection+profile, renamed tab triggers).
- glyph-spinner: assertions made compositor-honest for the CI runner
(steps() keyframes + layer promotion probed via the animation registry
instead of GPU-dependent screenshots).
- sidebar-states/tile-unread-bug: event-driven waits with mock-server
release handles replace wall-clock polls that lost races on loaded
runners.
- warm-resume-jitter/image-attachment-resume: real-session-builder harness
waits for the thread viewport before evaluating; failure path now dumps
per-surface pane state.
Local full-suite run on the CI-equivalent xvfb setup: 62 passed,
11 skipped, 1 flaky-passed (correction-session-switch live-correction spec,
passes on retry). No product code changed.
The 500ms setInterval in HudShell's band-measurement effect polled
forever, contradicting its own comment ("poll briefly until it exists,
then let the ResizeObserver own it") — the viewport was never actually
checked, so the timer never cleared. It kept re-running measure()
(DOM queries + getBoundingClientRect + a style write) every 500ms for
the life of the HUD window, one of several sustained per-window timers
reported in #98394 as sustained idle renderer CPU / repeated re-renders.
Extracted the effect into useHudTranscriptBand() (matching the
existing per-concern hook split in this file: useHudGlass,
useHudClickThrough, useHudThreadFocus) and made the interval check for
the viewport before re-measuring, clearing itself once found so the
ResizeObserver takes over as the comment always said it would.
- curly: brace all single-line if statements in profile-migration.ts and
profile-migration.test.ts (17 errors in CI check:lint)
- perfectionist/sort-imports: node:fs builtin import before vitest external
- padding-line-between-statements: blank lines after block statements
- prettier: normalize formatting (fmt script style) in the three touched files
All 967 electron project tests still pass.
Polish from antigravity review of the rebase resolution (GPT-OSS):
the previous comment said "BEFORE the first primaryProfileKey() /
primaryBackendIsRemote() read" but those two calls live at different
points — primaryBackendIsRemote() is the very next line, primaryProfileKey()
is inside the connection IIFE. Be explicit about which is where so a
future reader who moves one of them knows what to preserve.
Addresses teknium1's review (#64195) finding #2: the multi-rung resolver
needs Electron tests covering precedence, stale-PID rejection, fallback
behavior, and the remote boot path. The pure decision helpers are now
covered by 29 unit tests in `profile-migration.test.ts` (vitest electron
project).
Coverage:
- precedence: legacy > single-running-gateway > state.db heuristic
- stale-PID rejection: recycled PIDs not owned by hermes are dropped
- malformed pid files: JSON parse errors, non-integer PIDs, zero/negative
- scoring edge cases: ancient files (recency floored at 0.1), tiny files
(size floored at MIN_SIZE), larger DB beats smaller at similar recency
- single-profile fallback: best === 'default' suppresses the write
- no-op cases: preference file already exists, missing profiles root
The remote boot path is verified by code review of the call-site move
(commit preceding this one) — `migrateActiveProfileIfMissing()` now runs
before `primaryProfileKey()` is first read in `startHermes()`.
The pure decision logic that the orchestrator relies on is covered end-
to-end below; this matches the repo's testable-helper pattern (see
`profile-delete-routing.test.ts`).
Addresses teknium1's review (#64195) finding #1: the previous PR placed
the migration inside the connection IIFE, AFTER
`resolveRemoteBackend(primaryProfileKey())`. When the preference file
was missing, `primaryProfileKey()` resolved to 'default' and the remote
branch returned immediately without ever reaching the migration. Remote-
mode users got no migration at all.
Move the call site to the top of `startHermes()`, before the connection
IIFE that reads `primaryProfileKey()`. Both remote and local branches now
flow through this path before any profile-dependent resolution, so the
migration runs on first boot regardless of mode.
The inlined implementation is replaced with a thin wrapper that builds a
`MigrationDeps` bag and delegates to `migrateActiveProfileIfMissing` from
`profile-migration.ts`. No production behavior change beyond the call-
site move.
Tests added in a separate commit.
Addresses teknium1's review (#64195) — the migration decision logic should
be unit-testable without Electron. Pull the ladder (legacy sticky file,
running-gateway scan, state.db heuristic) into pure helpers in a new
`profile-migration.ts` module, following the dep-injection pattern
already established by `profile-delete-routing.ts`.
The helpers take an injected `MigrationDeps` bag so tests can exercise
precedence, stale-PID rejection, fallback behavior, and the single-profile
case without touching `/proc`, `ps`, or the host filesystem. The default
profile is explicitly rejected from the legacy rung because the regex
matches `default` and accepting it would suppress the heuristic that is
the whole point of the migration.
The wrapper in main.ts is unchanged in behavior — the same `MigrationDeps`
fields get filled in from `fs`/`path` and `isHermesProcess`. The atomic
write + parent-dir-create that the wrapper performs matches
`writeActiveDesktopProfile`'s semantics so the migration produces a file
indistinguishable from a user-driven profile switch.
No production behavior change; pure code organization.
Tests added in a separate commit.
When active-profile.json does not exist (fresh install or first boot after
update), seed it from the best available signal so the Desktop launches
into the user's primary profile instead of always defaulting to "default".
Priority ladder:
1. Legacy ~/.hermes/active_profile (explicit CLI choice via hermes profile use)
2. Running gateway (gateway.pid with verified liveness + hermes identity check
via /proc/cmdline or ps -o args= to avoid PID recycling false positives)
3. state.db heuristics — hybrid recency×size score picks the primary workspace
(e.g. a 409MB coder DB beats a 28MB default DB even if touched at similar times)
The stored JSON includes _migrated:true for priority 3 (heuristic guess) so
the renderer can optionally surface a one-time notification. Priority 1 and 2
are higher-confidence signals and skip the flag.
The migration is a no-op once active-profile.json exists, and only writes
when a non-default profile is confidently identified — preserving the legacy
fallback-to-default behavior for single-profile users.
Fixes#64160 (active-profile half).
The local expires_in countdown killed OAuth sessions with a bare
"Session expired" before the backend poller's enriched message (Portal
sign-in stalled in the opened tab, retry/API-key fallback) could reach
the UI, and the desktop onboarding poller had no local expiry at all —
a dead session polled forever. Both surfaces now lapse with guidance
naming the common cause, prefer the backend error_message when it has
one, and keep polling when the backend still reports pending (clock
skew).
- Translate leftover English 'toolset(s)' strings in ru.ts
- Cover ru aliases/config value in languages.test.ts (parity with ar)
- List Russian among desktop UI languages in website/docs/user-guide/desktop.md
Full desktop UI translation (ru.ts, 3076 string/function leaves,
mirrors en.ts 1:1) plus locale registration in types, catalog and
language list with aliases (ru, ru-ru, ru_ru, ru-by).
Russian plurals (1 / 2-4 / 5+, 11-14 exception) via RU_PLURAL/RU_NOUN
helpers; count accepts number | string to match en.ts signatures.
Verified against current main: structural validation 3076/3076 with
placeholder parity, typecheck (renderer/electron/e2e) clean,
i18n vitest 28/28, prettier + eslint clean.