Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
The sidebar reports a profile it could not scan as HTTP 200 with an empty
page and errors=[{profile}]. The renderer merges that page keeping only
working, pinned, and selected rows, so every idle Yesterday / This-week
session disappears until a later scan succeeds — and the 5s coalescing cache
then serves the same empty payload back for the rest of its TTL.
Carry the previous rows forward for exactly the profiles named in errors[],
keyed by profile::id so a twin id in another profile is never stitched in.
Profiles that scanned cleanly are still authoritative, so a genuinely empty
page with no errors still clears the list. Per-profile usage and truncation
flags follow the same rule rather than zeroing under a list that was kept.
The legacy per-slice fallback stamps errors on the slice that actually
failed, so a cron read failure can no longer blank recents.
Part of #73847
Part of #88528
Co-authored-by: AKAZIK-py <AKAZIK-py@users.noreply.github.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.
- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
install-state spinners off $hubActions, and an OfficialSkillDetail pane
(hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
Session mutations from the desktop sidebar silently no-op against the wrong
profile's state.db and reappear on the next refresh, unless the mutated
session happens to belong to the serving (primary) profile. Most visible in
the "All Profiles" view and on a remote-primary desktop (a registered remote
gateway as primary, every profile served from it): deleting a session flashes
away optimistically, then comes back after a profile switch.
Root cause: the mutation helpers passed the owning profile ONLY as
request.profile, which the Electron main process consumes for backend ROUTING
but which does not scope the request the backend actually receives. The
backend selects its target DB from ?profile= (DELETE) or body.profile (PATCH).
Reads (getSession, getSessionMessages) and renameSession already scope
correctly; the other mutations regressed / never did:
- deleteSession -> no ?profile= in the URL
- setSessionArchived -> no profile in the PATCH body
- setSessionPinnedRemote -> no profile in the PATCH body
- setSessionUnreadRemote -> no profile in the PATCH body
On a remote gateway whose connection has no remoteProfile alias, the main
process leaves such requests unscoped, so the backend opens its own (default)
state.db, cannot find another profile's row, and returns
{ok:true, already_absent:true} (DELETE) or no-ops (PATCH) — a fake success the
UI treats as done. deleteSession lost its ?profile= scope in the api/ module
split (it was present via PR #44138 / the pre-split hermes.ts path).
Fix: scope all four mutations the same way the working endpoints do —
deleteSession appends sessionScopeQuery(profile) to the URL; setSessionArchived
/ setSessionPinnedRemote / setSessionUnreadRemote include profile in the PATCH
body (mirroring renameSession). request.profile stays for per-profile
remote-override and global-remote routing. Single-profile / selected-profile
users are unaffected (the serving profile already matched).
Verified against a live remote-primary desktop: the DELETE now goes out as
/api/sessions/<id>?profile=<owner> and the row is actually removed from the
owning profile's state.db (confirmed server-side) instead of returning
already_absent.
Adds api/sessions.test.ts coverage: delete scopes ?profile= in the URL for
object and bare-string owners and omits it when no owner is known; archive /
pin / unread carry body.profile when owned and omit it otherwise.
Fixes#78836
Users reported no GUI switch for browser.use_real_profile — the only
desktop home was the generic Settings → Config editor, which nobody
found. The Browser toolset detail pane now renders a 'Use My Real
Browser Profile' ToggleRow above the backend/provider matrix.
- new BrowserRealProfilePanel: reads the shared profile-scoped config
record cache, optimistic write-through, rollback on failure
- saveHermesConfigRecord: capability-scoped PUT /api/config counterpart
of getHermesConfigRecord, so the Capabilities scope selector writes
the profile it points at (possibly another gateway)
- i18n: en/ja/zh/zh-hant keys (ar inherits en via defineLocale)
- docs: browser.md desktop pointer corrected to the real location
Live E2E on the built app over CDP: clicking the switch flipped
browser.use_real_profile true→false→true in the sandbox HERMES_HOME
config.yaml, GET reflected it, no layout glitches (screenshots in PR).
Since 1e9a12a71 (v0.20.6), a remote/cloud/ssh registry PRIMARY makes
globalRemoteActive() true, so the ambient v1 route — and with it every
unpinned Capabilities read — resolves to the remote gateway. The only
road back to this machine is an explicit connectionId:'local' pin, but
capabilityScoped() deliberately DROPPED that pin (pre-#91564, absent id
always meant the local pool), and profileScopeKey() collapsed
'local::<profile>' to the bare profile key.
Consequences on any desktop whose registry primary is remote:
- Capabilities -> MCP showed the REMOTE host's mcp_servers under every
scope; locally configured (and connected) MCP servers vanished from
the UI entirely.
- Picking 'default - This device' in the scope selector was a silent
no-op: the collapsed cache key equaled the current scope key, so
changeScope() early-returned and the selector snapped back.
Fix: capabilityScoped() forwards EVERY non-empty connection id, 'local'
included — Electron's apiRequestRegistryConnectionId/ensureRegistryBackend
already own 'local' pins (forced-local pooled child) and this is the
documented contract there. profileScopeKey() namespaces every explicit
pin ('local::<profile>') so a This-device pick and the ambient path never
share a cache row. profiles.ts profileOwnerScoped, which hand-patched
this exact hole for profile mutations, reduces to a named alias.
Live A/B (Electron + CDP, remote registry primary, local-only servers in
local config): v0.20.6 = local servers invisible, local pick no-op;
fixed = 'This device' lists local-server-alpha/beta, remote scope
unchanged.
The fail-closed owner ladder (#95407) is correct for new sessions, but
legacy unowned rows on registry-topology installs dead-ended in
SessionOwnerResolutionError (reporter's Error B) with their transcripts
fully intact in state.db.
- resolveLegacyOwnerBackfillScope: pick the single-match store for the
server-side owner backfill at enumeration time (serving registered
connection / primary pool); fail closed on multi-candidate topologies.
- maybeBackfillLegacySessionOwners: one-shot per scope per renderer,
fire-and-forget from the #95407 stamp path, logs the stamped count.
- Read-only stored-transcript resume: when session.resume fails closed,
fetch the transcript over id-only REST (ambient first, then registered
backends, read-only probes only) and open the session as a read-only
transcript instead of dead-ending; sends are refused with a notice and
a later successful live resume clears the latch. Wired into the main
pane resume recovery and the session-tile delegate (which now runs the
same fail-closed owner gate as the RPC dispatcher).
Refs #94724
20s withTimeout() on the boot() and soft-switch paths (use-gateway-boot.ts),
since these are IPC round-trips into the main process with no timeout of
their own — a wedged main-process round-trip hangs the awaiting caller
forever instead of surfacing a failure.
Every other production call site of the same IPC pair was still unbounded:
- store/gateway.ts's openSecondary() and sharedPrimaryRoute() — the actual
connection-establishment underneath requestGatewayForProfile/Agent,
ensureGatewayForProfile/Agent, and every other exported routing entry
point that opens a non-primary profile's socket.
- use-gateway-request.ts's on-demand reconnect (the primary gateway's
"not connected" retry path hit by every RPC).
- voice-playback.ts's resolveSpeakStreamUrl().
- api/plugins.ts's activeConnection() (pluginSocket's connect()).
Extracted RECONNECT_ATTEMPT_TIMEOUT_MS into the shared lib/with-timeout.ts
(previously local to use-gateway-boot.ts) so every call site uses the same
budget instead of duplicating the constant.
Regression tests mirror the existing use-gateway-boot.test.tsx hang-repro
pattern: wedge getConnection()/getConnectionFor() with a never-resolving
promise, advance fake timers past the 20s bound, assert the caller settles
instead of hanging. Mutation-verified: reverted the production fix (kept
tests) and confirmed the 6 new tests fail — 4 by genuinely timing out at the
vitest level, 2 by TypeError on the not-yet-exported activeConnection —
restored the fix and confirmed all 70 tests across the gateway/voice/boot/
plugins suites pass, with tsc -p . --noEmit clean throughout.
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.
- Every registered gateway's profiles sit on the one strip, in registry order
(This device first, then by label), each group headed by that gateway's
kind glyph. The active gateway's squares are unchanged; the others are
"at rest" (dimmed) with tooltips/accessible names qualified by machine
(`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
the statusbar switcher, landing on that exact (gateway, profile):
`selectConnection(id, { profile })`. The spinner sits on the clicked
square; the previous source stays painted until the target answers.
Groups keep their slots whichever gateway is active, so a square never
moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
SOUL.md / Delete, executed on the owning gateway (renameProfile,
getProfileSoul and updateProfileSoul accept the same scope deleteProfile
already had); the delete confirmation names the machine. The legacy
per-profile "Connect to a remote host…" item is hidden on multi-gateway
setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
two registrations of one backend collapse to one group; past thirteen
squares across the fleet the strip condenses into a menu sectioned by
gateway. Roster is fetched on mount / focus / registry change only — no
periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.
Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.
Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.
Docs: multi-connection-desktop.md describes the fleet rail.
Refs #89304, #92384, #91047, #94724
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reconcile #94901's API-layer row stamping with #94656's durable-owner
persistence: extract lib/session-owner-stamp.ts as THE canonical
stamp-untagged-rows write path (never clobbers an explicit owner,
never stamps `local`) and re-express api/sessions'
stampActiveConnectionOwner through it. #94656's writers (optimistic
row from the captured owner route, mergeSessionPage carry, cache
patch) are exact-owner writers and stay as-is; the helper's contract
documents why it must not overwrite them.
Credit: row-stamping concept from PR #94901 (joe-rodgers) and
PR #95007 (weismanfamily); persistence shape from PR #94656
(Zeus-Deus).
Co-authored-by: joe-rodgers <25499388+joe-rodgers@users.noreply.github.com>
Partial cherry-pick of PR #94901 (joe-rodgers). Surviving scope:
- api/sessions: stampActiveConnectionOwner — rows returned by the
active non-local gateway are stamped with its registry connection_id
(explicit owners from multi-source responses preserved), so a later
resume cannot fall back to a same-named local profile.
- store/projects: one-shot projects.tree retry when a remote source
switch leaves the first read RPC on a newly-opened socket without a
response (request timed out / gateway connection closed), only while
the same gateway/profile is still foreground. Component fix for the
live-confirmed #92352 sidebar-never-paints gap.
Dropped scope (superseded on main / by the #94656 anchor landed just
below): knownSessionOwner+SessionOwnerScope rewiring in session.ts,
session-states.ts, wiring.tsx (main and #94656 carry richer variants),
and the $connection-derived optimistic-row stamp in
use-session-actions/utils.ts (#94656 stamps the optimistic row from
the captured exact owner route instead of ambient state).
Original-PR: #94901
Dropped-scope: routing half of 2cb5bdbf1 (session.ts, session-states.ts, wiring.tsx, use-session-actions/utils.ts hunks)
Sidebar and legacy session-list helpers tagged the registry connection but
not the active profile, so Electron routed those reads to the wrong backend
after a profile or remote switch.
Keep hermesApi connection-only: stamp profileScoped on the list helpers
instead of every REST call.
Co-authored-by: noah <loahnisk@gmail.com>
Preserve the connection-bound Bot and session routing implementation while
adopting upstream's modular Desktop API split and all changes through
b2057c168.
2,248 lines of gateway REST client become twelve modules by domain, with
hermes.ts left as a barrel so all 144 importers stay put. The import
graph is a star — every domain module imports only ./client, and client
imports nothing back — so there are no cycles.
The barrel names client's public exports rather than re-exporting it
wholesale. Splitting a module forces its private helpers into exports so
siblings can reach them, and export * would then republish them:
profileScoped, connectionScoped and capabilityScoped were private to
hermes.ts and have to stay that way, or a call site can assemble its own
request scope and drift from the api layer.