setBackgroundMaterial on a transparent window permanently kills per-pixel
alpha on Win11 — every transparent pixel composites as opaque white, so the
HUD showed a white slab instead of the desktop behind it. Verified against a
minimal repro on Electron 40.10.2: the break happens with ANY material value
including 'none', which is exactly what the idle HUD asks for, and neither
'auto' nor a follow-up setBackgroundColor('#00000000') restores it.
The DWM backdrop and window transparency are mutually exclusive, so the
Windows HUD keeps the CSS tint its sheet already paints and skips the native
frost. macOS is untouched: setVibrancy composites correctly.
window-below asked node_modules for get-windows, whose lib/windows.js locates
its native binding through preGyp.find() — by HOST platform. When the tree was
installed on one OS and Electron is running on another (a WSL-hosted dev run
driving a win32 Electron), pre-gyp picks the host's slot, ignores the correct
binding sitting beside it, and upstream's fail-soft path returns no-op stubs.
Enumeration then reports 'unavailable' on a machine that answers perfectly
well, which silently disables read_window_below.
scripts/stage-native-deps.mjs already writes a staged lib/windows.js that
requires its binding directly, so prefer it and keep the bare import as the
fallback.
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:
- Desktop relay drain forwards bot_relay.deliver's error.data.reason
into bot_relay.reply (and prefers it for the attention badge over
free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
text, so the completion notification the sending agent receives is
machine-branchable.
Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
ensureGroupChatSession's resume loop caught ANY session.resume error
(stored sid, then title lookup) identically and fell through to
session.create — the same bug findExistingCanonicalChat was fixed for
hours earlier (87b645f52c) in the same file: a transient failure (the
backend still warming up after a restart, a network blip on a
cross-connection lookup, an oversized-resume refusal) read as "no
session, mint a new one". That forks the member's real session AND
silently overwrites room.sessions[key], making the original
unreachable from the room. ensureGroupChatSession is actually more
exposed than the 1:1 case: it runs every group turn
(runGroupChatMemberTurn), with two independent swallow points.
Distinguish "genuinely doesn't exist" from "transient failure" the
same way the gateway itself does: session.resume's own handler
(tui_gateway/methods_session.py) returns JSON-RPC code 4007 only when
the target truly isn't found; every other failure (including 4130,
"session too large to resume" — a session that DOES exist) now
surfaces instead of being silently swallowed. The existing outer
try/catch at the call site already treats a thrown error as "this
member passes the round" (recordGroupActivity kind: 'failed'), so
nothing new needs to catch it — a transient hiccup now costs one
skipped round instead of a permanent fork.
The toggle was gated on having 2+ registered sources, which hid it in exactly
the local-only state the drift produces — the state where a user most needs to
change what launch restores.
Reconciliation repairs the drift at its source, but it can still fail to
persist (read-only or full userData), which leaves a window live on a source
the registry cannot name. $activeConnectionId is null there, the preferred-id
guard misses, and the restore re-homes a working connection.
Return early when a connection is live but unnameable. The registry has no
claim on a source it does not know about.
migrateV1ToRegistry runs exactly once, only when connections.json is absent.
A user who was local at that moment and pointed Settings -> Gateway at a
remote afterwards gets a live remote the registry cannot name: the descriptor
resolves to no connectionId, primary still says 'local', and the boot-time
launch pick force-switches the window onto a fresh local backend seconds after
the sessions list paints. That backend has no provider, so onboarding pops.
Reconcile on read: when the v1 global route names a remote with no matching
registry entry, register it and adopt it as primary/last-used, then persist so
the repair happens once. Narrow on purpose — an already-registered route is
left alone even when primary names something else, because that is the user's
pick in the Connections panel, not drift.
Replaces the hand-edit-connections.json workaround users have been trading.
The showAllProfiles browse-mode flag is persisted to localStorage, but
every restart it was force-collapsed anyway: initializeConnectionsRegistry
restores the last-used source via selectConnection, and selectConnection's
post-activation path unconditionally ran $showAllProfiles.set(false).
That collapse is correct for a user click on the connection picker (a
concrete-source action), but the silent boot restore is not a user action.
Gate both reset sites on pendingTarget === null && activeConnectionId ===
null (the fresh-boot state) so the persisted preference survives restart,
while any user-initiated switch still collapses browse mode.
Regression tests cover both directions: boot restore preserves true, a
user switch collapses it.
Fixes#93197
The salvaged hardening tests matched main.ts source text to assert that
fetchConnectionStatus reaches for a bearer and that Apply preflights before
persisting. A rename breaks them while a real auth regression that keeps the
substrings passes.
Make the preflight a first-class option on applyConnectionConfigAtomically so
its ordering is observable, and assert it through the seam: preflight runs
before either write, and a rejected preflight leaves both stores and the
activation untouched.
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.
- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
_AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
of fallback-providers and the delegation /review section missed in
#93339
CI sibling-test blast radius from the cluster branch:
- singleFlightSessionResume crashed on run() doubles that return
non-promises (Cannot read 'finally'); wrap via Promise.resolve().then(run).
- Three use-session-tile-delegate tests pinned the pre-#92961 ambient
dispatch for default-profile sessions; the routing-authority change
intentionally routes every known owner through the profile router, so
the tests now assert requestGatewayForProfile('default', ...) instead.
Fixes#85834 (Electron REST intercept fall-through). The
/api/sessions/{id}[/messages] intercept in electron/main.ts required an
explicit ?profile= (or request.profile) to route a read to its remote owner;
callers without a hint fell straight through to the LOCAL backend and 404'd
on its state.db even though the session lives on a configured remote — while
the list endpoints happily showed the row (remoteSessionList tags s.profile).
When no explicit profile resolves, consult the same remote session lists the
list endpoints use to find the owning profile (matching id or lineage root
id), memoized for 30s so a transcript+messages burst costs one sweep. Only
when the id is genuinely unknown remotely does the request fall through to
local, exactly as before. Pure lookup lives in profile-session-routing.ts
with unit tests (owner hit, lineage-root match, null on miss/dead
remotes/no remotes).
Maintainer commit (cluster salvage).
Client half of #91684. The approval bar (approval.tsx) and the native
notification action path (native-notifications.ts) sent approval.respond on
the AMBIENT gateway socket. Ambient follows foreground focus; for an approval
raised by a cross-profile or tile-owned session it points at a backend that
never held the approval, so Run/Reject silently failed after a profile swap
or reconnect.
- New knownOwnerForSession/requestForOwnedSession in store/session-states.ts:
resolve the owner sync (tile owner route -> known session profile via row or
open-time hint; runtime ids translated to stored ids first) and dispatch via
requestForSessionProfile. Ambient only when no owner is known — never a
fall-back to "active".
- approval.tsx and native-notifications.ts respond through it, binding the
ambient dispatcher so the no-owner path keeps the exact 2-arg call shape.
- Tests: owner resolution (tile route first, row-profile fallback,
undefined for unknown/null) and ambient arity preservation in
session-states.test.ts; existing approval + native-notification suites
still pass unchanged on the ambient path.
Maintainer commit (cluster salvage).
After sleep/wake or a reconnect, many surfaces discover the same dead runtime
at once (submit recovery, slash/rewind recovery, tile resumes, the target
resolver, session switch) and each fired its own session.resume — the gateway
minted a runtime per call and the losers fed the orphan reaper (#91276 storm).
- New use-prompt-actions/single-flight-resume.ts: module-level in-flight map
keyed by storedSessionId; all resume call sites (utils.ts recovery, submit.ts
direct rung, resolve-target-session.ts, use-session-actions switch resume,
use-session-tile-delegate resumeTile) share one in-flight promise per stored
id. Failed flights are not cached.
- Drift-abort paths no longer abandon a freshly-minted runtime: utils.ts
SessionRecoveryAborted and submit.ts post-routed-resume / post-resume aborts
register it in a stored->runtime recovery cache; the next action for that
stored session adopts it (via onRecovered) or reuses it instead of minting
another. Cache entries are take-once and skip a known-dead id.
- Unit tests: one RPC for two concurrent callers of the same stored id,
drift-abort registers (not strands) the recovered runtime, independent
stored ids resume independently, cached-runtime adoption.
Maintainer commit (cluster salvage, part of the session-not-found-after-
reconnect consolidation).
Follow-up to PR #91357 (salvaged, author enwaiax): the committed #90428
explicit-target regression fixture started foreground B with a valid active
runtime and a positive B->runtime cache entry, so routedSessionNeedsResume was
false and the formerly broken foreground-recovery branch was never exercised —
the test passed even on the broken head (b9df1f9c2).
Strengthen it per the review: B now starts with activeSessionIdRef null and an
empty ownership cache, resumeStoredSession(B) fully publishes B's runtime and
cache binding, and the assertions still require no high-level resume of B, an
authoritative session.resume(C), exactly one queued prompt.submit to C's
recovered runtime, and no mutation of foreground refs/cache.
Salvaged-from: PR #91357 (author enwaiax); fixture hardening by maintainer.
Step 2 of removing 'active gateway' as a routing input. A session's backend
is a property of the SESSION (its profile), never of whatever the window is
currently showing. The active-profile fallback was the root cause of Bot Mode
'session not found' / hangs: a hidden/unlisted session with an unknown owner
was silently dispatched to the active profile's backend, which never owned it.
- sessionRpcNeedsProfileRoute: drop the active-profile comparison entirely. A
KNOWN owner (route or profile name) ALWAYS routes to its own profile's
socket; only a null/empty owner (fresh draft, global chrome) routes ambient.
A primary-profile owner collapses back to the primary socket inside
gatewayForProfile, so the reauth-aware reconnect path is unchanged.
- session.ts: split knownSessionProfile (row -> hint, undefined when unknown)
out of rememberedSessionProfile. rememberedSessionProfile keeps its active
fallback but is now documented as PRESENTATION-only (navigation keying),
never routing.
- wiring requestGateway: resolve the owner from the tile route -> known
profile -> a cross-profile REST probe (resolveSessionProfile, stamps
ownership) before dispatch; only a request with no session at all falls to
ambient. Never the silent active fallback.
Tests updated to the new contract + knownSessionProfile coverage asserting it
returns undefined (not active) for an unknown session. tsc 0 errors.
* feat(desktop): add OAuth sign-in to the connections registry editor
A gated remote gateway (OAuth, or username/password) never accepts a
session token — it authenticates with a browser sign-in and the desktop
keeps whatever the flow mints. The registry editor only rendered a token
field for 'token' mode and nothing at all for 'oauth', so a gated
connection could be created but never authenticated: selecting OAuth left
an empty row, and Test failed with no way to fix it.
Render an Authentication row in the oauth branch that calls the existing
oauthLoginConnectionConfig IPC — the same one first-run-remote-form and
the gateway panel already use. The URL is probed (debounced) so the row
can name the provider and use password-specific copy when every
advertised provider supports passwords, matching gateway-settings.
No new i18n keys; all strings already exist under settings.gateway.
Test needed no change: testDesktopConnectionConfig already skips the
token for oauth and mints a ws-ticket from the session.
* fix(desktop): keep profile picks on the source being browsed
$profiles is the ACTIVE gateway's list, so a profile picked while a
registry source is live names one of THAT source's profiles. Both
selectProfile and newSessionInProfile sent it through the profile-only
path, which resolves the descriptor with a bare name — and
getConnection(profile) is answered against the primary. Picking
"researcher" while browsing a remote source therefore opened a LOCAL
backend of that name and snapped the gateway home, so the pick looked
like it never took: the user could reach the agent from Bot Mode but
never from the profile switcher.
Route both through the live source instead: a non-null
activeGatewayConnectionId means a registry source owns the current
gateway, so activate the (connection, profile) agent. A null id means
the primary is live, which is exactly the legacy path — single-source
users keep their existing behavior unchanged.
* fix(desktop): cancel the registry auth probe on unmount; reset the signed-in pill on mode flips
Review follow-ups on the OAuth sign-in row: the debounced probe sets a
cancelled flag in its effect cleanup (probeSeq covers staleness but not
unmount), and oauthConnected resets when the auth mode flips as well as on
URL changes — a saved row edited token -> oauth no longer reports a stale
'Signed in' from an earlier oauth stint.
* fix(desktop): stop Inbox-style session cards from clipping glyph ink
leading-none plus truncate (overflow:hidden) made the line box equal the
em-square, so Segoe UI on Windows shaved letter tops and bottoms. Give
truncated sidebar text 1.35 line-height and tighten card gaps so the
taller lines still fit.
* test(desktop): lock Inbox card lines to a line-height that fits glyph ink
Assert the workspace, title, and footer lines keep leading-[1.35] and
never fall back to leading-none, which is what clipped the screenshot.
No token and no cookie cannot self-heal. Tag that throw with
isReauthRequired so boot stops retrying and Sign in stays clickable.
Leave needsOauthLogin-only ticket 401s retryable for AT/RT rotation.
The Bot Mode 'session not found' / bot-runs-on-wrong-backend bug. wiring's
requestGateway is ONE shared closure for every session-scoped RPC in the
window, but it derived the owning profile from the globally-FOCUSED tile
($focusedStoredSessionId). A bot chat is a background tile while another pane
is active, so its prompt.submit carried the bot's own session_id yet was
dispatched on the FOCUSED tile's backend — the default backend served the bot
via ?profile= from the default's state.db, or answered 4001 'session not
found' when it didn't hold the runtime session.
Route by the session the RPC TARGETS (params.session_id) instead. session_id
is a RUNTIME id while tiles/rows key on the STORED id, so translate via the
state cache then a reverse scan of the stored->runtime map (the same ladder
use-session-tile-delegate's storedSessionIdForRuntime uses); an unresolved id
is already a stored id (several RPCs pass stored ids directly). RPCs with no
session_id (ambient/config) keep the focused->selected fallback.
Pure helpers extracted to wiring-routing.ts so they're unit-testable without
importing the React controller; 6 tests cover the target-vs-focused routing,
the stored-id passthrough, and the no-session fallback. tsc 0 errors.
Diagnosis verified on a live install: the fix was present in source but the
running build still misrouted, and logs showed the bot's turn executing on the
default backend while its own per-profile backend sat idle.
A legacy remote primary carries no registry connectionId, so the scoped
reconnect reset could not name the restarted owner and fell back to
preserving every owner-routed Bot tile -- leaving the restarted backend's
own Bot Chat bound to its dead runtime (the original bug, persisting for
that one connection shape).
Unknown identity now fails toward recovery instead: preserve only Bot
runtimes owned by provably-live secondary connections
(liveSecondaryConnectionIds()); everything else drops its binding and
re-resumes. A reset only costs a re-resume, so this is safe for the
preserved-set survivors and correct for the dead one.
Brief transport blips often self-heal in 1–3 minutes. Raise the
non-blocking escalate toast from 45s to 5m so those windows stay quiet
while chat remains readable/draftable. Confirmed reauth still takes the
full-screen recovery overlay immediately.
Post-boot WebSocket ticket mint failures and prolonged reconnects were
promoting into the full-screen "Hermes couldn't start" overlay, locking
users out of reading/drafting during brief 1–3 minute remote flaps.
- Ignore non-reauth boot-progress errors after a healthy cold boot
- Escalate prolonged transport reconnects with a non-blocking toast
- Soft-reset remote liveness rebuilds (no boot UI reset)
- Retry transient ws-ticket mints; auth rejections still fail fast
requestGateway (contrib/wiring) resolved the owning backend from
$focusedStoredSessionId — the WINDOW's focused tile — for every RPC it
dispatched. A session-scoped RPC names its real target in
params.session_id; whenever that session's chat was NOT the focused
pane (any background bot chat — the normal Bot Mode case), the RPC was
dispatched on whichever backend the focused tile happened to own. A
profile bot's prompt.submit then executed on the DEFAULT backend
(sessions created in the default store, profile logs empty), or failed
with 4001 'session not found' when default didn't hold the session.
Root cause isolated by Teknium: submit for a Developer-profile bot ran
in root logs/agent.log while profiles/developer/logs sat empty, and
completed only when default happened to hold the session — proving the
misroute sits downstream of the tile-owner-route lookup, in the
routing-key choice itself.
requestGateway now routes by the RPC's own target first:
params.session_id (a RUNTIME id) is translated to the stored id via the
tile map — new storedSessionIdForRuntimeId() in session-states, where
tiles already carry both identities — and only session-less RPCs
(config reads, list refreshes, cron) fall back to the focused-tile key,
which is genuinely window-ambient. Stored-id claims win over runtime
bindings in the lookup so a stale tile's dead runtimeId can never
hijack a live tile's identity.
#90587 replaced the bundled Nous with a GitHub-theme fork. The earlier
glass-and-cream palette stays available under its own name; the default
does not change.
The status stack hydrates the goal indicator on every session open via
/goal status. A finished goal stays status=done in the DB permanently,
so hydration re-created the '✓ Goal done' chip on every mount — the 8s
linger only applies to the live completion event, not re-hydration. In
Bot Mode (one endless session) the completed overlay never went away.
applyGoalStatusText now takes a hydrate flag: terminal (done) goals are
treated as no-goal during hydration and any lingering done chip is
dropped, while live events keep the 8s linger behavior.
Drives the real `syncBotsHomeWorkspace` against a shell whose reveal
does not hand the tab its zone's active slot — modelled on
`revealTreePane`'s hidden-pane early return — and asserts the passive
reconcile re-fronts once rather than on every pass. Against the
unfixed code the first case remounts the view 21 times for 21 passes.
The other three keep the bound from becoming a regression of its own:
giving up must leave the tab OPEN (a closed home drops the Bots tab
through to the ownerless Sessions composer); a cooperative shell still
gets its legitimate re-front, and gets another one the next time the
tab is genuinely backgrounded, so the budget is per-attempt rather than
one-shot for the life of the process; and an explicit gesture re-fronts
even on a shell where passive reconciles have already given up.
Re-fronting the Bots home tab is a close followed by a re-open, which
tears down and rebuilds the entire Bots view. `openBotsHomeWorkspace`
took that path on EVERY passive reconcile that found the tab open but
not holding its zone's active slot, with nothing bounding the retries.
That condition is not always transient. `revealTreePane` returns early
for a pane in `$hiddenTreePanes` without ever activating it,
`isPaneVisible` is false for a minimized zone, and a pane the tree
never adopted has no group to be active in. Pinned in any of those
states, every signal that reaches a surface sync — sidebar visibility
flips, focus churn, group changes — bought one more full remount, and
the view visibly strobed.
A passive reconcile now gets one attempt. The reveal has already
granted or refused the active slot by the time `openWorkspace` returns,
so the budget settles on that answer directly instead of waiting for a
visibility notification that is not coming: a computed store stays
silent when the value does not change. Retiring the tab starts a fresh
budget, and an explicit gesture is never blocked.
Giving up keeps the surface rather than closing it — a closed home
drops the Bots tab through to the ownerless Sessions composer, which is
the hole the home exists to plug.
Covers the behavior, not the markup: a row exposes an activation target
that opens THAT job; the opener never contains the switch or the delete
control (a nested interactive element would swallow the toggle and is
invalid markup anyway); detail rows carry only fields the gateway
actually sent, so a job that has never run drops those rows instead of
rendering "undefined"; a paused job reports Paused and promises no next
run; the raw schedule appears only when the humanized label dropped
something; and a failing job explains itself in failure order — the run
that never happened outranks the delivery of a run that did.
`routine-owner.test.mjs` asserted the row's owner routing by matching
`function RoutineRow({ job, owner })` against the plugin source, so it
broke on a parameter addition that changed no behavior. Replaced with
the real invariant it was reaching for: toggling the switch sends
`cron.manage` for the owner that rendered the row and evicts that
owner's cache key — which a signature change cannot fake.
In the Bots pane the Cronjobs rows were inert. The only interactive
controls were the enable switch and the hover-only delete button, so
clicking a cronjob to see what it runs, when it runs next, or why it
stopped did nothing at all — while the same job on the main Cron page
opens a full detail panel.
The gateway already ships every one of those facts with
`cron.manage list` (schedule, repeat, next/last run, last status,
delivery target, model, workdir, prompt preview, and the
fire/delivery/pause failures). None of it had a surface in Bot Mode: a
job failing every run reads exactly like a healthy paused one.
The row title becomes a real button that opens a read-only inspector
rendered from the record the pane is already holding — no extra RPC,
and no second mutation path beside the row's own switch and delete. The
switch and delete button stay siblings of the opener, so a toggle can
never be swallowed by the open. The inspector tracks the job by id
rather than by object, so the 20s poll keeps an open panel live instead
of freezing the snapshot it opened with.