* feat(desktop): add OAuth sign-in to the connections registry editor
A gated remote gateway (OAuth, or username/password) never accepts a
session token — it authenticates with a browser sign-in and the desktop
keeps whatever the flow mints. The registry editor only rendered a token
field for 'token' mode and nothing at all for 'oauth', so a gated
connection could be created but never authenticated: selecting OAuth left
an empty row, and Test failed with no way to fix it.
Render an Authentication row in the oauth branch that calls the existing
oauthLoginConnectionConfig IPC — the same one first-run-remote-form and
the gateway panel already use. The URL is probed (debounced) so the row
can name the provider and use password-specific copy when every
advertised provider supports passwords, matching gateway-settings.
No new i18n keys; all strings already exist under settings.gateway.
Test needed no change: testDesktopConnectionConfig already skips the
token for oauth and mints a ws-ticket from the session.
* fix(desktop): keep profile picks on the source being browsed
$profiles is the ACTIVE gateway's list, so a profile picked while a
registry source is live names one of THAT source's profiles. Both
selectProfile and newSessionInProfile sent it through the profile-only
path, which resolves the descriptor with a bare name — and
getConnection(profile) is answered against the primary. Picking
"researcher" while browsing a remote source therefore opened a LOCAL
backend of that name and snapped the gateway home, so the pick looked
like it never took: the user could reach the agent from Bot Mode but
never from the profile switcher.
Route both through the live source instead: a non-null
activeGatewayConnectionId means a registry source owns the current
gateway, so activate the (connection, profile) agent. A null id means
the primary is live, which is exactly the legacy path — single-source
users keep their existing behavior unchanged.
* fix(desktop): cancel the registry auth probe on unmount; reset the signed-in pill on mode flips
Review follow-ups on the OAuth sign-in row: the debounced probe sets a
cancelled flag in its effect cleanup (probeSeq covers staleness but not
unmount), and oauthConnected resets when the auth mode flips as well as on
URL changes — a saved row edited token -> oauth no longer reports a stale
'Signed in' from an earlier oauth stint.
* fix(desktop): stop Inbox-style session cards from clipping glyph ink
leading-none plus truncate (overflow:hidden) made the line box equal the
em-square, so Segoe UI on Windows shaved letter tops and bottoms. Give
truncated sidebar text 1.35 line-height and tighten card gaps so the
taller lines still fit.
* test(desktop): lock Inbox card lines to a line-height that fits glyph ink
Assert the workspace, title, and footer lines keep leading-[1.35] and
never fall back to leading-none, which is what clipped the screenshot.
No token and no cookie cannot self-heal. Tag that throw with
isReauthRequired so boot stops retrying and Sign in stays clickable.
Leave needsOauthLogin-only ticket 401s retryable for AT/RT rotation.
The Bot Mode 'session not found' / bot-runs-on-wrong-backend bug. wiring's
requestGateway is ONE shared closure for every session-scoped RPC in the
window, but it derived the owning profile from the globally-FOCUSED tile
($focusedStoredSessionId). A bot chat is a background tile while another pane
is active, so its prompt.submit carried the bot's own session_id yet was
dispatched on the FOCUSED tile's backend — the default backend served the bot
via ?profile= from the default's state.db, or answered 4001 'session not
found' when it didn't hold the runtime session.
Route by the session the RPC TARGETS (params.session_id) instead. session_id
is a RUNTIME id while tiles/rows key on the STORED id, so translate via the
state cache then a reverse scan of the stored->runtime map (the same ladder
use-session-tile-delegate's storedSessionIdForRuntime uses); an unresolved id
is already a stored id (several RPCs pass stored ids directly). RPCs with no
session_id (ambient/config) keep the focused->selected fallback.
Pure helpers extracted to wiring-routing.ts so they're unit-testable without
importing the React controller; 6 tests cover the target-vs-focused routing,
the stored-id passthrough, and the no-session fallback. tsc 0 errors.
Diagnosis verified on a live install: the fix was present in source but the
running build still misrouted, and logs showed the bot's turn executing on the
default backend while its own per-profile backend sat idle.
A legacy remote primary carries no registry connectionId, so the scoped
reconnect reset could not name the restarted owner and fell back to
preserving every owner-routed Bot tile -- leaving the restarted backend's
own Bot Chat bound to its dead runtime (the original bug, persisting for
that one connection shape).
Unknown identity now fails toward recovery instead: preserve only Bot
runtimes owned by provably-live secondary connections
(liveSecondaryConnectionIds()); everything else drops its binding and
re-resumes. A reset only costs a re-resume, so this is safe for the
preserved-set survivors and correct for the dead one.
Brief transport blips often self-heal in 1–3 minutes. Raise the
non-blocking escalate toast from 45s to 5m so those windows stay quiet
while chat remains readable/draftable. Confirmed reauth still takes the
full-screen recovery overlay immediately.
Post-boot WebSocket ticket mint failures and prolonged reconnects were
promoting into the full-screen "Hermes couldn't start" overlay, locking
users out of reading/drafting during brief 1–3 minute remote flaps.
- Ignore non-reauth boot-progress errors after a healthy cold boot
- Escalate prolonged transport reconnects with a non-blocking toast
- Soft-reset remote liveness rebuilds (no boot UI reset)
- Retry transient ws-ticket mints; auth rejections still fail fast
requestGateway (contrib/wiring) resolved the owning backend from
$focusedStoredSessionId — the WINDOW's focused tile — for every RPC it
dispatched. A session-scoped RPC names its real target in
params.session_id; whenever that session's chat was NOT the focused
pane (any background bot chat — the normal Bot Mode case), the RPC was
dispatched on whichever backend the focused tile happened to own. A
profile bot's prompt.submit then executed on the DEFAULT backend
(sessions created in the default store, profile logs empty), or failed
with 4001 'session not found' when default didn't hold the session.
Root cause isolated by Teknium: submit for a Developer-profile bot ran
in root logs/agent.log while profiles/developer/logs sat empty, and
completed only when default happened to hold the session — proving the
misroute sits downstream of the tile-owner-route lookup, in the
routing-key choice itself.
requestGateway now routes by the RPC's own target first:
params.session_id (a RUNTIME id) is translated to the stored id via the
tile map — new storedSessionIdForRuntimeId() in session-states, where
tiles already carry both identities — and only session-less RPCs
(config reads, list refreshes, cron) fall back to the focused-tile key,
which is genuinely window-ambient. Stored-id claims win over runtime
bindings in the lookup so a stale tile's dead runtimeId can never
hijack a live tile's identity.
#90587 replaced the bundled Nous with a GitHub-theme fork. The earlier
glass-and-cream palette stays available under its own name; the default
does not change.
The status stack hydrates the goal indicator on every session open via
/goal status. A finished goal stays status=done in the DB permanently,
so hydration re-created the '✓ Goal done' chip on every mount — the 8s
linger only applies to the live completion event, not re-hydration. In
Bot Mode (one endless session) the completed overlay never went away.
applyGoalStatusText now takes a hydrate flag: terminal (done) goals are
treated as no-goal during hydration and any lingering done chip is
dropped, while live events keep the 8s linger behavior.
Drives the real `syncBotsHomeWorkspace` against a shell whose reveal
does not hand the tab its zone's active slot — modelled on
`revealTreePane`'s hidden-pane early return — and asserts the passive
reconcile re-fronts once rather than on every pass. Against the
unfixed code the first case remounts the view 21 times for 21 passes.
The other three keep the bound from becoming a regression of its own:
giving up must leave the tab OPEN (a closed home drops the Bots tab
through to the ownerless Sessions composer); a cooperative shell still
gets its legitimate re-front, and gets another one the next time the
tab is genuinely backgrounded, so the budget is per-attempt rather than
one-shot for the life of the process; and an explicit gesture re-fronts
even on a shell where passive reconciles have already given up.
Re-fronting the Bots home tab is a close followed by a re-open, which
tears down and rebuilds the entire Bots view. `openBotsHomeWorkspace`
took that path on EVERY passive reconcile that found the tab open but
not holding its zone's active slot, with nothing bounding the retries.
That condition is not always transient. `revealTreePane` returns early
for a pane in `$hiddenTreePanes` without ever activating it,
`isPaneVisible` is false for a minimized zone, and a pane the tree
never adopted has no group to be active in. Pinned in any of those
states, every signal that reaches a surface sync — sidebar visibility
flips, focus churn, group changes — bought one more full remount, and
the view visibly strobed.
A passive reconcile now gets one attempt. The reveal has already
granted or refused the active slot by the time `openWorkspace` returns,
so the budget settles on that answer directly instead of waiting for a
visibility notification that is not coming: a computed store stays
silent when the value does not change. Retiring the tab starts a fresh
budget, and an explicit gesture is never blocked.
Giving up keeps the surface rather than closing it — a closed home
drops the Bots tab through to the ownerless Sessions composer, which is
the hole the home exists to plug.
Covers the behavior, not the markup: a row exposes an activation target
that opens THAT job; the opener never contains the switch or the delete
control (a nested interactive element would swallow the toggle and is
invalid markup anyway); detail rows carry only fields the gateway
actually sent, so a job that has never run drops those rows instead of
rendering "undefined"; a paused job reports Paused and promises no next
run; the raw schedule appears only when the humanized label dropped
something; and a failing job explains itself in failure order — the run
that never happened outranks the delivery of a run that did.
`routine-owner.test.mjs` asserted the row's owner routing by matching
`function RoutineRow({ job, owner })` against the plugin source, so it
broke on a parameter addition that changed no behavior. Replaced with
the real invariant it was reaching for: toggling the switch sends
`cron.manage` for the owner that rendered the row and evicts that
owner's cache key — which a signature change cannot fake.
In the Bots pane the Cronjobs rows were inert. The only interactive
controls were the enable switch and the hover-only delete button, so
clicking a cronjob to see what it runs, when it runs next, or why it
stopped did nothing at all — while the same job on the main Cron page
opens a full detail panel.
The gateway already ships every one of those facts with
`cron.manage list` (schedule, repeat, next/last run, last status,
delivery target, model, workdir, prompt preview, and the
fire/delivery/pause failures). None of it had a surface in Bot Mode: a
job failing every run reads exactly like a healthy paused one.
The row title becomes a real button that opens a read-only inspector
rendered from the record the pane is already holding — no extra RPC,
and no second mutation path beside the row's own switch and delete. The
switch and delete button stay siblings of the opener, so a toggle can
never be swallowed by the open. The inspector tracks the job by id
rather than by object, so the 20s poll keeps an open panel live instead
of freezing the snapshot it opened with.
Final-diff pass: trimGroupChatLog drops entries from the FRONT once a room
crosses the history cap, so slicing the post-turn log at the pre-turn
LENGTH could overshoot after a mid-turn trim, read an empty tail, and
silently commit a stale turn — re-opening #93127's double delivery in
long-history rooms exactly. Anchor on the last pre-turn entry's id; if the
anchor itself was trimmed, every surviving entry is newer, so scanning the
whole log stays exact.
- Cross-thread supersession no longer discards finished work: an epoch bump
from a send in ANOTHER thread doesn't re-drive this thread's members (delta
filters are thread-scoped), so dropping the finished reply lost completed
work until someone revisited the old thread. shouldCommitMemberTurn now
drops only when a newer USER entry landed in the same thread; the caller
computes that from the log tail past the pre-turn length.
- '@all stop' now holds every member — it parsed to everyone:true with no
mentions and silently held nobody, the asymmetric twin of the tested
'@all resume'. classifyGroupHoldDirective gains holdAll; the send path
passes the room's member keys for expansion.
- Tests pin both: cross-thread commit preserved, @all-stop holds all
(mutation-checked: reverting either guard fails its test).
A user 'stop @member' was just log text: the next room delta (receipt
round completing, any later turn) re-dispatched the member and it
re-claimed the very task it was told to stop. Holds are now durable
room state: set by an explicit user stop mention, checked by the round
loop before dispatch (skip consumes the delta exactly once — no spin),
released only by an explicit resume, @all resume, or a direct non-stop
mention of the held member. Holds persist and rehydrate with the same
durability as room watermarks, and the activity feed shows WHY a held
bot is silent (⏸ held glyph + hint) the first time it is skipped.
Conservative parse documented in-code: any standalone stop/halt/pause
next to a mention holds — a wrongly-held bot is one mention away from
release; a wrongly-running one keeps doing forbidden work.
Final-diff pass: pin both envelope mtimes via os.utime relative to the
watermark (write_text alone is wall-clock/FS dependent), and reset
relayDrainRerun in stopBotRelay so a rerun remembered mid-drain can't
leak one stale drain into the next start/stop cycle.
Review follow-ups:
- A push signal landing while drainRelayOutboxes is mid-flight hit the
relayDrainBusy early-return and was gone forever — the gateway signature
is monotone (one event per new envelope, never re-broadcast), so the
envelope waited out the full 4s poll, exactly the latency the push path
removes. relayDrainRerun remembers the race and schedules one debounced
follow-up pass after the drain finishes.
- test_new_envelope_after_drain_fires_pending_again pins the untested half
of the monotone contract: the watermark must not eat genuinely NEW
envelopes (write -> drain -> write-newer fires twice). Mutation-checked:
a stale-signature regression fails it while the other three still pass.
Cross-connection DMs were pure polling: the Desktop drains every gateway's
bot_relay outbox on a 4s interval, so each hop eats up to 4s outbound plus
4s for the reply leg (#92760 'bots reply slowly').
Emission point: the gateway's existing change watcher (_CHANGE_WATCHES in
tui_gateway/server.py). Envelopes are written by the AGENT process
(message_agent -> tools.bot_relay.enqueue_envelope), not the gateway, so no
gateway RPC is on the enqueue path and an in-process emit is impossible.
That is exactly the situation the change watcher already solves for the
pairing store (pairing.changed: 'written by a different process; the files
are the only shared signal') - so a new bot_relay.outbox.pending entry in
the existing watch table is the smallest correct diff: one cheap 1s-interval
stat probe folded into the existing 0.5s watcher tick, no new thread, no new
RPC, and _broadcast_global_event fans it to every connected WS client for
free. The signature is monotone (newest envelope mtime ever seen) so a
drain emptying outbox/ never re-fires the event.
Desktop (hermes-bots plugin): subscribe via the existing host.onEvent tap
(feature-detected - older shells lack it) and run drainRelayOutboxes through
a 250ms trailing debounce so a burst of signals collapses to one drain.
The 4s interval poll is intentionally UNCHANGED as the backstop: the event
tap only hears the active gateway socket, so per-connection push detection
would be complex and wrong to trade the poll against - push simply makes
the common case near-instant while older backends keep working exactly as
before.
Tests: 3 new watcher contracts (fires on enqueue, monotone across drain,
silent with no outbox) and a new relay-push-drain.test.mjs (debounce burst
-> one drain, re-arm after window, disposed no-op, poll backstop intact).
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
Review follow-up: the relay drain records attention under
'<connectionId>::<profile>', but local/unannotated roster rows carry no
bot.connectionId — botRosterKey gives 'legacy::name' and botSelectionKey
bare 'name', so a failed relay DM to a bot on the ACTIVE connection never
rendered its badge. BotRow now also checks
'<bot.connectionId || activeConnectionId>::<name>', covering exactly the
rows the user is most likely looking at. Test pins all three lookup shapes.
Post-merge review follow-up for #93080: the isDisabled guard test
documented 'flipping back also notifies' but never asserted it — a
regression making the true->false transition silent (e.g. gating
disabledChanged on truthiness) would still pass. Pin the flip-back
notify with a beforeFlipBack capture (mutation-checked: gating the
seed on newDisabled truthiness now fails this test).
Review follow-up: __internal_setAdapter assigned this.isDisabled before
the fast path but never fed it into the new 'changed' flag, so an
isDisabled-only flip on an otherwise-identical adapter swap would have
been silently swallowed. Seed 'changed' with the isDisabled comparison
and add a guard test (mutation-checked: reverting the seed fails it).
Two independent desktop-boot/runtime bugs found driving the app over CDP
against current main, each pinned by a regression test:
1) Adapter no-op notify loop: IncrementalExternalStoreThreadRuntimeCore.
__internal_setAdapter's fast path (same isRunning + same messageRepository)
called _notifySubscribers() unconditionally. ChatRuntimeBoundary passes a
fresh adapter literal every render, so any subscriber whose notification
re-renders the boundary loops render->setAdapter->notify->render until
React kills the tile with 'Maximum update depth exceeded' (reproduced live
on every bot-profile switch; session tile dies behind its error boundary).
Now the fast path notifies only when extras/suggestions/capabilities
actually changed.
2) Dual-venv interpreter mismatch: findPythonForRoot() prefers .venv over
venv, but createPythonBackend() hardcoded venvRoot=root/venv for
PYTHONPATH. A checkout with BOTH venvs (dev .venv 3.12 + install venv
3.11) got a 3.12 interpreter with 3.11-compiled native wheels on
PYTHONPATH and died on the first import (pydantic_core) before the
gateway bound - the renderer then showed 'Gateway offline' on every
profile. venvRootForPython() now maps the selected interpreter back to
ITS venv; root/venv remains the fallback for system pythons only.
The real fix for Bot Mode 'session not found' / endless hang: dispatch
session-scoped RPCs on the OWNING profile's local gateway, using the route
the chat tile already carries — the same multi-connection machinery Sessions
mode uses, which has never had this problem.
Root cause chain:
- A bot chat is a persisted tile that records its exact owner (connectionId +
profile) in tile.ownerRoute; requestForSessionProfile already dispatches on
any (connectionId, profile) via the per-profile local gateway pool.
- But wiring's requestGateway resolved the owner via rememberedSessionProfile,
a $sessions row lookup. Canonical Bot Chats are born hidden (never listed),
so the lookup missed and fell back to the ACTIVE profile -> prompt.submit hit
the launch backend that never owned the session -> 4001, and the resume
ladder re-resolved through the same blind spot, so it hung.
- It also keyed off $selectedStoredSessionId, but a bot chat renders in a TILE
whose id is $focusedStoredSessionId (selected stays the primary pane), so
even the row path was reading the wrong session.
Fix:
- wiring requestGateway: resolve owner from the FOCUSED stored id, preferring
the tile's persisted ownerRoute; fall back to the list-derived profile only
when no tile route exists. One resolver, every session RPC (submit, resume,
attach, interrupt, compress) inherits it.
- sdk openSession: synthesize a local ownerRoute from for bot opens
that carry no explicit cross-connection route, so LOCAL bot tiles carry their
owner too (previously only remote routes did). Strictly routing metadata:
the dial path, all-profiles view, and the route-registry retry check all
still key off the EXPLICIT route, so a plain local open behaves exactly as
before (no registry-secondary dial, no forced all-profiles view).
Fixes already-open chats (tile route is persisted, needs no fresh open) and
survives relaunch. 3 tests for sessionTileOwnerRoute. tsc 0 errors.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
When the gateway reaps the runtime behind the open bot chat (idle TTL,
LRU cap, or the WS-orphan mass reap that killed every background bot's
handle at once in the Aug 23 incident), the plugin now hears
session.reclaimed and re-resumes the canonical chat immediately, instead
of leaving the dead handle for the user's next send to trip over.
Matched on the stored id against both claim identities; guarded by the
open generation so a user action mid-re-resume wins; a failed re-resume
is swallowed — the next-send recovery ladder (#92928) stays the
backstop. Feature-detected on host.onEvent; disposed with the other
listeners.
groupSpeakerLabel resolved friendly identity for exactly one case: the
literal profile name 'default' → 'Hermes'. A renamed default (core
display_name via 'hermes profile rename', e.g. Lucy) or a Bot Mode title
never reached the room's working line, activity feed, or transcript
speaker prefix — the community report was Lucy's group turns still
reading 'hermes thinking'.
The label now walks the same rungs as displayName(): Bot Mode title
first, then the ACTIVE gateway roster row's display_name (remote/thin
rows are skipped so another connection's default can't lend its name),
then the existing default→Hermes fallback.
Validation: group-chat.test.mjs 87/87, full hermes-bots suite 474/474.
* fix(desktop): route hidden Bot Chat RPCs to the owning profile backend
Bot Mode chats failed with 4001 'session not found' (then hung on retry)
for every bot except the launch profile. Root cause is an identity gap in
the session-RPC router for HIDDEN sessions:
- wiring's requestGateway resolves the owning profile via
rememberedSessionProfile($sessions, selectedStoredSessionId, active).
- Canonical Bot Chats are born hidden (hermes-agent#86797), so the sidebar
aggregator NEVER lists them; whenever the in-memory row is absent the
lookup misses and the resolver silently falls back to the ACTIVE profile.
- prompt.submit then lands on the launch backend, which never owned the
session -> 4001. withSessionNotFoundResume's session.resume ALSO resolves
the profile through the same blind spot, so the recovery re-registers the
wrong backend too and the ladder dies without ever reaching the bot's own
(healthy) gateway. Log fingerprint: the bot backend shows ws accepts with
messages=0 and no 'tui prompt accepted' after the reap, while the launch
backend answers 4001.
Fix at the resolver, so every session-scoped RPC (submit, resume, attach,
interrupt, compress) gets the same answer:
- rememberedSessionProfile: when no session row matches, consult the
session-owner hint (targetProfile over profile) before falling back to
the active profile.
- sdk openSession: record an owner hint for LOCAL plugin opens that carry
an explicit profile (remote routes already did via ownerRoute). Hidden
sessions have no sidebar row, so this hint is the only durable owner
record the router can consult.
3 focused tests: hint fallback for a hidden session, targetProfile
preference, and row-over-hint precedence. tsc 0 errors (baseline-equal).
* test: add activeGatewayConnectionId to the full-replacement gateway mock
sdk/index.ts now imports it for the local-open owner hint; the full
vi.mock('@/store/gateway') in profile-routing.test.ts must export every
symbol the module under test imports or all 18 suite tests fail at load.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The dialog body's grid used the implicit column, which sizes its track to
unbreakable content: the nowrap view-link <code> forced the track wider than
the dialog, so the description and URL were clipped and the Copy button sat
past the right edge behind a horizontal scrollbar. Both DialogContent body
boxes now pin the column to minmax(0,1fr) so children truncate instead of
widening the track — this hardens every dialog against long unbreakable
content, not just this one.
The view link is also now a real anchor (system browser on click, link
context menu on right-click) inside a data-selectable-text row so the URL
can be highlighted and copied by hand; truncation clips the paint only,
selection still carries the full URL.
The Bots home landing appeared on EVERY first click of a bot whose
canonical Bot Chat had been compressed; only a second click got through.
openRosterBot claimed the center with the durable registry id, but the
session-focus edge fired by the open itself reports the compression-
lineage TIP. releaseStaleOpenBotChat compared tip !== registry id,
declared the claim stale, released it, and the home reasserted over the
freshly opened chat. The second click worked only because the tip was
already focused — no new focus edge fired to sabotage it.
openBotCanonicalChat now returns both identities (registryId + openedId);
the claim carries both; a focus edge matching EITHER keeps it. Foreign
sessions still release, and the legacy no-id draft claim is unchanged.
/voice arms SERVER-side capture (voice.record → PortAudio on the backend
host) — meaningless on desktop, which has its own composer-native voice
conversation (mic menu / Ctrl+B). It was already suppressed from the
slash palette, but typing it got the generic 'advanced' shrug that never
mentioned the button exists. New composer-voice unavailability reason
with a message naming the actual surface.