Final-diff pass: trimGroupChatLog drops entries from the FRONT once a room
crosses the history cap, so slicing the post-turn log at the pre-turn
LENGTH could overshoot after a mid-turn trim, read an empty tail, and
silently commit a stale turn — re-opening #93127's double delivery in
long-history rooms exactly. Anchor on the last pre-turn entry's id; if the
anchor itself was trimmed, every surviving entry is newer, so scanning the
whole log stays exact.
- Cross-thread supersession no longer discards finished work: an epoch bump
from a send in ANOTHER thread doesn't re-drive this thread's members (delta
filters are thread-scoped), so dropping the finished reply lost completed
work until someone revisited the old thread. shouldCommitMemberTurn now
drops only when a newer USER entry landed in the same thread; the caller
computes that from the log tail past the pre-turn length.
- '@all stop' now holds every member — it parsed to everyone:true with no
mentions and silently held nobody, the asymmetric twin of the tested
'@all resume'. classifyGroupHoldDirective gains holdAll; the send path
passes the room's member keys for expansion.
- Tests pin both: cross-thread commit preserved, @all-stop holds all
(mutation-checked: reverting either guard fails its test).
A user 'stop @member' was just log text: the next room delta (receipt
round completing, any later turn) re-dispatched the member and it
re-claimed the very task it was told to stop. Holds are now durable
room state: set by an explicit user stop mention, checked by the round
loop before dispatch (skip consumes the delta exactly once — no spin),
released only by an explicit resume, @all resume, or a direct non-stop
mention of the held member. Holds persist and rehydrate with the same
durability as room watermarks, and the activity feed shows WHY a held
bot is silent (⏸ held glyph + hint) the first time it is skipped.
Conservative parse documented in-code: any standalone stop/halt/pause
next to a mention holds — a wrongly-held bot is one mention away from
release; a wrongly-running one keeps doing forbidden work.
Final-diff pass: pin both envelope mtimes via os.utime relative to the
watermark (write_text alone is wall-clock/FS dependent), and reset
relayDrainRerun in stopBotRelay so a rerun remembered mid-drain can't
leak one stale drain into the next start/stop cycle.
Review follow-ups:
- A push signal landing while drainRelayOutboxes is mid-flight hit the
relayDrainBusy early-return and was gone forever — the gateway signature
is monotone (one event per new envelope, never re-broadcast), so the
envelope waited out the full 4s poll, exactly the latency the push path
removes. relayDrainRerun remembers the race and schedules one debounced
follow-up pass after the drain finishes.
- test_new_envelope_after_drain_fires_pending_again pins the untested half
of the monotone contract: the watermark must not eat genuinely NEW
envelopes (write -> drain -> write-newer fires twice). Mutation-checked:
a stale-signature regression fails it while the other three still pass.
Cross-connection DMs were pure polling: the Desktop drains every gateway's
bot_relay outbox on a 4s interval, so each hop eats up to 4s outbound plus
4s for the reply leg (#92760 'bots reply slowly').
Emission point: the gateway's existing change watcher (_CHANGE_WATCHES in
tui_gateway/server.py). Envelopes are written by the AGENT process
(message_agent -> tools.bot_relay.enqueue_envelope), not the gateway, so no
gateway RPC is on the enqueue path and an in-process emit is impossible.
That is exactly the situation the change watcher already solves for the
pairing store (pairing.changed: 'written by a different process; the files
are the only shared signal') - so a new bot_relay.outbox.pending entry in
the existing watch table is the smallest correct diff: one cheap 1s-interval
stat probe folded into the existing 0.5s watcher tick, no new thread, no new
RPC, and _broadcast_global_event fans it to every connected WS client for
free. The signature is monotone (newest envelope mtime ever seen) so a
drain emptying outbox/ never re-fires the event.
Desktop (hermes-bots plugin): subscribe via the existing host.onEvent tap
(feature-detected - older shells lack it) and run drainRelayOutboxes through
a 250ms trailing debounce so a burst of signals collapses to one drain.
The 4s interval poll is intentionally UNCHANGED as the backstop: the event
tap only hears the active gateway socket, so per-connection push detection
would be complex and wrong to trade the poll against - push simply makes
the common case near-instant while older backends keep working exactly as
before.
Tests: 3 new watcher contracts (fires on enqueue, monotone across drain,
silent with no outbox) and a new relay-push-drain.test.mjs (debounce burst
-> one drain, re-arm after window, disposed no-op, poll backstop intact).
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
Review follow-up: the relay drain records attention under
'<connectionId>::<profile>', but local/unannotated roster rows carry no
bot.connectionId — botRosterKey gives 'legacy::name' and botSelectionKey
bare 'name', so a failed relay DM to a bot on the ACTIVE connection never
rendered its badge. BotRow now also checks
'<bot.connectionId || activeConnectionId>::<name>', covering exactly the
rows the user is most likely looking at. Test pins all three lookup shapes.
Post-merge review follow-up for #93080: the isDisabled guard test
documented 'flipping back also notifies' but never asserted it — a
regression making the true->false transition silent (e.g. gating
disabledChanged on truthiness) would still pass. Pin the flip-back
notify with a beforeFlipBack capture (mutation-checked: gating the
seed on newDisabled truthiness now fails this test).
Review follow-up: __internal_setAdapter assigned this.isDisabled before
the fast path but never fed it into the new 'changed' flag, so an
isDisabled-only flip on an otherwise-identical adapter swap would have
been silently swallowed. Seed 'changed' with the isDisabled comparison
and add a guard test (mutation-checked: reverting the seed fails it).
Two independent desktop-boot/runtime bugs found driving the app over CDP
against current main, each pinned by a regression test:
1) Adapter no-op notify loop: IncrementalExternalStoreThreadRuntimeCore.
__internal_setAdapter's fast path (same isRunning + same messageRepository)
called _notifySubscribers() unconditionally. ChatRuntimeBoundary passes a
fresh adapter literal every render, so any subscriber whose notification
re-renders the boundary loops render->setAdapter->notify->render until
React kills the tile with 'Maximum update depth exceeded' (reproduced live
on every bot-profile switch; session tile dies behind its error boundary).
Now the fast path notifies only when extras/suggestions/capabilities
actually changed.
2) Dual-venv interpreter mismatch: findPythonForRoot() prefers .venv over
venv, but createPythonBackend() hardcoded venvRoot=root/venv for
PYTHONPATH. A checkout with BOTH venvs (dev .venv 3.12 + install venv
3.11) got a 3.12 interpreter with 3.11-compiled native wheels on
PYTHONPATH and died on the first import (pydantic_core) before the
gateway bound - the renderer then showed 'Gateway offline' on every
profile. venvRootForPython() now maps the selected interpreter back to
ITS venv; root/venv remains the fallback for system pythons only.
The real fix for Bot Mode 'session not found' / endless hang: dispatch
session-scoped RPCs on the OWNING profile's local gateway, using the route
the chat tile already carries — the same multi-connection machinery Sessions
mode uses, which has never had this problem.
Root cause chain:
- A bot chat is a persisted tile that records its exact owner (connectionId +
profile) in tile.ownerRoute; requestForSessionProfile already dispatches on
any (connectionId, profile) via the per-profile local gateway pool.
- But wiring's requestGateway resolved the owner via rememberedSessionProfile,
a $sessions row lookup. Canonical Bot Chats are born hidden (never listed),
so the lookup missed and fell back to the ACTIVE profile -> prompt.submit hit
the launch backend that never owned the session -> 4001, and the resume
ladder re-resolved through the same blind spot, so it hung.
- It also keyed off $selectedStoredSessionId, but a bot chat renders in a TILE
whose id is $focusedStoredSessionId (selected stays the primary pane), so
even the row path was reading the wrong session.
Fix:
- wiring requestGateway: resolve owner from the FOCUSED stored id, preferring
the tile's persisted ownerRoute; fall back to the list-derived profile only
when no tile route exists. One resolver, every session RPC (submit, resume,
attach, interrupt, compress) inherits it.
- sdk openSession: synthesize a local ownerRoute from for bot opens
that carry no explicit cross-connection route, so LOCAL bot tiles carry their
owner too (previously only remote routes did). Strictly routing metadata:
the dial path, all-profiles view, and the route-registry retry check all
still key off the EXPLICIT route, so a plain local open behaves exactly as
before (no registry-secondary dial, no forced all-profiles view).
Fixes already-open chats (tile route is persisted, needs no fresh open) and
survives relaunch. 3 tests for sessionTileOwnerRoute. tsc 0 errors.
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
When the gateway reaps the runtime behind the open bot chat (idle TTL,
LRU cap, or the WS-orphan mass reap that killed every background bot's
handle at once in the Aug 23 incident), the plugin now hears
session.reclaimed and re-resumes the canonical chat immediately, instead
of leaving the dead handle for the user's next send to trip over.
Matched on the stored id against both claim identities; guarded by the
open generation so a user action mid-re-resume wins; a failed re-resume
is swallowed — the next-send recovery ladder (#92928) stays the
backstop. Feature-detected on host.onEvent; disposed with the other
listeners.
groupSpeakerLabel resolved friendly identity for exactly one case: the
literal profile name 'default' → 'Hermes'. A renamed default (core
display_name via 'hermes profile rename', e.g. Lucy) or a Bot Mode title
never reached the room's working line, activity feed, or transcript
speaker prefix — the community report was Lucy's group turns still
reading 'hermes thinking'.
The label now walks the same rungs as displayName(): Bot Mode title
first, then the ACTIVE gateway roster row's display_name (remote/thin
rows are skipped so another connection's default can't lend its name),
then the existing default→Hermes fallback.
Validation: group-chat.test.mjs 87/87, full hermes-bots suite 474/474.
* fix(desktop): route hidden Bot Chat RPCs to the owning profile backend
Bot Mode chats failed with 4001 'session not found' (then hung on retry)
for every bot except the launch profile. Root cause is an identity gap in
the session-RPC router for HIDDEN sessions:
- wiring's requestGateway resolves the owning profile via
rememberedSessionProfile($sessions, selectedStoredSessionId, active).
- Canonical Bot Chats are born hidden (hermes-agent#86797), so the sidebar
aggregator NEVER lists them; whenever the in-memory row is absent the
lookup misses and the resolver silently falls back to the ACTIVE profile.
- prompt.submit then lands on the launch backend, which never owned the
session -> 4001. withSessionNotFoundResume's session.resume ALSO resolves
the profile through the same blind spot, so the recovery re-registers the
wrong backend too and the ladder dies without ever reaching the bot's own
(healthy) gateway. Log fingerprint: the bot backend shows ws accepts with
messages=0 and no 'tui prompt accepted' after the reap, while the launch
backend answers 4001.
Fix at the resolver, so every session-scoped RPC (submit, resume, attach,
interrupt, compress) gets the same answer:
- rememberedSessionProfile: when no session row matches, consult the
session-owner hint (targetProfile over profile) before falling back to
the active profile.
- sdk openSession: record an owner hint for LOCAL plugin opens that carry
an explicit profile (remote routes already did via ownerRoute). Hidden
sessions have no sidebar row, so this hint is the only durable owner
record the router can consult.
3 focused tests: hint fallback for a hidden session, targetProfile
preference, and row-over-hint precedence. tsc 0 errors (baseline-equal).
* test: add activeGatewayConnectionId to the full-replacement gateway mock
sdk/index.ts now imports it for the local-open owner hint; the full
vi.mock('@/store/gateway') in profile-routing.test.ts must export every
symbol the module under test imports or all 18 suite tests fail at load.
---------
Co-authored-by: Teknium <teknium1@users.noreply.github.com>
The dialog body's grid used the implicit column, which sizes its track to
unbreakable content: the nowrap view-link <code> forced the track wider than
the dialog, so the description and URL were clipped and the Copy button sat
past the right edge behind a horizontal scrollbar. Both DialogContent body
boxes now pin the column to minmax(0,1fr) so children truncate instead of
widening the track — this hardens every dialog against long unbreakable
content, not just this one.
The view link is also now a real anchor (system browser on click, link
context menu on right-click) inside a data-selectable-text row so the URL
can be highlighted and copied by hand; truncation clips the paint only,
selection still carries the full URL.
The Bots home landing appeared on EVERY first click of a bot whose
canonical Bot Chat had been compressed; only a second click got through.
openRosterBot claimed the center with the durable registry id, but the
session-focus edge fired by the open itself reports the compression-
lineage TIP. releaseStaleOpenBotChat compared tip !== registry id,
declared the claim stale, released it, and the home reasserted over the
freshly opened chat. The second click worked only because the tip was
already focused — no new focus edge fired to sabotage it.
openBotCanonicalChat now returns both identities (registryId + openedId);
the claim carries both; a focus edge matching EITHER keeps it. Foreign
sessions still release, and the legacy no-id draft claim is unchanged.
/voice arms SERVER-side capture (voice.record → PortAudio on the backend
host) — meaningless on desktop, which has its own composer-native voice
conversation (mic menu / Ctrl+B). It was already suppressed from the
slash palette, but typing it got the generic 'advanced' shrug that never
mentioned the button exists. New composer-voice unavailability reason
with a message naming the actual surface.
The Electron main has routed a profile with a `profiles.<name>` remote
entry in connection.json to its own pooled backend since
profileRemoteOverride() landed — but the only way to WRITE that entry was
hand-editing connection.json (#91349, design intent from #90223 /
6170f844: this belongs on the profile rail, not the machine-level
Gateways page).
- New "Connect to a remote host…" action on the profile-rail square's
context menu, opening a URL + token dialog that writes the exact
`profiles.<name>` shape through the existing typed
getConnectionConfig/applyConnectionConfig bridge (renderer never
touches connection.json; tokens ride the existing safeStorage
encryption path, with the allowPlainTextToken opt-in surfaced on
keyring-less machines).
- First-time connect shows a one-time confirmation with a plain-language
risk note; editing an existing override (token rotation) skips it.
- Overridden profile squares carry a "remote" globe badge, and the
tooltip/aria label names the host. A "Remove remote connection" button
clears the override (mode: local via the existing coerce path).
- Registry name collision: the dialog warns when the profile name
matches a v2 connections-registry id/label.
- Token rotation: when switching to an overridden profile fails with an
auth-shaped error (401/forbidden/invalid token), a re-enter-token
toast opens the same dialog instead of leaving a silently dead
profile. Connectivity failures stay generic.
- All five locales updated.
Closes#91349
Design-intent analysis credit: @otfnfn
A Desktop per-profile alias (e.g. moxie with a Cloud override) routes to a
remote backend's root profile: route { connectionId, profile: 'moxie',
targetProfile: 'default' }. Once the hosted backend answers the roster
itself, the row's identity is (connection, 'default') — a different key
than the alias meta — so the friendly name regressed to the raw Cloud
hostname after activation, and Cloud-only rosters showed generic 'Hermes'
instead of the configured alias (#89131).
Add a connection-exact alias index built from the credential-free route
inventory, keyed by (connectionId, targetProfile). displayName,
botRosterMeta, and botFriendlyNames consult it so the claimed backend row
reads as the alias (and its title/meta), while:
- same-named defaults on OTHER connections never borrow the identity
- two aliases claiming one backend row fail closed
- the local default and un-aliased remote defaults keep existing behavior
Evidence: @TheAirick's controlled candidate testing on #89131.
The speak-stream WebSocket resolved its URL through the bare v1
getConnection/getGatewayWsUrl pair, which answers for the PRIMARY backend.
When a registry remote connection rides over a machine that also has a
local Hermes install (the common case — the installer always installs the
full agent), spoken replies dialed the LOCAL backend and hit its
unconfigured TTS, while chat (connectionScoped REST) correctly went remote.
Users saw 'configure STT/TTS' although their remote gateway had voice
fully configured.
Resolve the PCM socket through the same (connectionId, profile) bridges
store/gateway's openSecondary uses, and never overwrite a backend-namespace
profile the registry mint already wrote into the URL (SSH remoteProfile
aliasing, sharedRemote scoping).
Every REST audio call already carried connectionScoped(); this was the one
remaining self-built audio URL. Contract pinned by
voice-playback.routing.test.ts (sabotage-verified: 2/4 fail on the old
resolver).
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.
Backend:
- tools/voice_client_config.py: single resolver; per-provider client
wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
elevenlabs-tts). Server-host-only providers (local whisper, edge,
command/plugin) and missing credentials resolve to {mode: relay}.
xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).
Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
profile) with 60s TTL, provider-direct STT + TTS calls, sentence
cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
below it. Barge-in via the same stopVoicePlayback sequence bump.
Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.
Docs: voice-mode.md client-direct section ships in this PR.
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.
Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
Follow-up on the #91828 salvage: regression test for the subtle
resolveDiskPluginEntry branch that rejects a DIRECTORY literally named
plugin.js, and lint fix. The truncation-fix semantics from #92809 are
preserved: an oversize source keeps its error inventory row (returns
true — the file exists and was fully probed), while a vanished file
returns false so the scanner reconciles the ghost.
Runtime desktop plugins read their source through hermes:readFileText,
the preview IPC that silently truncates at TEXT_PREVIEW_MAX_BYTES
(512 KiB) — a larger plugin.js evaluated as a partial file (cryptic
syntax error, or worse, a half-module that parses).
- electron/main.ts: dedicated hermes:readPluginSource handler — full
read, 16 MiB cap enforced as a hard EFBIG via resolveReadableFileForIpc
(same path hardening + sensitive-file blocking), never truncation.
- preload.ts / global.d.ts: bridge + types (optional — older shells).
- runtime-loader.ts: loadDiskPlugin reads via readPluginSource; on older
shells falls back to readFileText but FAILS LOUDLY on truncated:true
(error toast + error inventory row) instead of evaluating a partial
file. Existence probe keeps the preview read (metadata is enough).
- tests: full-read path, loud old-shell truncation failure (sabotage-
verified), small-plugin fallback.
Builds on @mrsucesso's durable-marker recovery (previous commit):
- GET /api/hermes/update/receipt — the full durable receipt (steps,
skips, gateway restart outcome, fleet matrix) + compact summary; the
authoritative update-outcome record (written by every run since
#91283, including refused/failed).
- /api/actions/hermes-update/status now attaches the receipt summary,
and when BOTH the in-memory registries and the update.log marker are
gone (dashboard restarted + log rotated — the #81193 state), a
finished receipt reports the outcome: success→0, partial→1. A
still-running receipt proves nothing (clients keep polling).
- Desktop (updates.ts): the apply poll reads the attached receipt — a
finished receipt whose run started at/after this apply is
authoritative, replacing timeout-based failure inference across the
update's restart gap ('Backend update failed' on successful updates,
#81193; 'boot failed' during update restarts, #87359).
Live-verified: real uvicorn server + real UpdateReceipt writer (the
exact code hermes update runs) over real HTTP — receipt endpoint 200
with summary; #81193 state (no registries, no marker) reports success
from the receipt alone; partial receipt with a DOWN fleet row maps to
exit 1 (no false success).