Commit Graph

201 Commits

Author SHA1 Message Date
686f6c61 961635c19c fix(desktop): copy unsafe RPC rejections instead of mutating name
asRpcError now always wraps a non-string name in a fresh Error. In-place
assignment was a silent no-op on sealed objects in sloppy mode. Catch
only host.request / requestProfile so routing TypeErrors keep their stack.
2026-08-25 16:25:41 -07:00
686f6c61 a837c7aaa2 fix(desktop): coerce bot RPC rejections for React 19 error formatting
JSON-RPC/IPC can reject with a plain object whose name is a number.
React 19 then crashes on (error.name || '').trim, which takes down the
Routines pane instead of showing the cron.manage failure. requestForBot
now wraps those values in an Error with a string name, including
cross-realm Error-like objects from the plugin test vm.
2026-08-25 16:25:41 -07:00
Finn763 936a6ea8f7 fix(desktop): scope Cronjobs pane to the roster-clicked bot when the focused session has no owner (#94516)
The SDK's focusedSessionOwner store fails closed to null whenever the
focused session has no unique bot owner (a normal chat, ambiguous owner
hints) - the common case while the user browses the Bots pane.
resolveRoutineOwner treated that null as an error and returned null
before consulting the roster selection, so the Routines pane pinned
every agent on 'Cronjobs are unavailable until this agent appears in
the roster.'

Drop the fail-closed null gate and fall through to the existing
selection ladder (focusedBot || selectedBot || ...). An authoritative
focused owner still wins through its exact roster row and still fails
closed when that row is absent; a null owner with no matching selection
also still fails closed. Regression tests prove red pre-fix, green
post-fix.
2026-08-25 16:24:54 -07:00
Teknium 85958abf09 test: drop unused param in group-turn lease mock (lint) 2026-08-24 03:41:34 -07:00
Teknium c29c9d1790 fix(desktop): hold a per-turn socket lease so group-chat member turns survive the runtime-session reaper (#93602)
A group member turn is a session-scoped RPC sequence (resume → attach →
prompt.submit → poll) issued with the runtime id its first RPC minted, but
requestForBot routes every RPC through its own request-scoped socket lease
(retained:false secondaries in store/gateway). Between two RPCs the refcount
hits 0, the leased socket closes, the gateway detaches the runtime session on
WS disconnect, the orphan reaper frees it after grace, and the next RPC —
prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts
and the sub-profile bot goes silent in the room.

- store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted
  hold on the pooled socket with an idempotent release, mirroring the
  existing request-lease machinery.
- sdk: host.retainProfile(route) exposes the retain to plugins
  (feature-detected by consumers; older hosts keep working).
- hermes-bots plugin: runGroupChatMemberTurn acquires the lease before
  ensureGroupChatSession's first RPC and releases in finally, so the socket
  that minted the runtime id stays open across attach+submit+poll; and
  prompt.submit gets a one-shot catch-and-retry that re-resumes via the
  STORED session id on 4001-class failures (belt-and-braces for routes the
  lease can't cover). 4007 'never existed' keeps flowing to session.create.

Tests: simulated 4001 on first submit recovers via re-resume and delivers;
lease held across attach+submit (mock refcount never hits 0 mid-turn); lease
released after success AND failure; no-retainProfile host feature detection;
store-level retain/release + idempotent double-release + the unretained
disposal race.
2026-08-24 03:41:34 -07:00
Teknium 06fc941d23 fix(desktop): stop bot-relay drain loop from redialing a WebSocket per connection per tick
The bot relay's drain loop RPCs every registered connection through
requestGatewayForAgent's per-request lease. With no other consumer
holding the route, the refcount hit 0 after every tick and the pooled
secondary was disposed — a fresh WebSocket dial + teardown per
connection every 4s, flooding the gateway logs with connect/disconnect
pairs (#93594).

Two changes, both directions from the issue:

- Retained relay-route secondaries: retainGatewayForRelay pins a
  route's pooled socket with a counted retention (never clobbering the
  foreground 'retained' flag) for the relay's active lifetime, reusing
  the existing scheduleReconnect/full-jitter machinery on drops. The
  plugin pins each registered connection once via the new feature-
  detected host.retainProfileSocket door, reconciles pins with the
  current connection set on every drain, and releases everything in
  stopBotRelay/dispose. Local routes (null/'local') are exempt so the
  idle reaper can still reclaim spawned local backends. The live-work
  pruner also respects the pin.

- RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (#93091,
  bot_relay.outbox.pending) carries envelope latency, so the poll is
  purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS.

Tests: relay-push-drain updated to the new backstop semantics; new
gateway-relay-retention.test.ts proves one socket construction across
5 drain ticks (vs 3 constructions for 3 unretained ticks) and that
release/prune/local-exemption behave; new relay-socket-retention
plugin test pins the pin-once / release-on-departure / stop-releases
contracts.
2026-08-24 03:23:15 -07:00
Teknium 40ab950ae3 fix(desktop): guard render-reachable route lookups against orphaned rows
Audit of the remaining unguarded botConnectionRoute() callers a pane
render can reach (#93492 follow-up to the botRosterMeta split). Each now
uses the non-throwing resolveBotConnectionRoute() and degrades on an
owner_removed row instead of throwing into the pane's error boundary:

- botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility
  listener, Bots home open, and roster context menus recompute these on
  passive UI edges; an orphaned selection now yields the name-keyed owner
  and the blocked workspace target.
- durableGroupChatMembers: rebuilt on every group send over the whole
  seated roster; one orphaned member no longer aborts the room update,
  and a swept member's degraded mark now survives the rebuild.
- useModelOptions: hook body runs during render; the query is disabled
  for an orphaned row and the picker paints its error/disabled state.
- AdvancedProfileConfig: dialog falls back to the bot's own name scope.

Strict dispatch callers (requestForBot, session creation, deleteBot,
duplicateBot, ensureBotMetadata, routines) intentionally keep the
fail-closed throw — remote-routing-races.test.mjs still asserts it.

Adds orphaned-connection-members.test.mjs covering the removed-connection
sweep, the hydrate annotate (with/without a readable registry), the
degraded 'Gateway removed' rendering of swept rows, and every guarded
caller.
2026-08-24 03:22:28 -07:00
Teknium 725cfe2909 fix(desktop): annotate group-chat members already orphaned before hydrate
Rows poisoned before the removed-connection sweep existed (their
connection was deleted while an older Desktop ran, so no lifecycle push
ever swept them) are what made #93492 survive app restarts. After the
persisted 'group-chats' hydrate, run a pure annotate pass over the rooms:

- a descriptor that lost its connectionId (route unresolvable — the exact
  shape that threw on render) is always marked;
- a descriptor whose connectionId is absent from the live connection
  registry is marked only when the registry could actually be read —
  an unavailable registry must not read as 'everything is orphaned'.

Marked rows keep their identity and degrade to the existing 'Gateway
removed' state; nothing is deleted.
2026-08-24 03:22:28 -07:00
Teknium c509af689f fix(desktop): sweep group-chat rosters when a connection is removed
Root cause of #93492: deleting a cloud/remote connection disposed its
gateways (store/gateway.ts) but never touched the persisted 'group-chats'
storage, so every member descriptor referencing the deleted connection
stayed behind as a poisoned row (remoteSource: true, connection gone) that
render-path route lookups tripped over forever.

Subscribe to the connection registry's 'removed' lifecycle push
(window.hermesDesktop.connections.onChanged, feature-detected — older
Electron mains don't emit it) and annotate every persisted group-chat
member owned by the deleted connection. Rows are marked
(sourceMissing/sourceReachable), never silently deleted: the member keeps
its identity and panes render the existing degraded 'Gateway removed'
botSourceStatus state. Writes ride updateGroupChat so the durable record
keeps its full shape, and the listener unbinds on plugin dispose.
2026-08-24 03:22:28 -07:00
chelsealong ec013b76db fix(desktop): split strict connection routing from passive roster lookup
botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.

Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
2026-08-24 03:22:28 -07:00
chelsealong 09529afdd2 fix(desktop): group chats no longer crash when a member's connection is deleted
botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).

botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.

Fixes #93492
2026-08-24 03:22:28 -07:00
quexiaolong 7e67f64fce fix(desktop): bots group chat sends message on IME composition Enter
macOS Chinese pinyin IME: pressing Enter to confirm a candidate word in
the group-chat composer submitted the draft as a message mid-composition.
The GroupMentionInput onKeyDown checked only `event.key === 'Enter' &&
!event.shiftKey` with no IME guard, unlike the core composer which guards
isComposing + keyCode 229 (#44135).

Add the same guard to the three Enter handlers in the bots plugin:
- GroupMentionInput (group composer + reply box) — the reported bug
- GroupClarifyCard free-text answer input — same premature-submit
- skill-hub search input — same premature-trigger

Closes #93528
2026-08-23 23:33:33 -07:00
Gille 34feb375ac fix(desktop): clear stale group metadata on disband 2026-08-23 21:59:33 -07:00
Teknium c584d15cdc feat(bots): typed failure reasons reach the sending agent on A2A calls (#93091)
message_agent callers previously got provider prose (a raw 401
paragraph, a missing-provider essay) and could not branch on the
failure class. Now the #93091 item-1 reason enum rides the whole relay
roundtrip:

- Desktop relay drain forwards bot_relay.deliver's error.data.reason
  into bot_relay.reply (and prefers it for the attention badge over
  free-text re-parsing);
- write_reply already persisted reason / classified fallbacks;
- the sender-side waiter prints "[reason: <code>]" ahead of the free
  text, so the completion notification the sending agent receives is
  machine-branchable.

Additive everywhere: healthy replies unchanged, reasonless errors
classify to a code, old consumers keep working.
2026-08-23 20:07:21 -07:00
pierrenode 525597c9c3 fix(bot-mode): fail closed on transient group-session resume failures
ensureGroupChatSession's resume loop caught ANY session.resume error
(stored sid, then title lookup) identically and fell through to
session.create — the same bug findExistingCanonicalChat was fixed for
hours earlier (87b645f52c) in the same file: a transient failure (the
backend still warming up after a restart, a network blip on a
cross-connection lookup, an oversized-resume refusal) read as "no
session, mint a new one". That forks the member's real session AND
silently overwrites room.sessions[key], making the original
unreachable from the room. ensureGroupChatSession is actually more
exposed than the 1:1 case: it runs every group turn
(runGroupChatMemberTurn), with two independent swallow points.

Distinguish "genuinely doesn't exist" from "transient failure" the
same way the gateway itself does: session.resume's own handler
(tui_gateway/methods_session.py) returns JSON-RPC code 4007 only when
the target truly isn't found; every other failure (including 4130,
"session too large to resume" — a session that DOES exist) now
surfaces instead of being silently swallowed. The existing outer
try/catch at the call site already treats a thrown error as "this
member passes the round" (recordGroupActivity kind: 'failed'), so
nothing new needs to catch it — a transient hiccup now costs one
skipped round instead of a permanent fork.
2026-08-23 18:58:40 -07:00
fangliquanflq 8fca54f9f1 test(bots): pin draft sweep age boundary 2026-08-23 18:25:35 -07:00
fangliquanflq 1b18442f36 fix(bots): protect new drafts from title sweep 2026-08-23 18:25:35 -07:00
HexLab98 650cf3348f test(bot-mode): pin the Bots home re-front budget
Drives the real `syncBotsHomeWorkspace` against a shell whose reveal
does not hand the tab its zone's active slot — modelled on
`revealTreePane`'s hidden-pane early return — and asserts the passive
reconcile re-fronts once rather than on every pass. Against the
unfixed code the first case remounts the view 21 times for 21 passes.

The other three keep the bound from becoming a regression of its own:
giving up must leave the tab OPEN (a closed home drops the Bots tab
through to the ownerless Sessions composer); a cooperative shell still
gets its legitimate re-front, and gets another one the next time the
tab is genuinely backgrounded, so the budget is per-attempt rather than
one-shot for the life of the process; and an explicit gesture re-fronts
even on a shell where passive reconciles have already given up.
2026-08-23 15:35:48 -07:00
HexLab98 2350f4dc60 fix(bot-mode): bound the Bots home re-front so the view stops strobing
Re-fronting the Bots home tab is a close followed by a re-open, which
tears down and rebuilds the entire Bots view. `openBotsHomeWorkspace`
took that path on EVERY passive reconcile that found the tab open but
not holding its zone's active slot, with nothing bounding the retries.

That condition is not always transient. `revealTreePane` returns early
for a pane in `$hiddenTreePanes` without ever activating it,
`isPaneVisible` is false for a minimized zone, and a pane the tree
never adopted has no group to be active in. Pinned in any of those
states, every signal that reaches a surface sync — sidebar visibility
flips, focus churn, group changes — bought one more full remount, and
the view visibly strobed.

A passive reconcile now gets one attempt. The reveal has already
granted or refused the active slot by the time `openWorkspace` returns,
so the budget settles on that answer directly instead of waiting for a
visibility notification that is not coming: a computed store stays
silent when the value does not change. Retiring the tab starts a fresh
budget, and an explicit gesture is never blocked.

Giving up keeps the surface rather than closing it — a closed home
drops the Bots tab through to the ownerless Sessions composer, which is
the hole the home exists to plug.
2026-08-23 15:35:48 -07:00
HexLab98 127a72b223 test(bot-mode): pin the cronjob inspector, and drop a source-regex test
Covers the behavior, not the markup: a row exposes an activation target
that opens THAT job; the opener never contains the switch or the delete
control (a nested interactive element would swallow the toggle and is
invalid markup anyway); detail rows carry only fields the gateway
actually sent, so a job that has never run drops those rows instead of
rendering "undefined"; a paused job reports Paused and promises no next
run; the raw schedule appears only when the humanized label dropped
something; and a failing job explains itself in failure order — the run
that never happened outranks the delivery of a run that did.

`routine-owner.test.mjs` asserted the row's owner routing by matching
`function RoutineRow({ job, owner })` against the plugin source, so it
broke on a parameter addition that changed no behavior. Replaced with
the real invariant it was reaching for: toggling the switch sends
`cron.manage` for the owner that rendered the row and evicts that
owner's cache key — which a signature change cannot fake.
2026-08-23 15:35:43 -07:00
HexLab98 913de4eca6 fix(bot-mode): a cronjob row opens — clicking one no longer does nothing
In the Bots pane the Cronjobs rows were inert. The only interactive
controls were the enable switch and the hover-only delete button, so
clicking a cronjob to see what it runs, when it runs next, or why it
stopped did nothing at all — while the same job on the main Cron page
opens a full detail panel.

The gateway already ships every one of those facts with
`cron.manage list` (schedule, repeat, next/last run, last status,
delivery target, model, workdir, prompt preview, and the
fire/delivery/pause failures). None of it had a surface in Bot Mode: a
job failing every run reads exactly like a healthy paused one.

The row title becomes a real button that opens a read-only inspector
rendered from the record the pane is already holding — no extra RPC,
and no second mutation path beside the row's own switch and delete. The
switch and delete button stay siblings of the opener, so a toggle can
never be swallowed by the open. The inspector tracks the job by id
rather than by object, so the 20s poll keeps an open panel live instead
of freezing the snapshot it opened with.
2026-08-23 15:35:43 -07:00
kshitij a234e93654 fix(desktop): anchor the during-turn tail by entry id, not index
Final-diff pass: trimGroupChatLog drops entries from the FRONT once a room
crosses the history cap, so slicing the post-turn log at the pre-turn
LENGTH could overshoot after a mid-turn trim, read an empty tail, and
silently commit a stale turn — re-opening #93127's double delivery in
long-history rooms exactly. Anchor on the last pre-turn entry's id; if the
anchor itself was trimmed, every surviving entry is newer, so scanning the
whole log stays exact.
2026-08-24 02:54:57 +05:30
kshitij 2863e8fb5d fix(desktop): review follow-ups for the room-race fixes
- Cross-thread supersession no longer discards finished work: an epoch bump
  from a send in ANOTHER thread doesn't re-drive this thread's members (delta
  filters are thread-scoped), so dropping the finished reply lost completed
  work until someone revisited the old thread. shouldCommitMemberTurn now
  drops only when a newer USER entry landed in the same thread; the caller
  computes that from the log tail past the pre-turn length.
- '@all stop' now holds every member — it parsed to everyone:true with no
  mentions and silently held nobody, the asymmetric twin of the tested
  '@all resume'. classifyGroupHoldDirective gains holdAll; the send path
  passes the room's member keys for expansion.
- Tests pin both: cross-thread commit preserved, @all-stop holds all
  (mutation-checked: reverting either guard fails its test).
2026-08-24 02:39:11 +05:30
kshitijk4poor 0c474d7820 fix(desktop): make member stop sticky — per-member hold until explicit resume (#93129)
A user 'stop @member' was just log text: the next room delta (receipt
round completing, any later turn) re-dispatched the member and it
re-claimed the very task it was told to stop. Holds are now durable
room state: set by an explicit user stop mention, checked by the round
loop before dispatch (skip consumes the delta exactly once — no spin),
released only by an explicit resume, @all resume, or a direct non-stop
mention of the held member. Holds persist and rehydrate with the same
durability as room watermarks, and the activity feed shows WHY a held
bot is silent (⏸ held glyph + hint) the first time it is skipped.

Conservative parse documented in-code: any standalone stop/halt/pause
next to a mention holds — a wrongly-held bot is one mention away from
release; a wrongly-running one keeps doing forbidden work.
2026-08-24 02:08:53 +05:30
kshitijk4poor ae6baf333f fix(desktop): drop superseded group-chat turns and dedupe adjacent identical replies (#93127) 2026-08-24 02:08:53 +05:30
kshitij 1aadf863f9 test(bot-mode): deterministic watermark timestamps + reset rerun flag on relay stop
Final-diff pass: pin both envelope mtimes via os.utime relative to the
watermark (write_text alone is wall-clock/FS dependent), and reset
relayDrainRerun in stopBotRelay so a rerun remembered mid-drain can't
leak one stale drain into the next start/stop cycle.
2026-08-24 01:51:59 +05:30
kshitij bb63e0c4c7 fix(bot-mode): re-schedule a push that races an in-flight drain + pin the watermark's fire-again contract
Review follow-ups:
- A push signal landing while drainRelayOutboxes is mid-flight hit the
  relayDrainBusy early-return and was gone forever — the gateway signature
  is monotone (one event per new envelope, never re-broadcast), so the
  envelope waited out the full 4s poll, exactly the latency the push path
  removes. relayDrainRerun remembers the race and schedules one debounced
  follow-up pass after the drain finishes.
- test_new_envelope_after_drain_fires_pending_again pins the untested half
  of the monotone contract: the watermark must not eat genuinely NEW
  envelopes (write -> drain -> write-newer fires twice). Mutation-checked:
  a stale-signature regression fails it while the other three still pass.
2026-08-24 01:40:48 +05:30
kshitijk4poor 9c829f965d feat(bot-mode): push-notified relay drain with poll backstop (#93091)
Cross-connection DMs were pure polling: the Desktop drains every gateway's
bot_relay outbox on a 4s interval, so each hop eats up to 4s outbound plus
4s for the reply leg (#92760 'bots reply slowly').

Emission point: the gateway's existing change watcher (_CHANGE_WATCHES in
tui_gateway/server.py). Envelopes are written by the AGENT process
(message_agent -> tools.bot_relay.enqueue_envelope), not the gateway, so no
gateway RPC is on the enqueue path and an in-process emit is impossible.
That is exactly the situation the change watcher already solves for the
pairing store (pairing.changed: 'written by a different process; the files
are the only shared signal') - so a new bot_relay.outbox.pending entry in
the existing watch table is the smallest correct diff: one cheap 1s-interval
stat probe folded into the existing 0.5s watcher tick, no new thread, no new
RPC, and _broadcast_global_event fans it to every connected WS client for
free. The signature is monotone (newest envelope mtime ever seen) so a
drain emptying outbox/ never re-fires the event.

Desktop (hermes-bots plugin): subscribe via the existing host.onEvent tap
(feature-detected - older shells lack it) and run drainRelayOutboxes through
a 250ms trailing debounce so a burst of signals collapses to one drain.
The 4s interval poll is intentionally UNCHANGED as the backstop: the event
tap only hears the active gateway socket, so per-connection push detection
would be complex and wrong to trade the poll against - push simply makes
the common case near-instant while older backends keep working exactly as
before.

Tests: 3 new watcher contracts (fires on enqueue, monotone across drain,
silent with no outbox) and a new relay-push-drain.test.mjs (debounce burst
-> one drain, re-arm after window, disposed no-op, poll backstop intact).
2026-08-24 01:17:28 +05:30
kshitij 2eaa863112 Merge pull request #93102 from kshitijk4poor/feat/bot-envelope-ttl
feat(bot-mode): envelope TTL + offline fast-fail for bot relay (#93091 item 2)
2026-08-24 01:10:53 +05:30
kshitij e00d6c1995 fix(desktop): don't push a live connection as absent when its profile fetch blips
Review follow-up: relayAgentsOn() returned [] on ANY error, so a transient
profiles.list timeout pushed a fresh union roster missing a LIVE machine's
agents — and the gateway-side _target_liveness reads 'absent from a fresh
roster' as definitively offline, refusing enqueues with a false
runtime_offline during the ~60s window. Failure now returns null (distinct
from a genuinely empty list); syncRelayRosters reuses the last good rows
for that connection and prunes the cache when a connection truly leaves
profileRoutes. Source-contract test pins null-on-failure + cache fallback.
2026-08-24 01:07:33 +05:30
kshitij 387698a60d fix(desktop): badge active-gateway bots too — resolve the relay attention key for local rows
Review follow-up: the relay drain records attention under
'<connectionId>::<profile>', but local/unannotated roster rows carry no
bot.connectionId — botRosterKey gives 'legacy::name' and botSelectionKey
bare 'name', so a failed relay DM to a bot on the ACTIVE connection never
rendered its badge. BotRow now also checks
'<bot.connectionId || activeConnectionId>::<name>', covering exactly the
rows the user is most likely looking at. Test pins all three lookup shapes.
2026-08-24 00:35:39 +05:30
kshitijk4poor d4f426792b feat(desktop): needs-attention badge for background bot failures (#93091 item 3) 2026-08-23 23:27:34 +05:30
Teknium d5281f5981 feat(bot-mode): a reclaimed bot chat re-resumes itself — no stale-id error on the next send
When the gateway reaps the runtime behind the open bot chat (idle TTL,
LRU cap, or the WS-orphan mass reap that killed every background bot's
handle at once in the Aug 23 incident), the plugin now hears
session.reclaimed and re-resumes the canonical chat immediately, instead
of leaving the dead handle for the user's next send to trip over.

Matched on the stored id against both claim identities; guarded by the
open generation so a user action mid-re-resume wins; a failed re-resume
is swallowed — the next-send recovery ladder (#92928) stays the
backstop. Feature-detected on host.onEvent; disposed with the other
listeners.
2026-08-23 06:01:10 -07:00
Teknium f530cd2b54 fix(bot-mode): group rooms name renamed bots — 'Lucy is thinking…', never a stale 'Hermes'
groupSpeakerLabel resolved friendly identity for exactly one case: the
literal profile name 'default' → 'Hermes'. A renamed default (core
display_name via 'hermes profile rename', e.g. Lucy) or a Bot Mode title
never reached the room's working line, activity feed, or transcript
speaker prefix — the community report was Lucy's group turns still
reading 'hermes thinking'.

The label now walks the same rungs as displayName(): Bot Mode title
first, then the ACTIVE gateway roster row's display_name (remote/thin
rows are skipped so another connection's default can't lend its name),
then the existing default→Hermes fallback.

Validation: group-chat.test.mjs 87/87, full hermes-bots suite 474/474.
2026-08-23 05:45:24 -07:00
Teknium a4c6c6bddb fix(bot-mode): first click on a bot opens its chat — the home no longer bounces over it
The Bots home landing appeared on EVERY first click of a bot whose
canonical Bot Chat had been compressed; only a second click got through.

openRosterBot claimed the center with the durable registry id, but the
session-focus edge fired by the open itself reports the compression-
lineage TIP. releaseStaleOpenBotChat compared tip !== registry id,
declared the claim stale, released it, and the home reasserted over the
freshly opened chat. The second click worked only because the tip was
already focused — no new focus edge fired to sabotage it.

openBotCanonicalChat now returns both identities (registryId + openedId);
the claim carries both; a focus edge matching EITHER keeps it. Foreign
sessions still release, and the legacy no-id draft claim is unchanged.
2026-08-23 04:57:46 -07:00
Teknium 2ec229ec5a fixup: roster query keeps SDK ambient owner route; alias index refresh preserved 2026-08-23 04:19:55 -07:00
David Dudok de Wit 8523819fcf refactor(desktop): consume upstream Bot owner routing 2026-08-23 04:19:55 -07:00
David Dudok de Wit 9ff7f3325c fix(desktop): preserve cross-realm Bot registry errors 2026-08-23 04:19:55 -07:00
David Dudok de Wit a81854a2bd fix(desktop): preserve Bot tabs across owner lifecycles 2026-08-23 04:19:55 -07:00
David Dudok de Wit 856fc66d19 fix(desktop): bound Bot owner wake races 2026-08-23 04:19:55 -07:00
David Dudok de Wit b42d8279ed feat(desktop): retain Bot group drafts by room 2026-08-23 04:19:55 -07:00
David Dudok de Wit 9b36c2d43c feat(desktop): make new Bot tabs owner-aware 2026-08-23 04:19:55 -07:00
David Dudok de Wit 613244cbb1 feat(desktop): organize the global bot roster 2026-08-23 04:19:55 -07:00
Teknium 3638961da8 fix(desktop): Bot Mode keeps Cloud alias identity after hosted handoff
A Desktop per-profile alias (e.g. moxie with a Cloud override) routes to a
remote backend's root profile: route { connectionId, profile: 'moxie',
targetProfile: 'default' }. Once the hosted backend answers the roster
itself, the row's identity is (connection, 'default') — a different key
than the alias meta — so the friendly name regressed to the raw Cloud
hostname after activation, and Cloud-only rosters showed generic 'Hermes'
instead of the configured alias (#89131).

Add a connection-exact alias index built from the credential-free route
inventory, keyed by (connectionId, targetProfile). displayName,
botRosterMeta, and botFriendlyNames consult it so the claimed backend row
reads as the alias (and its title/meta), while:

- same-named defaults on OTHER connections never borrow the identity
- two aliases claiming one backend row fail closed
- the local default and un-aliased remote defaults keep existing behavior

Evidence: @TheAirick's controlled candidate testing on #89131.
2026-08-23 03:56:33 -07:00
Teknium d3e087fd8c feat(bot-mode): bots on every Desktop connection can message each other
Connections ARE the peer set: every gateway connected to the Desktop
(local, remote URL, SSH, Hermes Cloud, docker) is now message_agent-
reachable. The Desktop relays over the persistent sockets it already
holds — roster sync per connection, envelope drain/deliver/reply loops —
so cross-connection DMs work exactly like local ones, replies included.

Also fixes the legacy-SOUL gate bug: profiles whose SOUL.md carries the
old plugin-appended protocol silently lost the message_agent tool
because the injection/execution gates keyed on protocol-section
non-emptiness instead of managed-install.
2026-08-23 02:16:11 -07:00
Teknium e6962b818c style: eslint --fix across src/ + electron/ (CI lints the full tree); drop unused destructure 2026-08-22 22:35:55 -07:00
Teknium 0404020f7b Merge PR #90006: connection-bound Bot Mode actions, reconciled with name-identity + fail-closed canonical resolution
Salvage of saralilyb's remote-bot routing work onto current main:
- kept: immutable (connectionId, profile) owner capture, requestForBot
  routing, backendTargetProfile aliasing, group session owners,
  connection-qualified deletion, focused-owner atoms, remote roster
  merge, Electron profile-delete routing, sdk/store/transcript changes
- reconciled: canonical Bot Chat resolution stays NAME-identity (the
  'Bot Chat' registry row) and FAIL-CLOSED on lookup errors — now
  consulted on the bot's own source via the captured owner route, so
  remote bots get the same no-fork guarantees
- dropped: pointer-pin plumbing (preferredSessionIds, saveBotMeta chat
  writes, pin verification) — superseded by name-identity on main;
  renderer-side remote DM delivery (deliverRemoteRosterMentions /
  pollRemoteDmReply / ensureRemoteCanonicalChat) — superseded by the
  message_agent tool architecture (#91802/#91915: middleware identifies,
  never delivers); pointer-era test files deleted on main
- openStoredBotChat/createCanonicalChat: remote opens keep Desktop's
  chrome home (keepAllProfilesScope: true on routed opens); local bots
  keep the measured workspace re-home
- prepareBotSource: capability gate only — routed RPCs never require
  activation authority
2026-08-22 22:24:54 -07:00
Teknium 87b645f52c fix(desktop): a failed Bot Chat registry lookup no longer forks the bot's forever chat
findExistingCanonicalChat() swallowed every lookup error and returned
null — indistinguishable from 'this bot has no Bot Chat yet' — so a
transient RPC failure against a just-restarted backend (the exact
post-desktop-update window) sent createCanonicalChat() straight to
session.create, minting a fresh 'Bot Chat' while the real one (data
intact, hidden) still held the canonical title. Users experienced this
as bots losing all context after every desktop update.

The lookup now fails CLOSED: a failed registry consultation throws,
both open paths surface their existing 'try again' toast, and
session.create can never fire off an unknown ownership state.

Tests: two new VM-executed regression tests (sabotage-verified — both
fail with the old fail-open catch); hide-bot-chats source-shape regex
updated for the new layout. 364/364 plugin tests green.
2026-08-22 21:24:30 -07:00
Mauvis Ledford e95dd466b1 fix(bot-mode): persist canonical chat before opening 2026-08-22 02:35:47 -07:00
Teknium a9860d413d fix(bot-mode): the canonical Bot Chat is found by NAME — session-id pins removed
A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.

Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.

Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.

No migration shims: this IS the system.
2026-08-22 01:23:39 -07:00