Commit Graph

4128 Commits

Author SHA1 Message Date
Teknium e9dd0bf5d5 feat(desktop): polish bot roster sections — dialog rename, Undo delete, Esc-cancel drag, nested under gateways (salvage #100745)
Follow-up on @fortun8te's user-made roster sections:

- Sections start empty: no seeded General/Workforce/Clients. With no
  sections created the roster renders exactly as before.
- New section and Rename go through one Dialog + Input + Cancel/Save
  (the app's session-rename shape) instead of an inline caret; the row
  menu's "New section…" files the bot as it creates.
- Delete needs no confirmation: bots return to Unassigned and the toast
  offers Undo (restores the section in its slot and refiles its bots).
- Drag: single-row drag under a private MIME type, every valid target
  shows a faint outline while a drag is live, the hovered target lights
  up, the source section refuses the drop, Escape cancels, and the moved
  row no longer stays faded after it remounts under its new section.
- Multi-select (cmd/shift-click, querySelectorAll shift-range) dropped:
  the roster has no selection model. Per-bot saveBotMeta writes run in
  sequence, one per profile (membership IS a field on each profile).
- Section heading reuses RosterSectionHeader (gains `action` /
  `onDoubleClick`), so user sections fold and look like the gateway
  headings; ⋯ menu and right-click drive the same Rename / Move up /
  Move down / Delete. Empty sections show a dashed "Drag bots here" slot.
- Composes with gateway buckets: sections nest INSIDE each connection
  bucket, indented under a hairline rail (membership lives in the bot's
  profile on that gateway); empty sections repeat there only mid-drag.
- Full i18n parity (en / ja / zh / zh-hant) for every new string; icon
  toggle and the storage-async plumbing removed.
- Tests trimmed to the three invariants (membership persists through
  saveBotMeta + reload, remainder = Unassigned, delete returns bots +
  undo) plus a live Electron e2e covering the whole flow.
- Docs: "Organize bots into sections" in user-guide/bot-mode.md.
2026-09-02 06:06:32 -07:00
Michael Knaap bd9955d529 fix(desktop): section rename — Escape cancels, Enter commits once
Closing the rename field unmounts the input, and the unmount can still
fire onBlur, which committed the draft the user had just asked to throw
away with Escape. Enter also called commit() directly and then again from
blur. Route both through blur with a cancelled flag so the commit runs
exactly once and Escape never renames.
2026-09-02 06:06:32 -07:00
Michael Knaap 3d0ac691af feat(desktop): user-made sections in the bot roster, with drag-and-drop filing
The roster already has sections, but only automatic ones: one per gateway
connection plus the group-chat bucket. Those answer "where does this bot
run", which is not the question being asked when someone wants two client
bots filed together under "Clients" and the internal ones under "Team".

This adds a second axis that composes with the first: gateway sections keep
the top level whenever more than one connection is showing, and user
sections group the flat list underneath.

Design choices, each deliberate:

- Membership lives on the BOT (`ui_meta.sectionId`), not as a member list on
  the section. A bot can only be in one place, deleting a section cannot
  orphan anybody, and the assignment rides the same profile.yaml sync every
  other bot setting already uses, so it follows the profile to another
  machine. Section records (id, name, icon) live in plugin storage.
- "Unassigned" is not a section. It is whatever is left, always drawn last,
  and it is where members of a deleted section land. No record, so nothing
  to keep in sync.
- Three gestures, one rule: drag a row onto a section heading; cmd/ctrl-click
  and shift-click build a multi-selection (shift ranges in DOCUMENT order,
  anchored Finder-style); and the row's context menu gets "Move to section…"
  with the same targets the drag would use. Dragging a row that is part of
  the selection drags the whole selection.
- The drag uses a private MIME type, so a bot dropped on the composer or the
  transcript is simply not a valid payload there instead of pasting its key
  as text.
- Section headings rename inline (double-click / menu), reorder, hide their
  glyph, and delete (keeping their bots). Right-click and the ⋯ button open
  the same menu so neither can drift.

With no sections created the roster renders exactly as before.

Tests: user-sections.test.ts covers the pure model (normalisation, grouping
with unknown/deleted sections falling to Unassigned, drag payload
round-trip). The existing hermes-bots suite passes; tsc and eslint clean.
2026-09-02 06:06:32 -07:00
Teknium 648c664eb3 feat(desktop): gate cold-start restore on display.resume_last_session
- use-desktop-integrations: hold the restore latch until the config
  record answers; when false, stay on the fresh chat (route and session
  restore alike) while still remembering the open chat for next launch.
- wiring: read the shared config-record query; undefined while pending,
  fetch failure falls back to the historical behavior (resume).
- appearance-settings: ToggleRow writing through the shared config
  cache with rollback + notifyError on a failed save.
- ar/ru strings, docs line in user-guide/desktop.md, two hook tests.
2026-09-02 05:56:54 -07:00
jinglun010 7aff724e56 feat(desktop): add display.resume_last_session config toggle (#60812)
Config default (true) plus the Appearance-settings strings for a
"Reopen Last Chat on Launch" switch. Salvaged from PR #60816 onto
current main (defaults moved to config_defaults.py since the PR).
2026-09-02 05:56:54 -07:00
liguoyu 1592e48ac9 fix(desktop): a split-dragged Bot tab keeps its Bot workspace scope
Dragging a session tab to split in the Bot workspace committed through
openSessionTile with no scope, whose default { workspaceMode: 'sessions' }
was written onto the already-open tile before the pane move — so the tile
re-bucketed into the Sessions workspace and vanished from the Bot strip.

A scope-less open of an already-open tile is a MOVE, not a re-scope:
default the scope to the tile's current workspaceMode and only rewrite
scope when the caller passed one explicitly (sidebar/bot openers still
win). Brand-new tiles keep the 'sessions' default.

Closes #96865
Supersedes #96998

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-02 05:54:10 -07:00
Teknium d1f8c2ea11 test(desktop): group chat picker row label opts out of min-width:auto so names truncate
Renders CreateGroupChatDialog with a long-named bot and asserts the row
label carries min-w-0 (and the inner text column keeps min-w-0 flex-1 +
truncate). Sabotage-verified: fails against the pre-#90624 markup with
'expected [...] to include min-w-0'.
2026-09-02 05:42:14 -07:00
Jay c65c79a4a8 fix(bot-mode): group chat picker rows overflow and scroll names out of view
The New Group Chat picker's rows are `label` flex containers holding a
`min-w-0 flex-1` text column whose two lines are `truncate`. The label
itself has no `min-w-0`, so as a flex item it keeps its `auto` minimum
width and cannot shrink below its content. `truncate` therefore never
fires: the row grows to fit the longest secondary line instead, which is
`@handle · in "Group A", "Group B", …` and so scales with how many groups
a bot already belongs to.

The row is a grid item inside a Radix ScrollArea, whose viewport wraps
children in a `display: table` div that sizes to content, so the whole
list widens rather than clipping.

Measured on a 414px viewport with one bot in three groups:

  wrapper width  593  (viewport 414)
  widest row     585
  overflows      yes

Nothing is visibly wrong until the first click. The checkbox now sits
past the right edge, so focusing it scrolls it into view: `scrollLeft`
jumps 0 → 178.38 (= 593 − 414) and every row shifts to `left: -162px`,
clipping the bot names from the left — the user clicks a name and the
names disappear. The scroll offset persists after unchecking, until the
dialog is remounted.

Adding `min-w-0` to the label lets it shrink, so `truncate` engages as
the markup already intended. Same viewport, same data:

  wrapper width  414  (unchanged display: table)
  widest row     406
  overflows      no
  scrollLeft after clicking a row  0 → 0

`display: table` on the ScrollArea wrapper is untouched; the fix works
with it rather than around it. Verified against a packaged build via CDP,
before and after, on identical roster data.

No test. The rule for this is "extract the logic into a small
pure/DI-testable function and call it for real", but there is no logic
here — `min-w-0` is a class name, and the behaviour under test belongs to
the layout engine. jsdom does not lay out, so a unit test cannot observe
the overflow; the Playwright suite could, but has no bots-roster fixture,
which is a large scaffold to hang off a one-class change. A source-regex
assertion would pass without ever laying anything out, which is precisely
the false confidence AGENTS.md describes. The before/after measurements
above are offered as the evidence instead — happy to add a Playwright
case if you'd rather have the fixture.
2026-09-02 05:42:14 -07:00
Teknium 8d6a286fe8 fix(desktop): restored background tabs resolve their session title without a click
A restored session tile has no runtimeId and never mounts its pane until first
activation, so the by-id resolution effect inside SessionTilePane never runs.
When the row is also outside the recents page and project tree, tileTitle()
falls back to "New session" until the user clicks the tab (#94167).

Add a one-shot backfill, wired next to watchSessionTiles(): once the gateway
is open, look each unrestored, untitled, unlisted tile up via
resolveStoredSession(id, tile.ownerRoute). That call already upserts the row
into $sessions, which the tab strip watches, so the tab renames itself —
nothing new is persisted and workspaceTabTitle stays the Bot Chat marker.

Live repro (Electron e2e, target session pushed off the 50-row recents page
by 60 newer sessions, restored as a stacked background tab): main showed
"New session" after boot with no click; with this fix the tab reads the real
title while the pane is still unmounted.

Closes #94167
Supersedes #94212

Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 05:42:00 -07:00
Teknium ad800ea8cd fix(desktop): cold resume paints the prefetched REST transcript before session.resume settles
The REST prefetch and the gateway `session.resume` already ran concurrently,
but the prefetch result was held until the runtime resume settled. A cold
profile build (skills / MCP / memory) can keep `session.resume` pending past
the hydration budget while the complete transcript is already in hand, so a
Bot Chat sat on the loader and burned its retries with readable history
off screen.

- Publish the grafted REST snapshot as soon as the prefetch resolves and
  `isCurrentResume()` holds; the runtime path grafts only its live projection
  onto that same snapshot, and the post-resume `chatMessageArraysEquivalent`
  skip keeps reference identity when nothing changed (no second DOM build).
- Stamp the eagerly painted page with persisted-display provenance on the
  runtime state so the warm-path gate admits it on the next switch.
- REST fallback after a resume rejection skips the redundant re-publish when
  the early paint already shows the transcript.

Live repro (Electron e2e, session.resume stalled 25s via a temporary
backend shim): main never painted within 20s; with this fix the transcript
painted 150ms after the row click.

Supersedes #90130.

Co-authored-by: Alexandre Roumieu <269586168+alexandreroumieu-codeapprentice@users.noreply.github.com>
2026-09-02 05:41:47 -07:00
Teknium 1bf2cf57ef fix(desktop): project tree follows sessions.changed so external session create/delete/rename/cwd changes show up (#100354) 2026-09-02 05:40:18 -07:00
Teknium 209de12d5f fix(desktop): Bot Mode tabs caption a Bot Chat with the bot's name, not "Bot Chat"
Every bot's canonical chat is stored under the same title ("Bot Chat" — the
name the gateway resolves it by, and an invariant roster-actions.ts's stale-tile
probe and #90102 rely on), so the main tab strip captioned every open bot chat
identically and two bots' tabs were indistinguishable (#99152).

Fix at the presentation layer, leaving the stored title and tabTitle untouched:
- workspace-scope.ts gains `$workspaceOwnerLabels` + `workspaceOwnerTitle()`:
  a bots-mode tab whose resolved title still equals its registered placeholder
  reads its owner's label instead. Side threads / Sessions tabs are untouched.
- session-tile.tsx captions tiles through it (and the drag payload); the main
  `workspace` tab (controller.tsx) does the same via `$botChatScopes`, the
  bot-mode scope the main tab was last opened under (it has no tile).
- The hermes-bots roster publishes displayName() per owner key through the new
  `host.setWorkspaceOwnerLabel` (feature-detected), so renames follow.

Supersedes #99177, which set tabTitle at open time — that reverts after mount
because tileTitle() prefers the stored row's title once the hidden row is
upserted, and breaks the `workspaceTabTitle === 'Bot Chat'` invariant.

Tests: one unit test on workspaceOwnerTitle() (bot chat → bot name; side
thread / sessions tab / unlabeled owner untouched) and one Electron e2e
(tab strip reads "Alpha", not "Bot Chat"); both fail on main, pass here.

Closes #99152
Supersedes #99177

Co-authored-by: twotnguyen <nguyenngoctinh011258@gmail.com>
2026-09-02 05:38:10 -07:00
Teknium 7b951e46ab fix(desktop): route a shared-id unpin to the row the pull adopted
With every slice feeding the pin sync, two profiles can legitimately hold
the same session id. The pull already tie-breaks toward the active gateway's
row (rowsByPinId), but the push resolved the profile by first match — so an
unpin PATCHed the other profile and the next page re-adopted the pin.
Resolve the write's row with the same active-gateway preference.

Case and test from #92609.

Co-authored-by: Jake Vincent <45184202+jakewvincent@users.noreply.github.com>
2026-09-02 05:37:58 -07:00
fangliquanflq d417aee387 fix(desktop): sync pins across all session slices 2026-09-02 05:37:58 -07:00
Teknium 9b84a98e29 test: trim salvaged #99763 to the two invariant-pinning tests
Drop the no-field fallback vitest case and the get_compression_chain
unit test: the fallback is implied by the predicate (Boolean(undefined))
and the chain walk is covered by test_list_serves_full_lineage_ids_for_projected_rows
through the real list projection.
2026-09-02 05:37:05 -07:00
Alonso 202997b51d fix(desktop): one conversation never opens as two tabs after compaction
Compression rotates a conversation's tip id while tiles stay keyed by
whichever segment id they were opened with. focusOpenSession and
openSessionTile tested exact ids, so right after a rotation the same
chat read as 'not open' and opened again in a second tab — and a tile
keyed to a MIDDLE segment (the tip when it was opened) could no longer
prove it names the conversation at all, rendering as an untitled ghost.

The projected list row now carries the full chain
(SessionDB.get_compression_chain, served as _lineage_ids by
list_sessions_rich and the sidebar tree row), lineageAliases indexes
every segment, sessionMatchesStoredId accepts membership, and the tab
focus/open paths dedupe through the lineage instead of the exact id.
Older gateways omit the field and degrade to today's root/tip pairing.
2026-09-02 05:37:05 -07:00
Gille 2599793271 fix(bot-mode): hand off group tabs to remote bots 2026-09-02 05:36:54 -07:00
chelsealong 87d5e40f53 fix(desktop): fail closed when Bot Chat lookup returns zero rows
session.list can succeed with an empty sessions array during a profile
backend restart instead of throwing, and findExistingCanonicalChat's
`rows.find(...) || null` mapped that to the same value as "this bot
never had a chat". The click path then minted a replacement Bot Chat
and re-fired the kickoff intro on an intact, hidden canonical row,
orphaning in-progress work each time (#98383).

When the roster's own canonical_session already confirms this profile
has a Bot Chat, treat a zero-row result as unconfirmed absence and
fail closed the same way a thrown RPC error already does, instead of
minting.
2026-09-02 05:36:54 -07:00
Teknium 6d061ede58 fix(desktop): open transcript catches up on every reconnect
Closes #94779. A turn that finished while the gateway socket was down never
replays its sessions.changed tick, so the open transcript stayed stale
until the user reopened the session. The gateway-open effect now also
requests one signature-gated tail of the active transcript on every
(re)connect (connection-scoped, so a plain session switch adds no read;
messaging transcripts already refresh on open via their own effect).

Reported-by: Kkkkkuro
2026-09-02 05:36:41 -07:00
Teknium a0e4ce252c fix(desktop): open cron-run transcripts stay live
Reimplements PR #89728 on top of the ownerRoute rework. The active
transcript resolver only looked in $sessions and $messagingSessions, so an
open cron-run transcript resolved to nothing and every sessions.changed
tick was dropped — the pane froze until the user reopened it. Resolve
through ownerLookupSessionRows() (recents + cron + messaging) instead.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-02 05:36:41 -07:00
Teknium a53286999b fix(desktop): bot tile transcripts read from their owner backend
Salvages the client half of PR #99333. A workspace tile pinned to an exact
owner (connection + target profile) was reconciled via a bare
getLatestSessionMessages(storedId) — the foreground profile's backend — so
a bot tile on another profile never saw its new turns (or saw the wrong
session's). reconcileTileTranscripts now derives a ProfileScope from
tile.ownerRoute and keys the per-tile signature by that route; route-less
tiles keep the legacy local read.

The tui_gateway/server.py `_sessions_sig` cross-profile scan from #99333 is
intentionally not taken: the desktop runs one backend per profile, each with
its own watcher, so scanning sibling profiles would misattribute ticks.

Co-authored-by: stods21 <stods21@users.noreply.github.com>
2026-09-02 05:36:41 -07:00
Dolverin 51953a302f fix(desktop): keep busy tile transcripts stable 2026-09-02 05:36:41 -07:00
Teknium aaa34b0e08 fix(desktop): model picker no longer hardcodes --global; one persist policy for every surface (#90235)
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).

Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).

Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
  session-only rule: when neither model.default nor model.provider is
  configured yet, persist. This preserves #86414's first-pick motivation
  for CLI, gateway and Desktop alike. With a default configured, a plain
  pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
  Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
  lists --global.
- Docs: desktop.md picker note + slash-commands /model row.

Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
2026-09-02 05:33:33 -07:00
Ayush Nangia 3a980a431b fix(desktop): keep drafts editable while connecting 2026-09-02 04:36:51 -07:00
Teknium 6e7c7c7da9 fix(desktop): a bot row click always lands on the Bot Chat the row previews
A plain roster click fronted whatever bots-workspace tab the user last had
active for that bot (#96649). A '+' side thread persists in Local Storage
across restarts, so it won every click forever while the row kept previewing
the canonical Bot Chat (profiles.list canonical_session) — sidebar and center
described two different conversations; a message typed there landed in the
side thread and the row never moved. Support thread "[Bots] - Sessions is not
in sync again" (bundle 7dfff039), reproduced live on origin/main.

- roster-actions: the open-tab shortcut may front only the canonical chat
  (registry id or lineage tip, via a new onlyStoredIds allowlist on
  focusWorkspaceOwnerSessionTile); anything else resolves the registry and
  opens in place. Side tabs stay open beside it. "Open Bot Chat" in the row
  menu is the same action; the `canonical` option goes away.
- roster-actions: when the FOCUSED Bot Chat's canonical session advances on
  the gateway (cron bot-chat delivery, message_agent, group round, CLI turn —
  none reach this window's stream), re-open it in place so the transcript
  refreshes instead of waiting for an app restart (#99393 class).

Tests: the fronting-shortcut unit file and its e2e spec pinned the reversed
behavior; replaced by one unit file (5 tests) and one e2e spec that fails on
main and passes here. group-to-local-bot-handoff e2e still passes.
2026-09-02 03:41:44 -07:00
hermes-seaeye[bot] 254158f453 fmt(js): npm run fix on merge (#101150)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 10:12:08 +00:00
Teknium 3e63367175 feat(desktop): status bar can show live cache-hit rate and tokens/sec (off by default)
Two new right-click-toggleable status bar items, mirroring the CLI/TUI
Pantheon status bar upgrades: prompt-cache hit rate ("87%") and rolling
output throughput ("42 t/s"). Both are hidden by default and enabled from
the bar's existing 'Show in status bar' context menu, like the context meter.

Renderer-only: the tui_gateway already emits cache_hit_pct and avg_tps in
every session.usage tick and message.complete payload, so the items ride
the same UsageStats the context meter reads — no new RPC, no polling.
Labels show a placeholder until the backend has data, never self-hide.
2026-09-02 03:06:20 -07:00
Teknium 5d4aa4fcb2 fix(desktop): group chat rooms are serial again; keep only the push-woken turn poll
#101112 made round members take their turns concurrently. That changed what
a group chat IS: later speakers in a round no longer saw earlier speakers'
replies, so bots answered the user independently instead of building on
each other. Group rooms are serial round-robin by design — this restores the
pre-#101112 round engine (group-rounds.ts, group-chat.ts, group-chat-view.tsx,
their tests, and the docs) byte-for-byte.

What stays from #101112: the per-turn poll wakes on the member session's
terminal frame (message.complete / error via host.onEvent) instead of
sleeping a fixed 2s between session.resume reads; 5s timer kept as backstop.
That is a pure latency fix with no change to room semantics.

Live A/B (real tui_gateway over WS, 4 members, one serial round):
2s poll 32.5s -> push-woken 22.5s. The remaining time is model latency.

Refs #92760
2026-09-02 02:54:46 -07:00
Teknium fb5023950e perf(desktop): group chat rooms answer in the time of one bot, not the sum of all
Bot Mode group rooms were slow by construction: the round engine ran every
member's turn one after another, and each turn found out its bot had finished
by re-reading session.resume on a fixed 2s timer. A 4-bot room paid
4 x (model latency + up to 2s) per round, serially.

- group-rounds: members of a round now take their turns concurrently
  (Promise.all). Rounds stay serial so bots still build on each other's
  replies. Each member's delta is computed at its own turn start and its
  watermark advances only to the pre-turn log length, so sibling replies
  that land while it thinks are delivered next round exactly once; a
  member's own replies are excluded from its delta by author (they are
  already in its session). Message cap enforced per round; stop path
  interrupts every member mid-turn (room.turn -> room.turns map).
- group-turns: the poll wakes on the member session's terminal frame
  (message.complete / error via host.onEvent), then re-checks at 250ms
  until session.running clears. The timer poll stays as a 5s backstop for
  hosts without the event tap. Feature-detected; node test harness unaffected.
- group-chat-view: "X is thinking..." lists every member mid-turn.
- docs: bot-mode.md describes concurrent rounds + push-woken replies.

Live A/B (real tui_gateway over WS, 4 members, one round, same model):
serial+2s poll 35.0s -> concurrent+push 8.6s; every turn woke on the event.

Refs #92760
2026-09-02 02:36:26 -07:00
hermes-seaeye[bot] fbd40c907d fmt(js): npm run fix on merge (#101107)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 08:51:02 +00:00
hermes-seaeye[bot] 6840bb02e8 fmt(js): npm run fix on merge (#101102)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-02 08:42:10 +00:00
Teknium aa80626764 fix(tui-gateway): adopt late compute-host compress acks instead of a false 120s timeout (#97948)
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.

- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
  registered when the waiter times out; control.ack/control.error/error
  frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
  crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
  `compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
  cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
  applies the metadata mirror and emits the same `session.info` a normal
  compress does plus the existing `status.update kind=compacted` edge; a
  late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
  on waiter timeout answer `status: pending` (not 5019) and register the
  late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
  `status: 'pending'` renders as an info notice, not `error:`; the
  `compacted` status edge rehydrates an idle active session's transcript
  (mid-turn compaction still defers to the turn settle path).

Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.

Refs #97948

Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
2026-09-02 01:35:59 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium df4b3733ba fix(cron): every last_status consumer renders delivery_failed explicitly (dashboard badge, Desktop inspector, /cron list, docs)
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):

- web dashboard CronPage: last_status was never rendered at all — a
  delivery_failed job showed a green 'scheduled' badge and only a small red
  'delivery: ...' line. New pure cronLastResult() helper maps the closed
  literal set to tones (ok=success, delivery_failed/blocked_config=warning,
  error/unknown=destructive) and the card now shows an amber
  'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
  literal; routineLastResult() spells out each one ('Ran, but delivery
  failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
  appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
  on this branch; no consumer compared == 'ok' for success apart from the
  cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
  detail field carries the reason.

Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
2026-09-02 00:52:58 -07:00
Teknium 55d8c054eb test(desktop): pin sole-local registry approval routing through the event owner (#96394)
Regression for the single-connection/single-profile report: hasRegistryTopology()
is true on every modern Desktop, so the ambient escape hatch stays closed; the
approval.request event's own (connectionId, profile) stamp is what routes
approval.respond back to the primary socket.
2026-09-02 00:37:00 -07:00
liuhao1024 7f53ad1741 fix(desktop): route approval responses through the runtime event's exact owner
recordSessionEventScope already captures the exact (connectionId, profile) a
runtime's inbound events proved, but knownOwnerForSession never consulted it:
with no tile/hint/row binding for the runtime id, approval.respond failed
owner resolution (SessionOwnerResolutionError) even though the event source
itself named the owner.

Add a structured owner twin of the scope ledger, written and cleared with it,
consumed as the LAST rung of knownOwnerForSession so durable stored identity
still outranks it and untagged/unknown runtimes keep failing closed.
2026-09-02 00:37:00 -07:00
Teknium fdb2e10a8e fix(desktop): refuse the build when ANY declared non-optional dep is missing
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).

Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
2026-09-02 00:07:58 -07:00
kshitijk4poor ae0418599d chore: export BUILD_CRITICAL_PACKAGES for the test, drop dead default export
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
2026-09-02 00:07:58 -07:00
Jack Lau 15e74fa625 fix(desktop): guard the whole build-critical dep set, before clean
Refs #86443

assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.

Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.

Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.

The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.

The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
2026-09-02 00:07:58 -07:00
kshitijk4poor 1df78c2da8 refactor(desktop): review follow-ups for the stale-RPC salvage
- ambientRequestFor(gateway): one adapter for the six copy-pasted
  `<T>(method, params) => gateway.request(method, params ?? {})` lambdas
  the pollers built to call requestForOwnedSession.
- resetBackgroundPollingGuard() with no argument (primary reconnect,
  use-gateway-boot.ts) now also clears healsByStoredId. The latch and the
  heal budget share a lifetime: a respawned backend re-mints every runtime
  id, so a stored session that had exhausted its 3 heals must be healable
  again. Previously only the latch was cleared there.
- Leaf module comment states the actual import convention.
2026-09-02 12:31:33 +05:30
kshitijk4poor 9ed8331ced style(desktop): eslint --fix + prettier on the salvaged files; null-guard respondToApprovalAction latch 2026-09-02 12:31:33 +05:30
kshitijk4poor d68739c2fa fix(desktop): stop approval.pending replays against a gone active runtime
The approvals loop and broadcast_session_info fan out UNSCOPED session.info
frames for every live session. The event router attributes an unscoped
frame to the active session, and approvalReplaySessionId then re-pulls
approval.pending for it. When the active runtime is already gone (4001),
that made every fan-out tick a fresh dead-id request — the dominant source
of the 1,369 post-restart approval.pending rejections in #100639.

approvalReplaySessionId now takes the frame's explicitness and a gone
predicate and returns null for an unscoped replay onto a latched runtime.
An explicitly scoped frame is the runtime speaking for itself and is never
skipped.

Refs #100639
2026-09-02 12:31:33 +05:30
kshitijk4poor 3f87a8090c fix(desktop): fold the stale-RPC guard into runtime-gone and refund heal budget on rebind
The salvaged #95647 commits predate runtime-gone.ts and shipped their own
gone-latch (session-rpc-guard.ts) beside the one main already had. Fold
them: one latch, one classifier, one clear seam.

- session-gone-latch.ts is a dependency-free leaf holding the latch, the
  4001 classifier (now also rejecting "mentions session not found" tool
  strings and unwrapping IPC bridge prefixes), and the rebind seam. It
  exists because session-request-router — imported by every store — must
  clear the latch after a successful session.resume/activate without
  pulling the session/tile stores into its import graph.
- runtime-gone.ts re-exports the leaf and keeps the heal logic.
- A successful rebind now also refunds the stored session's heal budget.
  markRuntimeGone caps consecutive heals at 3 per stored id and only a
  successful process.list refunded it, so a backend that reaps a detached
  runtime a few times left the view stuck on a phantom id with every
  poller latched and the socket-reconnect global clear (removed by the
  salvaged commit) re-arming the storm. #100639: 1,230 approval.pending
  4001s on one runtime id in 42 minutes, zero recovery.

Refs #100639
2026-09-02 12:31:33 +05:30
Shinji Kaneko c37181a4ec fix(desktop): preserve transient approval errors 2026-09-02 12:31:33 +05:30
Shinji Kaneko 0a80adc127 fix(desktop): harden stale session RPC guard 2026-09-02 12:31:33 +05:30
Shinji Kaneko 16021395f9 fix(desktop): stop stale session RPC retries 2026-09-02 12:31:33 +05:30
Teknium 94db305b1a fix(desktop): drop the 'Switch models mid-thread' tip bubble
The model pill already reads as a button; the tip is noise over the composer.
Removes the catalog entry and its strings in all 5 locales.
2026-09-01 23:29:32 -07:00
Teknium 4e3feb8bbb feat(tts): speech toggles warm up and unload local TTS engines (#100881)
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.

- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
  acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
  registry; piper/kittentts loaders extracted so warm-up and synthesis
  share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
  failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
  intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.

Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
2026-09-01 21:43:59 -07:00
Brooklyn Nicholson 2659c917c1 refactor(desktop): make the tile owner ladder testable behavior
session-tile-owner-route.test.ts asserted against the TEXT of
session-tile.tsx, so it passed on a broken implementation whose call site
merely looked right, and failed on this refactor, which changed nothing the
tile actually does. AGENTS.md bans the pattern outright.

Extracted the ladder as tileOwnerRoute() and replaced the three regex
assertions with six that call it: tile route wins, row falls back, hint
falls back, targetProfile carries through, a bare profile narrows away, an
untagged session stays ambient. 12s of source-matching becomes 1.1s of
behavior.
2026-09-01 22:58:48 -05:00
Brooklyn Nicholson 7f27ca2701 test(desktop): two connections serving one parent id are not one branch
The coalescing key now carries the owner, so two backends that both expose a
session called `parent` get their own create instead of the second caller
receiving the first's child. Both creates are held open, which is the only
state the key guards — a sequential version passes even with the owner
stripped out.

Co-authored-by: Ahmett101 <ahmet.tunc@gmail.com>
2026-09-01 22:58:48 -05:00