Every caller tears the SSH connection down right after `await
cancelAndWait(scope)`, so resolving on this call's own barrier alone let a
connection-apply teardown overlap the pool stop's still-running afterStop
teardown for the same key. Wait for the composed barrier instead; chain the
barriers (drain promises never reject, so this equals allSettled). The test
now waits a macrotask and asserts that neither the apply nor the new
bootstrap runs before the first teardown completes.
A pool stop blocked in teardownSshConnection() holds drains[scope]; a
concurrent connection apply calling cancelAndWait() for the same scope
replaced that barrier with its own and, having nothing to drain, cleared the
map entry as soon as it finished. start() then saw no drain and began a new
bootstrap while the first SSH teardown was still running.
cancelAndWait() now composes with the drain already in flight and the entry
is cleared only when the composed barrier settles, so start() waits for
every active teardown. Regression test reproduces the pool-stop / apply
race; it fails on the previous coordinator.
Reported by ehz0ah on #110025 (follow-up to #106935).
A dead tunnel was redialled every 2 s until the scope was torn down;
delays now double from the base to a 30 s cap and reset on open. The
generation counter duplicated what the entry map and entry.socket
already say (stop() removes the entry, connect() replaces the socket),
so abandonIfStale reads those instead. afterStop's unused entry
parameter is dropped.
afterStop awaited cancelAndWait(teardownSshConnection) with no catch; the
idle reaper calls stopPoolBackend un-awaited, so a teardown rejection
became an unhandled rejection on the main process. Log it through the SSH
log and let the fence release.
Four registry tests overlapped: sibling independence is a two-line
assertion inside hold-until-stop, and the reject-missing-target case is
the setup half of the empty-scope primary case. Folding them keeps every
invariant covered (hold, no-reconnect-after-stop, sibling isolation,
''-is-a-real-key, baseUrl/token required) with fewer fixtures to keep in
sync. The pool-stop teardown fence is already covered by
pool-stop.test.ts::afterStop and ssh-bootstrap-coordinator.test.ts.
Electron bundles a global WebSocket in the main process; a missing
constructor is a build/environment error, not a runtime state to route
around. The two silent `typeof WebSocketImpl !== 'function'` returns in
start()/connect() would turn that build error into an armed-looking
registry that never opens a socket, letting web_server_idle_exit retire
an owned SSH-isolated sibling with no log line — exactly the bug this
module exists to prevent. Let `new WebSocketImpl(url)` throw into the
existing catch, which logs and schedules a reconnect.
Review banned source-reading main.ts greps and asked the remaining registry
cases down to ~4: hold-until-stop with no reconnect, empty-string primary,
reject missing target, and sibling isolation. Prettier the helper.
Co-authored-by: Cursor <cursoragent@cursor.com>
Process-less SSH entries finish child exit immediately. Keep inFlight
and the bootstrap drain up until keepalive teardown completes, and
stop reading main.ts from tests.
Co-authored-by: Cursor <cursoragent@cursor.com>
Desktop keeps live sockets on the profile-less backend while the --profile
sibling only sees short RPC sockets; idle-exit then retires the sibling and
the app restarts broadly. Arm a main-process /api/ws keep-alive for every
published sshConnections scope (including '') until teardown, without
treating spawn artifacts as liveness (#101626).
Co-authored-by: Cursor <cursoragent@cursor.com>
A renderer built after d9834a3e86 listens for srq- request frames; a v6 backend
still emits <kind>.request notifications, so every approval/clarify card would
silently never render. The skew toast now points the user at the backend update.
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
The staleness test regenerates in the Python lane, which has no
node_modules; prettier-dependent output would make the check pass locally
and fail in CI (or the reverse). Single-quoted literals, bare identifier
keys, no trailing commas or whitespace — prettier --check is clean on the
committed file.
apps/shared/src/gateway-events.ts is now a thin layer over
gateway-contract.generated.ts (client-local synthetic events + the
GatewayEvent envelope); gateway-events.json, its two rendezvous tests and
the duplicated BillingBlock / SessionInfo / ProjectInfo hand copies are
gone. Desktop, TUI, web and shared typecheck against the generated
RpcMethods / ServerRequestMap / BackendGatewayEventMap.
What tsc found once the types were honest: three phantom fields the
backend never sent (tool.start.todos, error.reason,
voice.transcript.voice_stopped) - the TUI todo tests were driving the
list through the phantom and are retargeted to tool.complete, where the
wire actually carries it; nullable fields (`None` on the wire) were typed
as plain optionals in eight places and now coerce at the boundary;
SessionResumeResult had a stale generic.
Contract fixes from the consumer pass: TranscriptMessage is the gateway
projection (text/row_id/context/args), not the stored row; SkinPayload
matches HermesSkin (empty-string defaults, never null); SessionLiveInfo
model/tools/skills are required (always emitted); BillingBlock.billing_url
is required-nullable (dataclass asdict).
tui_gateway/AGENTS.md documents the declare -> regenerate -> tsc loop.
Handlers own their documented domain codes (4006 missing session_id, 4015 bad
url, 4009 orphan claim); the contract's job on the way in is the one check no
handler performs — an unknown key (4000 with the key path). Missing/mistyped
fields are re-checked AFTER a successful handler answer under the strict
test policy, so a contract narrower than the wire still fails the suite.
Two models widened from the suite: SeedMessage (clients forward stored rows
verbatim), tool.complete.args (mirrored child rows omit it). Tests that
drove session.activate with prompt params (and vice versa) or stubbed
_live_session_payload with a bare {session_id} now send the real shapes.
215 methods, 13 server→client requests and 67 notifications now have Pydantic
contracts under tui_gateway/contracts/<topic>.py, rendered to
apps/shared/src/gateway-contract.generated.ts (616 types) and
gateway-contract.openrpc.json. tests/contracts/test_generated.py pins both
files to an in-memory regeneration and asserts catalog completeness from the
CODE side (every registered handler / emitted event / sent request has a
contract, nothing orphaned). scripts/ci/classify_changes.py runs the Python
lane when either generated file changes.
Phantom fields the hand-typed TS carried and no emitter ever set:
tool.start.todos, error.reason, voice.transcript.voice_stopped.
The old suites asserted the deleted wire (`*.request` events, `*.respond`
RPCs, `pending_clarify` snapshots, `_pending`/`_answers` teardowns). Each
test keeps its invariant against the new shape: a seeded live request's
`respond` spy receives the answer object, `hasOpenServerRequest` flips, the
`approval.respond` RPC fallback is asserted ONLY for queue entries restored
without a socket, and Bot Mode rooms answer via `request.answer` /
`clarify.lock`. The group-turns test that polled forever for a
`clarify.respond` that no longer exists (20-minute hang) now completes.
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.
`tui_gateway/server_requests.py` owns the one mechanism:
send() block the agent thread until the response frame
(`srq-<n>` ids; ints belong to the client)
send_async() fire-and-callback variant (bot relay)
cancel*() withdraw with ONE `request.cancel {id, method, reason}`
event (timeout / interrupt / process exit /
answered elsewhere) instead of per-kind *.expire
open_requests() the still-open frames, replayed by session.resume,
session.activate and session.events.since so a
reconnecting client re-renders every kind, not two
clarify.lock stays a real client→server RPC (locks one batch
answer early); locked answers merge into the final
set even when the closing response carries only the
tail the user answered last
A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.
Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.
Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
The backend never sent a JSON-RPC request; when it needed an answer from the
renderer it hand-correlated a `*.request` notification with a later `*.respond`
method through four module-level dicts, a timeout thread and 13 derived
`*.expire` names, plus a separate reconnect snapshot per prompt kind. That is a
second request/response layer built on a protocol that already has one.
`tui_gateway/server_requests.py` sends `{id: "srq-…", method, params}` and
blocks on the response frame with that id (string ids never collide with the
clients' integer ids). One `request.cancel {id, method, reason}` notification
withdraws a request on timeout / interrupt / session close. `open_requests` on
`session.resume` / `session.activate` / `session.events.since` re-delivers
unanswered requests after a reconnect; the shared TypeScript channel does that
itself before the caller sees the result. Batch clarify keeps its per-question
locks as a normal `clarify.lock` RPC (the last lock resolves the request).
Approvals stay queue-backed (`tools.approval` owns the timeout, `/approve all`,
coalescing): the request resolves the queue entry and the entry's own
resolution withdraws the request through `register_gateway_settle`.
Deleted: `_block`, `_respond`, `_pending`, `_answers`,
`_pending_prompt_payloads`, `_batch_clarify`, `_EXPIRING_REQUESTS`, the
`*.respond` methods, every `*.request` / `*.expire` event, `pending_clarify`.
Compute-host (turn isolation) mirrors the child's open request and relays the
response frame / lock to it. Desktop, TUI and shared clients register
`onRequest` handlers where they used to switch on `*.request` events; answers
are response frames over the socket the request arrived on, so #91684's
owner-routing class cannot recur for prompts.
Two remaining halves of #89184 (Desktop Settings saves rewriting unrelated
config):
- The `fallback_providers` structured editor normalized every entry down to
`{provider, model}`, so any edit (remove a row, pick a model) re-emitted a
hand-written local-gateway chain without its `base_url` / `api_key` /
`key_env` / `api_mode` — the next autosave persisted bare pairs and the
fallbacks silently routed to the public provider. Entries now carry every
key through; the editor only owns the two selects.
- `PUT /api/model/moa` did `cfg = load_config(); cfg["moa"].update(...);
save_config(cfg)`: the whole default-expanded snapshot went back to disk,
so a Desktop MoA autosave re-persisted every other section too (the
2026-09-10 repro: `fallback_providers: []` written alongside the MoA block
the user had just edited). It now saves `{"moa": ...}` with
`merge_existing=True`, the same section-scoped write every other sparse
writer uses since #110535. Hand-edited moa keys (#58819) still survive.
The `model.default not persisted / base_url cleared` symptom from the 0.20.4
report no longer reproduces on main through the real REST path (Config page
diffs against a baseline since 5361867c6d32; `_denormalize_config_from_web`
keeps the on-disk `model:` block).
Three consumers still read `result === undefined` as an open call or an error without text after the result and display-hint split.
The tool row lost the event's error explanation when the result carried none: `toolErrorText` now reads `toolResultMetadata.error` / `.message` when `isError` is set, and `toolStatus` runs the error path for sealed error rows so an envelope-only read miss stays on the notice tier. The turn-activity signature and wait narration count a sealed call as settled. Onboarding's start-with-connections offer withdraws once a connection wait is sealed.
Four tests fail on 235ec0f and pass here.
The delegate list, image card and delivery notice read pending-ness from `result === undefined`. A call sealed by a stopped turn or a lost completion event now carries `completedAt` and no result, so those parts kept rendering as running: a delegate row never left Running, image_generate showed a permanent Rendering image placeholder, and a delivery call showed a pending notice. Route that state to ToolFallback, which already renders it as Result unavailable.
Three component tests drive the real Thread render for each part; all three fail on the previous head.
Applying a reasoning/speed default in Desktop Settings -> Model reset an
auxiliary slot a user had pinned via CLI back to provider "auto" / model ""
while leaving reasoning_effort intact (#95460). POST /api/model/set was
never the writer; writeAgentDefault was: it round-tripped the whole
default-expanded config record (loaded when Settings opened) through
PUT /api/config, so every key another surface changed since the snapshot was
echoed back with its stale, default-filled value. The pinned slot's
provider/model existed only as defaults in the snapshot; reasoning_effort was
already in it, hence the asymmetry the report observed.
PUT /api/config deep-merges onto disk, so a writer only needs to send the key
it changed. Every desktop single-key writer now does exactly that (the
config-settings page already diffed against a baseline): Model defaults
(agent.reasoning_effort / service_tier), Appearance resume_last_session,
terminal font, session auto-archive, the two browser.use_real_profile
toggles, and the Capabilities voice fields (diffConfig against a baseline).
The optimistic shared-cache write keeps the full merged record so sibling
surfaces repaint without a refetch. The dashboard's ReasoningPicker had the
same read-modify-write shape and now sends the sparse patch too.
Tests pin the wire contract: only the edited key is sent, a sibling pin that
is not in the snapshot cannot be echoed back.
The Desktop AUX_TASKS registry drifted from web_server._AUX_TASK_SLOTS:
triage_specifier, kanban_decomposer, and profile_describer were served by
the backend (stale-aux warnings referenced them) but had no row in
Settings > Models > Auxiliary models, leaving "Reset all to main" as the
only way to change them. Add the three rows with localized labels/hints
(en/ar/ja/zh/zh-hant) mirroring the CLI _AUX_TASKS descriptions, and pin
the rendered set with test assertions.
Fixes#97297
- The Kanban plugin cannot import `@/lib/ime` (plugin fence), so its
`./ime-enter` twin duplicated the predicate. Export `isSubmitEnter` from
`@hermes/plugin-sdk` and drop the twin so one helper owns the policy.
- `BoardNameField` in kanban/board-switcher.tsx (main's refactor of the
create/rename board dialogs) submitted on composition Enter; guarded.
- Telegram allowed-ID input and the clarify-card textarea submitted on the
post-compositionend keyCode-229 Enter; both now use `isSubmitEnter`.
Widens the salvaged Kanban fix (PR #94611 by @huklaa) to the whole bug
class, following cloudflare-os PR #291 which centralized one
isImeComposing predicate and swept every Enter-submit site.
- New shared helper src/lib/ime.ts (isImeComposing / isSubmitEnter):
handles nativeEvent.isComposing (React), event.isComposing (DOM), and
the legacy Chromium/Safari keyCode 229 commit-Enter that arrives after
compositionend.
- Sweeps 21 previously unguarded Enter-submit sites: dialogs (project,
worktree, profile-remote-override, pet rename), settings fields
(credential keys, model API key, quick-entry shortcut, attachment
size, combobox), quick entry, pet overlay composer, pet generate +
hatch, review ship bar, file rename, preview browser address bar,
model catalog menu, MCP setup + approval strips, Kanban drawer.
- Upgrades two partial guards (onboarding, session-actions-menu) that
checked isComposing but missed keyCode 229.
- find-in-page: Enter step no longer fires mid-composition (typed CJK
search queries jumped the viewport on every conversion commit).
Validation: tsc clean, eslint clean, 135 tests green across ime/find-bar/
composer/kanban suites; sabotage run (guard removed) fails the new test.
Pasting exactly one http(s) link while composer text is selected now
turns the selection into a markdown link ([selected text](url)) instead
of replacing it — the behavior every rich text editor ships.
- resolveExactLinkPaste(): recognizes a clipboard payload that is exactly
one supported link (bare or <...>-wrapped, host required, no prose or
trailing punctuation) and returns the href.
- selectionLinkLabel(): the selected composer text eligible for linking;
rejects collapsed selections, selections spanning ref chips or line
breaks, and whitespace-only selections.
- markdownLinkFor(): builds the markdown link, escaping square brackets.
- handlePaste wires the three together ahead of the @url: chip path;
any non-qualifying paste falls through to existing behavior.
Adapted from Buzz's TipTap link-mark approach to Hermes' contenteditable
composer: we emit a markdown link (the composer's native rich construct)
rather than a ProseMirror mark.
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.
The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.
Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.
Closes#89259. Salvages #90649 by @higgs1729.
Settings → Model → Auxiliary gets a reasoning-effort selector per task next to the
provider/model pick (inherit / Off / level), sent as reasoning_effort on
POST /api/model/set and read back from GET /api/model/auxiliary.
(cherry picked from commit f09d008f10a81f57ed2426f835898c8e8ae595d7, resolved onto main; the backend half lives in
hermes_cli/web_server_config.py since the routers split and lands in the next commit)
Both stale-pin detections exempt only '' and 'auto':
- desktop persistentStaleAux banner (model-settings.tsx)
- switch-time stale_aux response (hermes_cli/web_server.py)
'main' is a backend-supported alias (auxiliary_client._normalize_aux_provider)
meaning "follow the active main provider", so aux slots pinned to it can
never be stale. The false positive fires for users following Moonshot's
official Hermes integration guide, which prescribes
auxiliary.vision.provider: main.
Exempt the alias in both places and add a regression test.
The contextSwitching early return in isRouteSessionMismatch sat above the
same-id short-circuit, so a profile or connection switch while the route
already pointed at the selected session reported a mismatch and blanked the
chat to the splash. On main that call returned false.
Move the selected-session check ahead of the contextSwitching guard: when the
selected view already owns the routed conversation there is no prior context
to leak, so nothing needs hiding. The guard still denies only the
transcript-retention fallback, which is the case it was added for.
Adds the exact regression to route-session-state.test.ts.
The node context menu now uses Radix, whose menu items are focusable
`div[role=menuitem]` elements. The window-level Space handler in
star-map.tsx only skipped INPUT/TEXTAREA/BUTTON/contentEditable, so
pressing Space on a focused menu item both activated the item and toggled
playback.
Extract the guard into `shouldIgnorePlaybackHotkey`, which additionally
bails when the event was already `defaultPrevented` or when the target or
active element sits inside a `[role=menu]`, and cover the menuitem case
with a small vitest.
The Star Map right-click menu was a hand-rolled `position: fixed` card
placed at the raw `clientX/clientY`, so a star within ~75px of the bottom
(or ~144px of the right) edge clipped the `Delete memory` / `Archive skill`
row off-window while `Edit …` stayed visible — the destructive action
silently disappeared.
Reuse the shared Radix `DropdownMenu` anchored to a zero-size fixed span at
the click point — the exact pattern `AppContextMenu` already uses — so the
menu gets the same flip/shift collision handling (and `collisionPadding`,
keyboard navigation, Escape/outside-click dismissal) as every other menu in
the app, instead of adding a second bespoke measure-and-clamp path.
`Edit …` keeps the menu open while the node content loads (`onSelect`
`preventDefault`) exactly as before; `openEdit` closes it on success.
Refs #109288. Supersedes the measure+clamp approach of #109301 (credit
@KoNit-K for the diagnosis). #100894 routes the gesture to this menu and is
untouched.