Hardening follow-up to #95396 (#95628): ignore primary connection-id writes
while the active key is a composite secondary scope, so future
presentation-layer writes cannot relabel the primary socket and poison
new-session routing. Regression test is sabotage-proven (fails with the
guard reverted).
The owner ladder's row rung (tile route -> hint -> session row) only searched
$sessions (recents). Cron- and messaging-sourced sessions are fetched as their
own sidebar slices ($cronSessions / $messagingSessions), so a scheduler-minted
cron session had no tile, no hint, and no row the rung could see. On a
registry-topology install every session-scoped RPC for such a session then
failed closed:
Session owner could not be resolved for "xxxx" (approval.respond): no owner
route, hint, connection-tagged row or profile probe named the backend ...
which made command approvals raised inside a cron chat impossible to answer
(the approval bar errors on every choice), even on a single-local-backend
install whose cron row - with its profile stamp - was already loaded for the
sidebar's cron section. The approval-bar path (requestForOwnedSession) has no
async REST-probe rung, so nothing downstream recovered.
Fix: ownerLookupSessionRows() returns the union of the three source-scoped
slices for OWNER lookups, keeping $sessions' array identity in the common
recents-only case so reference-keyed memo caches still hit. All four row-rung
call sites (session-rpc-dispatcher, knownOwnerForSession, tile delegate, tile
actions) now search it.
Repro: any cron session that raises approval.request (open the cron chat,
reply with a prompt that triggers a gated command, click Run).
createCanonicalChat gained { kickoff } on main (#95326) after #91227 was
filed with a bare positional openingStillCurrent; the salvage merges both
into one options object.
Close All only dismissed layout-tree panes. Bot tiles live in the
shared __bots_workspace__ bucket, so clicking a bot or swapping
profiles rehydrated the closed tabs. Persist-close those tiles
before dismissing the rest of the strip.
Rewrite the handoff proof to ≤150 lines on the hermes-bots VM seam so
base fails because the registered group closer is not invoked, not a
missing helper marker. Cover BotRow/Active Now local close-before-open,
remote no-dismiss, and no-group/old-host safety.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Exercise real BotRow and Active Now open handlers so close-before-open
and remote stay-put are load-bearing, not source-regex false greens.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Local BotRow and Active Now opens now share a handoff that retires the
selected group workspace before canonical chat open, so the bot chat can
own the main surface after leaving a group tab.
The #95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
Reproduces the reported Bot↔Default switch shape at the gateway.ts
activation seam: a switch-back that lands while the outgoing switch's
WS handshake is still pending keeps the route, the late-completing dial
neither steals the foreground nor breaks its socket, and re-activating
the bot works without an app restart. The guard (activation epochs +
open-socket-publish, landed via #89622/#92265/#81094) already prevents
the reported permanent break; this pins it so it cannot regress.
'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
- normalizeRegistry now preserves every malformed entry (unknown kind,
url-less remote/cloud, host-less ssh, mangled non-object items, and
any entry whose normalization throws) under a capped 'quarantined'
key that survives write cycles — healthy entries keep loading and
user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
surfaces reason+label summaries (never raw entries/token envelopes).
Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
These pre-existing tests fake a successful first connect by calling
only _ready.set(), which is what the real code did before this PR.
Now that run() gates the initial-vs-reconnect branch on the new sticky
_ever_connected flag instead, their later simulated reconnect failures
were misclassified as never-connected and hit the 3-attempt ladder,
failing test_reconnect_counter_resets_after_successful_session,
test_parked_server_self_probes_and_revives, and
test_retry_attempts_log_debug_transitions_warn in CI. Set the flag
alongside _ready.set() to mirror the real success sites, same as the
new test added in tools/mcp_tool.py's own PR.
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.
Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.
Fixes#94654
With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.
- Every registered gateway's profiles sit on the one strip, in registry order
(This device first, then by label), each group headed by that gateway's
kind glyph. The active gateway's squares are unchanged; the others are
"at rest" (dimmed) with tooltips/accessible names qualified by machine
(`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
the statusbar switcher, landing on that exact (gateway, profile):
`selectConnection(id, { profile })`. The spinner sits on the clicked
square; the previous source stays painted until the target answers.
Groups keep their slots whichever gateway is active, so a square never
moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
SOUL.md / Delete, executed on the owning gateway (renameProfile,
getProfileSoul and updateProfileSoul accept the same scope deleteProfile
already had); the delete confirmation names the machine. The legacy
per-profile "Connect to a remote host…" item is hidden on multi-gateway
setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
two registrations of one backend collapse to one group; past thirteen
squares across the fleet the strip condenses into a menu sectioned by
gateway. Roster is fetched on mount / focus / registry change only — no
periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.
Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as #91047). The restore now yields when
a switch is pending or already landed.
Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.
Docs: multi-connection-desktop.md describes the fleet rail.
Refs #89304, #92384, #91047, #94724
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fixes#89843. On a shared-remote connection every profile is served
through the primary socket, so waitForFocusedSessionHydration's
profileMatches gate could never become true — a bot chat whose stored
transcript painted within seconds still burned the whole 20s hydration
budget and then stranded the pane with 'Timed out loading <bot>'s
session history'.
The wake now resolves paint-first: once the stored transcript is painted
on exactly the target session, the content is its own proof — the pane
opens immediately and a subtle 'Syncing…' badge (new $hydrationSyncProfile
atom + ChatSyncBadge) shows until the profile gate catches up in the
background. Fail-closed everywhere content is not its own proof: a
superseded/conflicting concurrent wake still rejects, and an
expected-empty chat still waits for the full runtime gate.
- staleness guard in applyConfirmedSwitch: bail and dismiss if
current model/session no longer matches snapshot this warning was
created for, preventing stale Confirm from clobbering newer pick
(Enough1122 review #92492)
- neutral fallback for missing confirm_message: 'Confirm this model
switch?' instead of modelSwitchFailed
- document selectModel boolean: false means not-applied (pending
confirmation or failed) — pending already shows warning, not an error
- test mock: add dismissNotification mock for guarded flow
Addresses https://github.com/NousResearch/hermes-agent/pull/92492#issuecomment-5387151656
Desktop composer ignored gateway's confirm_required response for
contributor / expensive models. It painted the target optimistically
then invalidated model-options and refetched the still-active session,
so muse-spark-1.2-contributor appeared to instantly snap back to
gpt-5.6-sol with no explanation.
Now handle the confirmation protocol: on confirm_required rollback the
optimistic state, surface the backend's confirm_message as a warning
notification with a Confirm action, and on Confirm retry config.set with
confirm_expensive_model:true. Preserves data-training consent, keeps
the fix session-scoped and test-covered.
The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.
Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
Live tool-use A/B for session_search schema changes: arms are git refs
(tools/session_search_tool.py extracted per ref), tasks run a minimal
agent loop over OpenRouter against a freshly seeded temp session DB with
programmatic oracles — discovery, forced forward-scroll, AND-miss
broadening, verbatim link emission, profile-link resolution, browse.
Checked-in results/pr95570/ holds the 108-run battery (3 models, 3 reps,
2 arms) that validated the PR #95570 schema diet before merge:
base 49/54 vs diet 52/54, avg tokens/task -25%.
The old assertions pinned the phrasing that HID serve backends — the
exact asymmetry #81564 reports. Re-pinned to the new message and
strengthened: a serve-mode row must now appear, tagged [serve].
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.
Built on the spawn ledger (positive identity, never argv guessing):
- process_identity.py: LedgerEntry gains structured host/port/profile
(backward-compatible — readers .get()); register_self accepts detail=;
argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
manual backends inventory as supervisor=manual-serve with
restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
are stopped for the update and relaunched via an idempotent atexit
token built from structured identity (same contract as the gateway
pause/resume); receipts record serve_pause/serve_relaunch.
Desktop-owned backends keep the refusal (the app respawns what we
kill).
- dashboard_procs.py: the process scan is augmented with live ledger
rows, so profiled launches (`hermes --profile p serve ...`) that match
no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
— closing the #81564 status/stop asymmetry.
Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Slots below glm-5.3, above glm-5.2 in both curated lists; regenerates
model-catalog.json. No new metadata entries needed: context resolves via
the existing glm-5.3 fuzzy key (1,048,576 — matches OpenRouter live), and
both routes bill via official_models_api (live pricing).
Enough1122 review points on #94417:
1. Precedence hazard fixed: the busy-guard assertion now locates the
rebind helper body precisely and asserts the guard INSIDE it, instead
of a 2000-char window with an (m and X) or Y precedence trap.
2. stored_session_id guarantee: documented + pinned — the gateway always
stamps it ('stored_session_id': session_key or "" in server.py), and
the rebind's typeof check refuses non-string/empty values, so an
unnamed rebuilt runtime is never adopted as lineage proof.
3. New third assertion pins that refusal contract.
Structural smoke tests remain structural by design; the behavior
contract for the rebind is exercised end-to-end by the model-switch
manual repro path — a vitest harness driving handleSessionInfoEvent is
the follow-up candidate noted in the reply.
A mid-conversation model/provider switch rebuilds the agent runtime. The
rebuilt runtime emits session.info (and all later events) under a NEW
explicit session_id while the pane still holds the dead one as its
active id — isActiveEvent is false for the same conversation from that
moment on, so view-scoped updates stop and the chat freezes until a
full resume (#93942 scenario B; backend even logs 'client should resume
the stored session', but the client never does).
Fix: when a session.info event lineage-matches the selected conversation
(sessionMatchesStoredId over stored_session_id) but carries a different
runtime id, adopt the new runtime id as the active session id — keeping
the durable selection untouched — so every subsequent isActiveEvent gate
keeps matching without a resume. Guarded: the old runtime must show no
live turn (not busy/awaiting/streaming) or the adoption is refused, so
an overlapping manual switch can never split one conversation across
two panes.
The existing compression-rotation path does not cover this case: it
fires when the SAME runtime's stored id rotates, while a rebuild
produces a NEW runtime with a NEW stored id.
Together with #94255 (tile reconcile on sessions.changed), closes
#93942.
Regression tests verified failing pre-fix on 41447a6d70.
Enough1122 review points on #94255, all addressed:
1. Source-grep Python tests replaced with real vitest behavior tests:
- a tile whose stored transcript gained a background delivery IS
reconciled (updater invoked, correct stored id fetched)
- an unchanged transcript is skipped entirely (no updater call —
the signature gate proven, not asserted by regex)
The structural smoke test in test_bots_chat_live_append.py stays as
a cheap drift alarm; the contract now lives here.
2. Shared-sequence latest-wins semantics documented in the reconcile
docstring (review point 2).
3. Signature map pruning: a tile closed/superseded mid-read now deletes
its signature entry instead of leaking one map slot per ever-opened
tile for the app lifetime (review point 3).
4. Test harness: typed updateSessionState mock via Parameters<> instead
of the {}-as-state cast; no more silent spread corruption.
No production behavior change beyond pruning: 18/18 hook suite green,
typecheck clean.
CI caught two lint issues in the tile-reconcile path:
- no-restricted-syntax: don't mirror the $busy atom into a ref via
useEffect (stale-read hazard); reconcileTileTranscripts now receives a
live getter view so the loop reads the current value at tick time.
- react-hooks/exhaustive-deps: add updateSessionState to the
sessions.changed effect deps (stable useCallback from
useSessionStateCache, so no extra re-subscription).
No behavioral change: regression tests 2/2, adjacent hook suite 58/58,
typecheck clean.
Bot canonical chats open as workspace tiles (workspaceMode: 'bots') and
are deliberately hidden from $sessions/$messagingSessions, so the
sessions.changed transcript refresh skipped them twice over: it covers
only the main pane's selection, and its resolveSession() bails on hidden
sessions. A background delivery (bot-to-bot DM via bot_relay.deliver, a
cron run's output, another machine) therefore never reached an open bot
chat — the roster updated but the pane stayed stale until remount
(#93942 scenario A).
Fix: the sessions.changed tick now also reconciles every visible
workspace tile through a dedicated signature-gated path. Each tile
carries its own stored↔runtime id pair so no resolution step is needed;
per-tile signatures make no-change ticks free; busy tiles are skipped
(their own stream owns the view); closed/superseded tiles discard their
in-flight read.
Slice 1 of 2 for #93942 (scenario A only). Scenario B (stream re-key
after mid-conversation model switch) follows separately.
Fixes part of #93942