A shared dashboard's launch HERMES_HOME is not the selected profile. model.options now runs under @_profile_scoped, and global-remote REST keeps ?profile= even for the primary label.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Managed SSH maps a Desktop profile label onto a different remote name.
The sidebar filter lives in recents_profile, so rewriting only ?profile=
left those reads on the remote default and the Sessions list came back empty.
Co-authored-by: noah <loahnisk@gmail.com>
Post-boot WebSocket ticket mint failures and prolonged reconnects were
promoting into the full-screen "Hermes couldn't start" overlay, locking
users out of reading/drafting during brief 1–3 minute remote flaps.
- Ignore non-reauth boot-progress errors after a healthy cold boot
- Escalate prolonged transport reconnects with a non-blocking toast
- Soft-reset remote liveness rebuilds (no boot UI reset)
- Retry transient ws-ticket mints; auth rejections still fail fast
eslint --fix output: blank lines before statements and the import-order
spacing in connection-config.test.ts that the check:lint gate rejects.
Formatting only — no logic change.
The original implementation classified 502/503/504 only inside the readiness
loop, but for OAuth-backed Cloud connections the WebSocket-ticket mint runs
before waitForHermesReady. A server fault there was wrapped by
gatewayTicketFailure into a generic message and the Cloud-down classifier was
never reached. This closes that boundary and fixes a latent regex defect.
- isServerSideHttpError: structured-first (err.statusCode for 502/503/504),
legacy 'NNN:' prefix as fallback, non-Error inputs rejected. Also fixes the
committed '\d' (double-escaped, matched a literal backslash) that made the
function never detect a status prefix.
- makeNousCloudBackendDownError: single factory for the actionable Cloud-down
error (isCloudBackendDown/statusCode/detail/cause), shared by both the
ticket-mint boundary and readiness exhaustion.
- main.ts: run the Cloud classifier at mintGatewayWsTicket before the
gatewayTicketFailure wrap; 401/403 still route to reauth.
- connection-config.ts: gatewayTicketFailure preserves an integer statusCode
from the source error; auth semantics unchanged.
- boot-progress/IPC: carry isCloudBackendDown and statusCode through
DesktopBootProgress so the renderer overlay (a PR-body promise) can key on
the structured result rather than re-classifying the message string.
Tests: backend-health (structured detection, non-Error rejection, factory
shape/cause/guards, legacy fallback), connection-config (statusCode preserve,
401/403 reauth, integer-only copy), and an OAuth ticket-mint integration
regression (Cloud 503 -> actionable Cloud-down; 401 -> reauth). Connection-
config suite 80/80 green; backend-health sync tests green; the async readiness
loop tests cannot run on this host (pre-existing local-run limitation) and are
the CI gate. PR #85373 (#85335).
Three fixes for the "Install on this agent" pipeline, covering the whole
split-brain class between action-spawning endpoints and their status polls:
1. electron/connection-config.ts — the /api/actions/{name}/status poll family
now routes to the same backend as every action-spawning route. Before,
POST /api/skills/hub/install ran on the PRIMARY backend (scoped route)
while the follow-up status poll for a non-default profile routed to the
profile's POOLED backend, which never registered the dynamic action name
(skills-install-<slug>-<hash> lives only in the spawning process's
memory) -> 404 "Unknown action" toast even though the install succeeded.
POST /api/mcp/catalog/install joins the scoped table for the same reason.
2. src/store/hub-actions.ts — a non-zero subprocess exit now rejects with the
action log tail so the caller's catch toasts it. Before, a failed install
(scan gate, network, bad identifier) stopped silently: no toast, no row
flip, and the unchanged skills list read as "install did nothing".
3. src/contrib/runtime-loader.ts — a disk plugin copy shadowed by a bundled
twin now publishes a visible "(stale disk copy)" inventory row carrying
the folder path, instead of a console.info nobody sees. Stale
desktop-plugins/ leftovers from dev deploys are the same folders that
actively break the feature on shells without the bundled twin.
Clicking Mac Mini / Spark (the device default row) passed the desktop
pool key as the remote Hermes profile. That profile does not exist, so
the chat never opened. Named profiles (bob, dixie) already sent a real
name and worked.
Also refuse to fall back to this-device's default chat pin when the
remote source did not actually become active.
A user typed their root password into the Desktop SSH host field
(root@IP:PASSWORD form). Three failures compounded:
1. validateSshTarget() only checked for option injection (leading dash),
control chars, and port range — commas in an IP, whitespace ("ssh "
prefix pastes), and non-numeric ":<segment>" leftovers all dialed ssh
with garbage and failed silently five times.
2. normalizeSshConfig() only strips a ":<segment>" when it is numeric, so
a pasted password stayed glued to the hostname all the way into ssh
argv and the desktop.log connect line.
3. redactSecrets() had no pattern for ssh targets, so the password landed
verbatim in desktop.log and then in a PUBLIC debug-share paste.
Changes:
- validateSshTarget(): reject whitespace, commas, non-numeric colon
segments (with a "never put a password in the host field" hint that
does NOT echo the credential), and garbage hostnames; still accepts
bare IPv6 (::1, fe80::1%eth0). Reject whitespace/@ in user.
- redactSecrets(): new pattern masks any non-numeric segment where a
port belongs in user@host:... strings — defense in depth so future
parse gaps can't leak credentials into logs or debug shares.
- normalizeSshConfig(): strip a pasted leading "ssh " prefix.
- Tests for all three, including the exact incident shapes.
When Hermes Desktop works against a REGISTERED gateway connection, cron
jobs execute on that gateway and persist their run sessions in the
gateway's state.db. But every REST call in the app — the cron surface
included — carried only `profile`, so `hermes:api` routed it through the
local profile pool and `_list_cron_job_runs_sync` read a local state.db
with zero `source='cron'` rows. Every job showed "No runs yet" while the
same endpoint on the gateway returned the real runs (#87882).
Fix at the routing seam:
- HermesApiRequest gains an optional `connectionId`. The renderer's cron
helpers (list/get/runs/delivery-targets/create/update/pause/resume/
trigger/delete/blueprints) now tag the active registry connection via a
new connectionScoped() twin of profileScoped(), fed from the same
setApiRequestConnection seam store/gateway already maintains for the
plugin socket.
- The hermes:api main-process handler resolves a tagged request through
ensureRegistryBackend — the SAME pool the job list and WS traffic use —
instead of the legacy profile route. Shared remote/cloud hosts (one
gateway, many profiles) get the path scoped with ?profile= via the new
pathWithProfileScope helper, factored out of pathWithGlobalRemoteProfile.
- '' / 'local' / absent connectionId keep the byte-identical v1 route, so
single-source and connection-config-remote users are unaffected.
This covers the run-history panel, the sidebar cron peek, and every other
cron surface in one place, since they all funnel through the same helpers.
Fixes#87882
Fixes#73495. Two cold-start defects made the configured Hermes Cloud
agent vanish after a Desktop restart even though the persisted Portal
session was still renewable:
1. hasLivePortalSession() trusted the FIRST cookies.get() on the lazy
`persist:` partition. It now reuses the warmOauthCookieStore()
warm-up + bounded reread that hasLiveOauthSession() gained in
PR #67769, so a single hydration false-negative no longer clears the
agent list and flips the panel to signed-out.
2. Discovery required the short-lived `privy-token` access cookie but
treated its absence as a full interactive re-login, even when the
30-day `privy-session` / `privy-refresh-token` renewal cookies
survived the process exit. New cookiesHavePrivyAccessToken() splits
"signed in (renewable)" from "discovery can succeed right now";
discoverCloudAgents() and cloudAgentSilentSignIn() now mint a fresh
access token via one bounded, hidden, deadline-capped portal load
(renewPortalAccessSilently) before or after a 401, and only surface
needsCloudLogin when renewal genuinely cannot complete.
PRIVY_SESSION_COOKIE_VARIANTS also learns `privy-refresh-token` so a
renewal-only jar still counts as signed in rather than demanding an
interactive login while usable refresh material sits in the partition.
Tests: connection-config.test.ts covers the access/session split,
including the exact renewal-only cold-start jar from the issue repro.
Users pasting a Tailscale IP or LAN host as 'host:port' (no http://) hit
either a hard 'URL is not valid' error in the main process or, worse, a
silent dead probe in the renderer: the ^https?:// gates in the settings
and first-run forms never fired, so the field sat idle with no feedback.
- normalizeRemoteBaseUrl() (electron/connection-config.ts) now prepends
http:// when the input has no scheme:// prefix; explicit non-http
schemes (ws://, ftp://) still reach the protocol check and get a clear
rejection.
- New renderer twin coerceRemoteUrlScheme() (src/lib/remote-url.ts),
wired into both probe gates (gateway-settings.tsx and
first-run-remote-form.tsx) so the debounced /api/status probe, sign-in,
test, and save all see the coerced URL.
- Tests for both sides (electron/connection-config.test.ts,
src/lib/remote-url.test.ts).
Three helpers each re-derived part of the same decision: which backend
serves profile P, and does its REST path need a `?profile=` scope.
profileUsesPrimaryBackend answered the first half, pathWithGlobalRemoteProfile
answered the second, and ensureBackend re-checked globalRemoteActive() around
both. Splitting one table across three predicates is how the global-remote
case ended up registering reapable pool entries for a backend it never owned.
resolveProfileBackendRoute() states the four routes in one place and returns
the backend, the descriptor scope, and whether the path needs a query
parameter. The call sites read the answer instead of recomputing it.
One behavior change falls out: `hermes:api` now passes the primary profile
through, so the primary no longer sends itself a redundant `?profile=<self>`
on a global remote that already serves it.
Keep non-primary profiles that inherit the app-global remote on the primary connection descriptor instead of creating processless pool entries that the idle reaper repeatedly removes.
Preserve per-profile remote overrides and local pooled backends, and cover the routing policy with behavioral tests.
Co-authored-by: Rodrigo Fernandez <rodrigo@nxtlevelsaas.com>
The SSH modules predate the stricter lint config that landed on main (curly, no-empty, perfectionist sorting, prettier). Mechanical lint:fix + fmt pass, empty catch blocks filled with the codebase's void-0 convention, and inline no-control-regex disables on the three deliberate control-char patterns (same pattern as lib/ansi.ts).
Sync with main after the contribution-shell refactor (#60638) and the backendConnectionState extraction (#65885) landed. Seven conflicts, all resolved in favor of main's new architecture with the SSH lifecycle ported on top:
- electron/main.ts: adopt backendConnectionState (attempt tokens, attachProcess, clearPromiseForAttempt) as the sole owner of the primary backend; drop the branch's raw hermesProcess/connectionPromise globals and commitConnectionFailure call site. SSH liveness-streak classification, teardown, and the SSH quit path are layered onto the new ownership model; before-quit combines SSH transactional teardown with the Windows sandbox marker (#38216).
- desktop-controller.tsx: deleted on main; its SSH beforeConnectionSwitch cleanup (preserved-route fresh draft, overlay return-route reset, project-tree reset, close-all-terminals) moves to contrib/wiring.tsx's useGatewayBoot call.
- use-session-actions/index.ts: keep main's onFreshDraftRouteIntent (fires unconditionally); preserveRoute gates only the navigate.
- gateway-settings.tsx: keep main's acceptSavedConfig/connectedCloudUrl; every save/sign-in/sign-out success path routes through acceptSavedConfig with the branch's stale-async seq guards intact.
- boot-failure-reauth.ts: keep both sshFailureMessage and main's isRemoteReauthFailure formatting.
- Both test harnesses take the union of props.
Validation: tsc -b clean; desktop suite 1704 passed / 1 skipped; electron SSH/connection modules 225 passed; python SSH + web_server suites 508 passed.
Add SSH as a separate saved connection shape while preserving Cloud URL/OAuth semantics, inactive SSH drafts, strict host/port normalization, and profile-specific precedence.
Adds a third "Hermes Cloud" gateway mode to the desktop app: one portal
sign-in auto-discovers the agents on your account and connects to any of
them with no second interactive prompt.
- Electron: widen connection mode to 'local' | 'remote' | 'cloud', routed
through a centralized modeIsRemoteLike() so every resolution site treats
cloud exactly like remote; portal discovery (GET /api/agents over the
OAuth partition), Privy-cookie liveness, multi-org picker (NAS 409), and a
silent per-agent /oauth cascade (load protected root, not /login).
- Persist a cloudOrg on the cloud block; unselect cloud on mode switch.
- Renderer: Hermes Cloud ModeCard + agent picker (signed-out/loading/empty/
list), org picker, Change-org, connected-highlight + Connected pill.
- i18n (en + zh full; ja/zh-hant inherit via defineLocale), Cloud icon.
- IPC: hermes☁️{status,login,logout,discover,agent-sign-in}.
Salvage of #55402 onto current main: the original branch predates the
desktop electron .cjs -> .ts migration (39d09453f), so the electron half
was re-authored against the .ts files. Authorship preserved.
cloud-auto-discovery Phases 3 + 4.