Commit Graph

752 Commits

Author SHA1 Message Date
Teknium dfa74b5815 fix(desktop): boot session pop-out/watch windows against the session's owning profile
`openSessionInNewWindow` → IPC `hermes:window:openSession` →
`buildSessionWindowUrl` emitted no `profile`, so a secondary window (⇧⌘-click
pop-out, subagent watch) was a full renderer that adopted the PRIMARY
backend's profile and resolved the session id against the wrong store —
blank/wrong session for any non-primary profile (#82768, #61286).

The owning profile now rides the URL as `&profile=`, exactly the carry the
HUD already does (buildHudWindowUrl / windowProfileOverride in
use-gateway-boot); the renderer picks it with the same ladder openHud uses:
the session's stamped owner wins, an unstamped/uncached id (a brand-new
subagent child) inherits the profile the user is looking at.

Diagnosis credit: @DomGrieco (#82794).

Co-authored-by: DomGrieco <6556434+DomGrieco@users.noreply.github.com>
2026-09-02 06:47:55 -07:00
SZWzz ff8c1aca7a test(desktop): cover combined remote route precedence 2026-09-02 06:47:55 -07:00
SZWzz ae81786580 fix(desktop): let a per-profile remote override win over the forced-local route (#90477)
resolveRegistryLocalRoute collapsed globalRemote and profileRemoteOverride
into one forced-local branch. The two cases are different:

- globalRemote: forcing "This device" to spawn genuinely-local children is
  the intended migration behavior — unchanged.
- profileRemoteOverride: the per-profile SSH/remote override is an explicit,
  authoritative routing decision for that profile. Forcing local made the
  roster enumerate the profile via its override but open the thread in a
  forced-local child, which dies with 'Profile "x" no longer exists' when
  the profile only exists on the remote — reproduced on a macOS Desktop in
  global SSH mode where mythony-agent/q-agent exist only on the NAS.

The registry 'local' entry now delegates to the legacy profile route when a
per-profile override is present, so the override stays authoritative.

Tests: the override case now pins delegation, and a new witness pins that
globalRemote alone still forces local; both contracts are asserted together.
88 connection-registry + 91 remote-lifecycle + 73 routing tests pass;
tsc --build clean.
2026-09-02 06:47:55 -07:00
Brooklyn Nicholson 5fe9e209dc fix(desktop): offer a restart when the updated app is already on disk
When the bundle was swapped under a running process, the About banner sent the
user to the installer — a download and a reinstall for a state that a plain
restart repairs, and the reason reinstalling never helped these reports.

Report bundleSwapPending on hermes:version and give that case its own copy and
a "Restart Hermes" button. It gets its own headline too: reusing "App build out
of date" over a body that says the app is already installed repeats the
contradiction with the Updates card that the banner is supposed to resolve. The
installer link stays for the genuinely-stale-bundle case.

Packaged builds only — a dev `--build-only` rewrites the stamp under a running
`npm start`, and that is a rebuild the developer asked for, not a torn install.

Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
2026-09-01 22:33:29 -05:00
Brooklyn Nicholson c19537fb03 fix(desktop): relaunch into the swapped bundle instead of booting a torn renderer
A user who reopens Hermes while an update is running lands on the boot gate,
which is what it is for. But the updater swaps the packaged bundle on disk
after `hermes update` exits, and its `open` leg only focuses this already-
running process, so nothing ever loads the new build. The parked instance then
passes the gate and boots the new runtime under the old renderer — the "App
build out of date" banner immediately after a fully successful update, over an
Updates card that says "You're on the latest version" and so offers no remedy.

Compare the install stamp this process loaded at boot with the one on disk when
the gate clears. On positive proof of a swap — different commit, or a different
builtAt at the same commit — relaunch instead of starting a backend. Detection
fails quiet like bundle-skew, so a swap that never happened (the Windows
locked-binary case) is unchanged. A one-shot argv flag makes a relaunch loop
impossible and a 15s failsafe falls back to the old behavior.

Co-authored-by: tk-pkm111 <133480534+tk-pkm111@users.noreply.github.com>
Co-authored-by: aeonsong <aeonsong@users.noreply.github.com>
2026-09-01 22:33:29 -05:00
Brooklyn Nicholson 94e21c495e fix(desktop): don't warn about a bundle that isn't behind
detectBundleSkew() trusted `git rev-list --count <stamp>..HEAD -- apps/desktop`
outright, which claims skew in two states where the install is not torn.

Ancestry: `A..HEAD` only measures how far HEAD is ahead of A when A is an
ancestor of HEAD. A ZIP-fallback update rewrites the tree onto a synthetic
root, so the stamp still resolves but is unreachable; the range degenerates to
HEAD's own history and reports a permanent >= 1 while apps/desktop is
byte-identical. Ask `merge-base --is-ancestor` first and go quiet unless it
answers yes.

Scope: the pathspec counted every file under apps/desktop/, so a docs- or
e2e-only commit produced a banner promising missing UI features that do not
exist. Count only the paths that reach the shipped app.

Co-authored-by: jackulau <jackulau@users.noreply.github.com>
Co-authored-by: kokhlo <kokhlo@users.noreply.github.com>
2026-09-01 22:33:29 -05:00
Jeffrey Quesnelle c56f8cdd48 Merge pull request #100667 from NousResearch/feat/local-models-squash
feat: local models — managed llama.cpp runtime with one-click desktop  setup
2026-09-01 17:53:28 -04:00
hermes-seaeye[bot] 3ca096de5f fmt(js): npm run fix on merge (#100685)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 20:48:06 +00:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Sergey Tiraspolsky f6c9cb7b90 fix(desktop): re-score heuristic active-profile.json pins
_migrated:true files were skipped on later boots, so #100576 installs
stayed stuck on the named profile. Re-evaluate those files only; leave
user-selected pins (no _migrated) alone. If default now wins, write
{profile:null} instead of pinning default.
2026-09-01 15:24:33 -04:00
Sergey Tiraspolsky c90f8e8cf6 fix(desktop): score default ~/.hermes/state.db in first-boot profile migration
First-boot migrateActiveProfileIfMissing only listed ~/.hermes/profiles/*
and scored profiles/<name>/state.db. Default's real DB is ~/.hermes/state.db,
so a tiny named profile could be pinned after an update.

Always candidate default, score/pid-check it at HERMES_HOME, and do not
write active-profile.json when the winner is default.

Fixes #100576
2026-09-01 15:24:33 -04:00
hermes-seaeye[bot] e5e71d8c46 fmt(js): npm run fix on merge (#100517)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 16:59:53 +00:00
David Metcalfe 75f8ab2f87 style(desktop): satisfy eslint and prettier in profile migration
- curly: brace all single-line if statements in profile-migration.ts and
  profile-migration.test.ts (17 errors in CI check:lint)
- perfectionist/sort-imports: node:fs builtin import before vitest external
- padding-line-between-statements: blank lines after block statements
- prettier: normalize formatting (fmt script style) in the three touched files

All 967 electron project tests still pass.
2026-09-01 09:54:22 -07:00
David Metcalfe b6a0360f81 fix(desktop): clarify which profile reads the migration precedes
Polish from antigravity review of the rebase resolution (GPT-OSS):
the previous comment said "BEFORE the first primaryProfileKey() /
primaryBackendIsRemote() read" but those two calls live at different
points — primaryBackendIsRemote() is the very next line, primaryProfileKey()
is inside the connection IIFE. Be explicit about which is where so a
future reader who moves one of them knows what to preserve.
2026-09-01 09:54:22 -07:00
David Metcalfe 285ced674f test(desktop): cover active-profile migration helpers
Addresses teknium1's review (#64195) finding #2: the multi-rung resolver
needs Electron tests covering precedence, stale-PID rejection, fallback
behavior, and the remote boot path. The pure decision helpers are now
covered by 29 unit tests in `profile-migration.test.ts` (vitest electron
project).

Coverage:
- precedence: legacy > single-running-gateway > state.db heuristic
- stale-PID rejection: recycled PIDs not owned by hermes are dropped
- malformed pid files: JSON parse errors, non-integer PIDs, zero/negative
- scoring edge cases: ancient files (recency floored at 0.1), tiny files
  (size floored at MIN_SIZE), larger DB beats smaller at similar recency
- single-profile fallback: best === 'default' suppresses the write
- no-op cases: preference file already exists, missing profiles root

The remote boot path is verified by code review of the call-site move
(commit preceding this one) — `migrateActiveProfileIfMissing()` now runs
before `primaryProfileKey()` is first read in `startHermes()`.

The pure decision logic that the orchestrator relies on is covered end-
to-end below; this matches the repo's testable-helper pattern (see
`profile-delete-routing.test.ts`).
2026-09-01 09:54:22 -07:00
David Metcalfe 5760c65be2 fix(desktop): migrate active profile before first primaryProfileKey() read
Addresses teknium1's review (#64195) finding #1: the previous PR placed
the migration inside the connection IIFE, AFTER
`resolveRemoteBackend(primaryProfileKey())`. When the preference file
was missing, `primaryProfileKey()` resolved to 'default' and the remote
branch returned immediately without ever reaching the migration. Remote-
mode users got no migration at all.

Move the call site to the top of `startHermes()`, before the connection
IIFE that reads `primaryProfileKey()`. Both remote and local branches now
flow through this path before any profile-dependent resolution, so the
migration runs on first boot regardless of mode.

The inlined implementation is replaced with a thin wrapper that builds a
`MigrationDeps` bag and delegates to `migrateActiveProfileIfMissing` from
`profile-migration.ts`. No production behavior change beyond the call-
site move.

Tests added in a separate commit.
2026-09-01 09:54:22 -07:00
David Metcalfe a0511fd884 refactor(desktop): extract pure active-profile migration helpers
Addresses teknium1's review (#64195) — the migration decision logic should
be unit-testable without Electron. Pull the ladder (legacy sticky file,
running-gateway scan, state.db heuristic) into pure helpers in a new
`profile-migration.ts` module, following the dep-injection pattern
already established by `profile-delete-routing.ts`.

The helpers take an injected `MigrationDeps` bag so tests can exercise
precedence, stale-PID rejection, fallback behavior, and the single-profile
case without touching `/proc`, `ps`, or the host filesystem. The default
profile is explicitly rejected from the legacy rung because the regex
matches `default` and accepting it would suppress the heuristic that is
the whole point of the migration.

The wrapper in main.ts is unchanged in behavior — the same `MigrationDeps`
fields get filled in from `fs`/`path` and `isHermesProcess`. The atomic
write + parent-dir-create that the wrapper performs matches
`writeActiveDesktopProfile`'s semantics so the migration produces a file
indistinguishable from a user-driven profile switch.

No production behavior change; pure code organization.

Tests added in a separate commit.
2026-09-01 09:54:22 -07:00
David Metcalfe 4980a5ceae fix(desktop): migrate active profile preference from legacy signals on first boot
When active-profile.json does not exist (fresh install or first boot after
update), seed it from the best available signal so the Desktop launches
into the user's primary profile instead of always defaulting to "default".

Priority ladder:
1. Legacy ~/.hermes/active_profile (explicit CLI choice via hermes profile use)
2. Running gateway (gateway.pid with verified liveness + hermes identity check
   via /proc/cmdline or ps -o args= to avoid PID recycling false positives)
3. state.db heuristics — hybrid recency×size score picks the primary workspace
   (e.g. a 409MB coder DB beats a 28MB default DB even if touched at similar times)

The stored JSON includes _migrated:true for priority 3 (heuristic guess) so
the renderer can optionally surface a one-time notification. Priority 1 and 2
are higher-confidence signals and skip the flag.

The migration is a no-op once active-profile.json exists, and only writes
when a non-default profile is confidently identified — preserving the legacy
fallback-to-default behavior for single-profile users.

Fixes #64160 (active-profile half).
2026-09-01 09:54:22 -07:00
Teknium da96a9c204 fix(desktop): never lose HERMES_BACKEND_READY emitted before the port wait attaches
The spawn-time output tail (#93608) puts child stdout into flowing mode at
spawn. Both backend spawn paths in main.ts then await claimBackendChild
(whose Windows Get-Process probe cold-starts in 2-8s) and boot-progress IPC
BEFORE waitForDashboardPortAnnouncement attaches its stdout listener. Node
streams never replay consumed chunks to late listeners, so a READY sentinel
printed during that window was lost forever — the wait hit its 90s timeout
and a healthy backend was killed (deterministic on Windows, racy on
macOS/Linux; still firing on v0.21.0 incl. concurrent multi-profile boots).

Fix (belt and suspenders, both spawn paths — primary and profile pool):
- create the port-announcement promise immediately after spawn, before any
  await
- new bufferedOutput option on waitForDashboardPortAnnouncement: after
  attaching its own listener, waitForDashboardPort scans the output tail's
  already-buffered text for the sentinel, making listener-attach ordering
  irrelevant regardless of call-site shape

The readyFile path was already ordering-safe (it polls a file, not the
stream). Approach follows stale PR #60986 by @ParaWheeler, rebased onto the
output-tail/readyFile plumbing added since.

Fixes #60323
2026-09-01 08:30:21 -07:00
Teknium a1c25d393a feat(desktop): built-in optional-skills catalog in Capabilities → Skills with one-click install
The Skills tab now lists the entire official optional-skills catalog
(optional-skills/ shipped with the repo) below the installed skills.
Each catalog row has an Install button that routes through the standard
hub action pipeline; once the install finishes the row flips into the
installed list with the normal enabled/disabled toggle.

- backend: GET /api/skills/hub/official — OptionalSkillSource.list_local()
  scan (no network) + per-profile installed flags from the hub lock
- desktop: catalog section in SkillsView with search/scope integration,
  install-state spinners off $hubActions, and an OfficialSkillDetail pane
  (hub preview: frontmatter + full SKILL.md + Install)
- CapRow gains an optional action slot (button instead of the Switch)
- electron: route the new endpoint with the skills family (primary backend)
- i18n: officialCatalog/officialPill keys across en/ja/zh/zh-hant
2026-09-01 01:39:59 -07:00
hermes-seaeye[bot] 9366ba3afb fmt(js): npm run fix on merge (#100128)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 08:34:56 +00:00
Sergey0515 f85edd3d94 fix(desktop): stop plugin-launched gateways before the Windows release gate (#70337)
releaseBackendLock tree-kills only the Desktop's own backends
(backendConnectionState + backendPool). A messaging gateway launched by
the gateway-launcher desktop plugin via /api/gateway/start lives outside
those structures; on Windows its launcher (venv\Scripts\python.exe)
keeps the venv mandatory-locked, so the 15s release gate aborts the
hand-off BEFORE the venv-blocker scan — the scanner's pausable-gateway
exemption and the CLI updater's pause machinery never get their chance.

Delegate to `hermes gateway stop --all` (launcher + worker discovery
across every profile; gateway.pid records only the uv WORKER and
taskkill /T from the worker never reaches its parent). Per review, adds
the drain-semantics counterpart: every applyUpdates abort path
(lock-held, venv-blocked, probe-failed, updater-spawn-failed) restores
gateways via `gateway start --all`, so a failed update no longer
strands every profile's gateway stopped. Pure DI'd module + tests.

Salvaged from PR #76057 (issue #70337; overlap credit: #70477 by
@JonthanaHanh targeted the same symptom earlier via ZIP-dir preservation).
2026-09-01 01:30:09 -07:00
xy952666680 6b18224df5 fix(desktop): reap detached hindsight venv daemon before update handoff (Windows)
The pre-handoff teardown tree-kills only the backends the Desktop owns
(backendConnectionState + backendPool). The memory plugin's hindsight
daemon is spawned DETACHED off venv\Scripts\pythonw.exe, so it survives
the teardown, keeps venv files mapped, and either dead-ends the
venv-blocker scan with no in-app remedy or (pre-#74805 shim-only gate)
raced the updater into a half-updated venv.

Add a narrowly-scoped reap: kill only processes whose exe lives under
venv\Scripts (ordinal case-insensitive prefix — no PowerShell -like
wildcard hazards) AND whose cmdline references hindsight_api.main.
External holders (user terminals, unrelated scripts) are never killed —
scanVenvBlockers still reports them and the hand-off aborts, per existing
design. Selection logic is a pure DI'd module with unit tests.

Salvaged from PR #75477 (scoped per review: the narrow daemon kill; the
PR's generic every-exe kill was rejected as over-broad, its updateInFlight
half was superseded by #75778/#73822, and its generic lock-probe half by
the #74805 release gate + #99724 scanner classification).
2026-09-01 01:30:09 -07:00
liuhao1024 46ac31e84c fix(desktop): sanitized deferral evidence for ledgered manual serve blockers
The Desktop venv-blocker scan (since #99724) defers ledger-verified
serve/dashboard holders to the CLI updater's stop+relaunch rungs, but the
scan output only carried an opaque deferred_backends count — nothing
explained WHICH holders the deferral consumed or why they vanished from
processes.

Add sanitized decision evidence (#98350): deferred_backend_evidence lists
structured ledger identity only (pid, purpose, recorded port) — never the
command line, which can carry tokens or private endpoints. Adds a desktop
parser contract fixture proving the consumer tolerates the diagnostics
while keeping blocked/processes authoritative.

Salvaged from PR #98350; the exemption half of that PR was independently
consolidated on main via #99724 (_is_updater_owned_backend).
2026-09-01 01:30:09 -07:00
hermes-seaeye[bot] 04224b2f82 fmt(js): npm run fix on merge (#100096)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 07:41:45 +00:00
Teknium 383f827493 fix(desktop): satisfy import-sort and ref-mirror lint rules in comment mode
Auto-fix perfectionist import/export ordering across the new
preview-annotate files, and extract the conversation-switch annotate
reset into a useCallback so the effect body carries no .current writes
(no-restricted-syntax ref-mirror rule).
2026-09-01 00:30:56 -07:00
Adolanium 10f2a20966 feat(desktop): add comment mode to the in-app browser
Click or drag on the preview page to pin a numbered comment. Saving a pin only adds it to the stack. Add comments drops the crops and a short prompt into the composer and never sends the turn.

Capture takes the visible page then crops in bitmap space, because Electron's rect crop on guest webviews is empty on Windows and shifted on high-DPI.
2026-09-01 00:30:56 -07:00
xxxigm bcecd675f7 fix(desktop): unwrap Models-page code-skew 503 and recycle the owned backend (#97046)
Show a Restart backend action that kills the SSH serve before the local child
so reconnect cannot reuse a stale lockfile, instead of dumping raw IPC JSON.
2026-08-31 20:43:36 -05:00
Brooklyn Nicholson 0e7eebc266 fix(desktop): scope remote model catalog and primary-label REST to the focused profile
A shared dashboard's launch HERMES_HOME is not the selected profile. model.options now runs under @_profile_scoped, and global-remote REST keeps ?profile= even for the primary label.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-08-31 20:40:09 -05:00
Gille f98f5e74e0 fix(desktop): preserve terminal startup output 2026-08-31 19:43:31 -05:00
hermes-seaeye[bot] 66b844c967 fmt(js): npm run fix on merge (#99871)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 00:41:26 +00:00
Brooklyn Nicholson a2907a8bcd fix(desktop): skip cold process probes for dead backend-ownership PIDs
Windows Get-Process and macOS ps exit 1 on a missing PID, so reapOrphans
kept stale records and the next launch paid another 2-8s spawn each.
Throw ESRCH from the existing isPidAlive helper before any shell-out.

Closes #92875

Co-authored-by: Jackal991 <139240222+Jackal991@users.noreply.github.com>
Co-authored-by: jonotonfoto <126111813+jonotonfoto@users.noreply.github.com>
Co-authored-by: foras910521-lab <268267187+foras910521-lab@users.noreply.github.com>
2026-08-31 19:32:14 -05:00
hermes-seaeye[bot] e22f8a7fbd fmt(js): npm run fix on merge (#99862)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-01 00:28:40 +00:00
686f6c61 37fc92a3d6 fix(desktop): keep SSH serve teardown across re-entrant quit
Window X calls app.quit(); backendShutdown.finally() calls it again.
teardownSshConnection deletes the map entry before SSH exec kill, so
the second before-quit saw an empty map and exited while disconnect
was still running. Latch the in-flight teardown so the re-entrant quit
still waits.
2026-08-31 13:59:51 -07:00
Teknium 2e9a39d28e fix(desktop): heal v1 SSH gateway routes into the v2 connections registry
reconcileRegistryDrift only healed remote/cloud v1 routes. A v1 global
mode:'ssh' route (host, no url) written by Settings after the one-shot
migration had no registry identity: resolvedConnectionId returned null,
primary stayed 'local', and every launch re-homed the window onto a
fresh local backend. Because the heal skipped SSH entirely, the two
config files re-drifted after every update relaunch instead of
converging once.

Normalize the v1 SSH descriptor into a v2 kind:'ssh' entry (via the
same validated normalizeConnectionInput path the editor uses) and align
primary/lastUsed, with the same narrow-heal rules as remote: already-
registered targets and deliberate primary picks are left alone, and
unusable hosts never touch the registry.

Diagnosis credit: mgallmur-glitch (root cause) and jakewvincent
(re-drift after update relaunch) on #93888.
2026-08-31 12:17:31 -07:00
Teknium 6e41ab3460 fix(desktop): bounded auto-restart for no-mux SSH tunnel flaps instead of instant connection death (#96266)
A no-mux tunnel is a single persistent `ssh -N -L` child. On main, ANY
death of that child after readiness immediately set tunnel.alive=false,
which poisons SshConnection.isAlive() forever; upstream lifecycle probes
then treat the whole SSH connection as dead, tear down the scope, and
SIGTERM a perfectly healthy backend (~10s after HERMES_BACKEND_READY in
the #96266 logs: '[ssh] connection closed (no-mux tunnels killed)' ->
'Ignoring stale Hermes backend exit (SIGTERM)' -> 90s port-announcement
timeout, with retry/repair looping the same failure).

Now a post-readiness child death is a tunnel FLAP: the child is
restarted with a bounded budget (5 attempts, 1s delay by default,
injectable for tests) and only an exhausted budget marks the tunnel —
and thus the connection — unhealthy. Deliberate teardown (cancelForward
/ close) sets tunnel.stopping, cancels any pending restart timer, and
never restarts. Pre-readiness deaths keep failing fast with classified
stderr (auth/bind errors unchanged).

Fixes the kill chain of #96266.
2026-08-31 12:16:48 -07:00
andyst-dev 4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00
hermes-seaeye[bot] 446563262e fmt(js): npm run fix on merge (#99323)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-31 10:15:30 +00:00
dyreckt 2f261f99ff fix(desktop): keep profile scope when reusing primary remote backend
ensureRegistryBackend() reuses the ambient primary descriptor for a
non-local, non-ssh registry primary, but the reused descriptor was
missing sharedRemote: true. The request router only appends
?profile=<profile> for sharedRemote backends, so Capabilities/Skills
requests went out unscoped and the gateway served the default profile
while a named profile was selected.

Add the flag and a regression test asserting the scoped URL.
2026-08-31 03:09:53 -07:00
hermes-seaeye[bot] bdb8b1603f fmt(js): npm run fix on merge (#98548)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-30 11:59:33 +00:00
Teknium 3528a3bfc4 fix(desktop): satisfy perfectionist/sort-imports for mcp-oauth-callback-ipc import 2026-08-29 19:18:01 -07:00
Teknium 0f5dd5c46e feat(mcp-oauth): Desktop MCP OAuth now completes against remote backends (client-side callback relay)
The gateway's session-backed MCP OAuth flow (mcp.servers.oauth.start) binds
its browser-callback listener on the BACKEND machine's 127.0.0.1. When the
Desktop app connects to a remote backend (SSH/Tailscale), the user's browser
resolves that loopback to the user's machine, the redirect dies, and every
OAuth catalog server (ClickUp, Hospitable, ...) fails in-app with no working
path — the exact topology from the 'MCP Recurring erros' support thread.

Fix mirrors the Desktop's native gateway login (native-oauth-login.ts):

- gateway: mcp.servers.oauth.start accepts client_redirect_uri (loopback-only,
  RFC 8252-style validation); when supplied no gateway listener is bound and
  the OAuth redirect_uri pins to the client's listener.
- gateway: new mcp.servers.oauth.callback RPC relays the client-captured
  code/state into the flow; state verification stays in
  DashboardOAuthFlow.deliver_callback (constant-time compare, replay-safe).
- desktop: mcp-oauth-callback-ipc.ts hosts a one-shot 127.0.0.1 listener in
  the main process (hermes:mcp-oauth:listen/wait/cancel via preload bridge).
- desktop: hermes-bots mcp-setup.tsx prefers the client listener for local
  AND remote backends, falling back to the legacy gateway-listener flow on
  older gateways (feature-detect via start rejection).
- docs: remote-host MCP OAuth section documents the automatic Desktop path.

Validation: 19 new gateway tests (validator allowlist, listener skip, relay
accept/reject/replay) — sabotage-verified; 5 new desktop tests against a real
ephemeral listener; E2E through the real session registry + flow bridge with
a stubbed provider probe; tsc electron+renderer builds clean.
2026-08-29 19:18:01 -07:00
Amandeep Khurana a7e7de6407 fix(desktop): make gateway file saves failure-atomic so a failed download never destroys an existing file
`pumpStreamToFile` opened the user-chosen destination with
`fs.createWriteStream`, which truncates the target the instant it opens,
and its error path then unlinked that same path. When a user picked an
existing file in the Save dialog (and confirmed the overwrite) and the
gateway dropped mid-stream, the original was gone: truncated first,
deleted second, with nothing written in its place. The data-URL
compatibility fallback (`saveGatewayFileViaDataUrl`) had the same class
of bug via `fs.promises.writeFile`, which truncates before the write
completes.

Both paths now go through one failure-atomic primitive. Bytes land in a
short, randomly named sibling temp file (`.hermes-download-<hex>.part`,
same directory so the final step is a same-volume rename), created with
`flags: 'wx'`, and are renamed onto the destination only after the whole
body has been written and the descriptor released. The destination is
never opened before that point, so a failed download leaves whatever was
there untouched.

- Ownership-gated cleanup: the temp file is unlinked only after the
  stream's 'open' event proved THIS operation created it. An exclusive
  create that fails before open (EEXIST collision, EACCES, missing
  parent) never removes a file that belongs to someone else.
- `WriteStream.close(cb)` rather than `end(cb)` before renaming: `end`'s
  callback fires on 'finish' while the fd may still be open, and Windows
  refuses to rename a file with an open handle. Falls back to `end` for
  stream shapes without `close`.
- The failure path waits for 'close' (bounded by a 2s grace period)
  before unlinking, for the same reason: `destroy()` releases the fd
  asynchronously and an unlink racing the open handle would leak the
  `.part` file on Windows.
- A rename failure (destination locked, permissions) removes the owned
  temp file and rejects; nothing is left behind.
- Fixed-length temp name so a long user-chosen filename cannot push it
  past the filesystem's name limit.
- `fsPumpDeps()` is the single production deps factory (`'wx'` create,
  `fs.promises.rename`, `fs.promises.unlink`); `writeBufferToFile()`
  routes the data-URL fallback through the same pump. `PumpDeps` gains
  `rename` and a `tempPathFor` test seam.

Tests. Fakes: temp-then-rename on success, close-before-rename ordering,
the regression itself (destination neither opened nor unlinked when the
response fails mid-stream), write-error cleanup, close-before-unlink
ordering, rename-failure cleanup, pre-open EEXIST leaves the colliding
file alone, `writeBufferToFile` success and post-open write failure, the
temp-name length bound, and the `main.ts` wiring. Real filesystem
(`gateway-file-download.fs.test.ts`, exact production deps in a scratch
dir): completed download replaces the destination with no temp left;
mid-stream failure leaves the pre-existing destination byte-for-byte
with no `.part`; failure into a fresh name leaves nothing; seeded temp
path survives a pre-open EEXIST with no rename; rename failure (directory
at the destination) cleans the owned temp; data-URL fallback success and
missing-directory failure.

Adds the contributor email mapping required by the attribution check.

Fixes #96597

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015u8q2pHVPZmxpSrgkt94jC
2026-08-29 11:33:05 +05:30
Brooklyn Nicholson 240790af60 fix(desktop): let HUD prompts take clicks on solid X11
Ignore-mouse cannot restore on X11, so a visible band that still has
pointer-events:none swallows clarify options and links. Held prompts
and solid-window bands now take the pointer without composer focus.
2026-08-29 00:51:43 -05:00
Brooklyn Nicholson e60983a697 fix(desktop): let the HUD drag onto another monitor
Composer drag added renderer CSS-pixel deltas onto a window AppKit
clamps to the current display, so the bar could not follow the cursor
onto a second monitor (and drifted on mixed-DPI Windows). Track the OS
cursor in main and lift that clamp.

Co-authored-by: Biotrioo <biotrioo@protonmail.com>
2026-08-28 15:33:19 -05:00
Casey beb212dcc5 desktop: fix two managed-SSH-spawn bugs that break every fresh remote backend
1. Quoting: the spawn payload wrapped expandRemotePath() output -- already
   a shell-quoted fragment like "$HOME"'/...' -- in shq() again, so the
   reservation/lock/owner_file variables hold the quote characters
   literally and every mkdir "$reservation" fails forever (~5 min per
   attempt spinning in the reservation loop while holding the box-global
   update mutex; queued spawns starve behind it). The same double quoting
   sits in the stale-reaper identity guards, making every reap REFUSE.
   The lockfile-reuse path masks the bug for existing backends, so it
   only bites on fresh spawns.
2. Bashism: lockfile publication used ${var//__PID__/$child} -- bash-only
   substitution in a payload run under plain sh (dash on Ubuntu), which
   aborts the script AFTER the serve was spawned. The client then saw an
   unknown failure, ran its error cleanup (deleting the token file), and
   the just-booted serve died on the missing token -- orphaning one serve
   per attempt. Replaced with a POSIX sed substitution.

Adds two regression tests: payload variables must keep $HOME expandable
(no re-quoting), and the pid substitution must be POSIX sh. Both fail
against the previous code; all 89 remote-lifecycle tests pass with the
fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 02:48:21 -07:00
kshitijk4poor 9f05b06589 Revert "Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay"
This reverts commit df7d7f6e8d, reversing
changes made to 1a66134404.
2026-08-27 11:26:57 +05:30
kshitij df7d7f6e8d Merge pull request #94245 from kshitijk4poor/feat/gw-event-replay
feat(gateway): slim WS-only server — remove FastAPI/uvicorn from desktop boot path
2026-08-27 11:22:41 +05:30
hermes-seaeye[bot] 36b0a96dcb fmt(js): npm run fix on merge (#96076)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-27 04:43:58 +00:00
Finn763 00988f3b89 fix(desktop): widen pool keepalive-fresh window to absorb WSL2 IPC stalls (#95189)
The renderer pings each pool backend every 60s (`hermes:backend:touch` →
`touchPoolBackend` → updates `lastActiveAt`). The LRU eviction cap used a
keepalive-fresh window of 90s — only 1.5× the ping cadence — to decide
whether a backend was "plausibly still alive". One missed or delayed ping
pushed a live backend past the threshold and the cap-driven eviction killed
the active profile's backend mid-session, restarting the gateway and
re-minting runtime ids. On WSL2, where the renderer→Electron IPC roundtrips
through 9p, brief 9p hiccups commonly stretch a ping to seconds of observed
silence, producing the ~80–90s exit / ~2 min cycle reported in #95189
(122 gateway starts on 2026-08-26 alone, driving renderer OOM via reconnect
churn at ~5GB/day).

Widen POOL_KEEPALIVE_FRESH_MS to 4 minutes (3× ping cadence + IPC stall
headroom, still bounded well below POOL_IDLE_MS=10min). Backends with one or
even two missed pings are now spared; truly idle backends (multiple lapses,
minutes idle) remain eligible for eviction by the cap and the idle reaper.
The constant is also overridable via HERMES_DESKTOP_POOL_KEEPALIVE_FRESH_MS
to make this tunable without a rebuild.
2026-08-26 21:38:20 -07:00