1. Quoting: the spawn payload wrapped expandRemotePath() output -- already
a shell-quoted fragment like "$HOME"'/...' -- in shq() again, so the
reservation/lock/owner_file variables hold the quote characters
literally and every mkdir "$reservation" fails forever (~5 min per
attempt spinning in the reservation loop while holding the box-global
update mutex; queued spawns starve behind it). The same double quoting
sits in the stale-reaper identity guards, making every reap REFUSE.
The lockfile-reuse path masks the bug for existing backends, so it
only bites on fresh spawns.
2. Bashism: lockfile publication used ${var//__PID__/$child} -- bash-only
substitution in a payload run under plain sh (dash on Ubuntu), which
aborts the script AFTER the serve was spawned. The client then saw an
unknown failure, ran its error cleanup (deleting the token file), and
the just-booted serve died on the missing token -- orphaning one serve
per attempt. Replaced with a POSIX sed substitution.
Adds two regression tests: payload variables must keep $HOME expandable
(no re-quoting), and the pid substitution must be POSIX sh. Both fail
against the previous code; all 89 remote-lifecycle tests pass with the
fix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Desktop can drive a full update of a REMOTE SSH instance: claim the
connection (ManagedConnectionUpdateGate pauses dials/mutations while the
update owns it), terminate the owned backend with an identity re-proof
at the signal boundary (terminateOwnedDashboardForUpdate — argv+creation
time re-read in the same remote shell that signals), run the updater
under an update-in-progress marker/mutex on the remote install root,
bootstrap the new backend, and fence its publication so a rollback
cannot leak a half-published serve (fenceManagedSshBootstrapPublication
+ waitForManagedSshBootstrapFence barrier).
Extraction notes (campaign #91277 Phase 4; rollout/canary engine stays
behind per sequencing):
- managed-ssh-update.ts + test: clean cherry-pick (new files).
- windows-remote-lifecycle.ts + test: applied as-is (zero main drift).
- remote-lifecycle.ts: PR hunks stitched AROUND current main's #95532
skew guards and #91668 SIGKILL-escalating cleanupStale, which are
PRESERVED verbatim — the PR's python identity-re-proof termination is
wired for the managed-update path only; connect's stale replacement
keeps main's proven kill. One PR test assertion re-pinned accordingly.
- One PR test fixed: floating coordinator.start() promise whose
rejection IS the contract under test now has an explicit handler
(vitest flagged it as an unhandled rejection).
tsc clean; 131/131 across the three touched electron suites. main.ts
wiring (IPC + deps bag) follows as a separate commit.
A backend.lock.json that exists but doesn't match what this build writes
(unknown/future schemaVersion, truncated JSON, missing or foreign
ownershipId, malformed shape) was previously indistinguishable from 'no
lockfile': connect() would spawn a fresh backend on top of it and
overwrite the record, and cleanup paths could drop foreign state —
disarming the #78872 ownership guard exactly when another (e.g. forked)
desktop build shares the remote. readLockfile now returns a skew
sentinel for existing-but-foreign lockfiles; connect() refuses with a
'remote-lockfile-skew' error and a skew warning instead of
reaping/overwriting, disconnect() and cleanupStale() skip entirely.
Refs #95532
The #95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
teardownSshConnection closed the tunnel and SSH transport but never
killed the detached serve --isolated process. Spawn uses setsid/nohup,
so the backend reparents to pid 1, keeps state.db open, and accumulates
across Cmd+Q. Reuse cleanupStale via disconnect while SSH can still
exec, sequence remote kill before close, and seal the bootstrap
coordinator so reconnect during a prevented first quit cannot respawn.
The quit race is 6s to cover cleanupStale's 5s wait-for-exit loop.
Clicking the local agent left connectionId null, so the roster treated
the registry primary (often an SSH box) as active and dropped its
profiles while inventing a "This device" shadow of default.
Inventory undialed SSH sources with a cached ls of ~/.hermes/profiles
instead of requiring the window to switch onto that machine. Hostile
HERMES_HOME values are rejected before the listing command runs.
Replaces the canonicalization test (which pinned the behavior #74425
removes) with wrapper-preservation coverage for auto-detection and an
explicit remoteHermesPath, both asserting no python3 -c parser call is
issued. Verified both fail against the pre-fix implementation.
The SSH modules predate the stricter lint config that landed on main (curly, no-empty, perfectionist sorting, prettier). Mechanical lint:fix + fmt pass, empty catch blocks filled with the codebase's void-0 convention, and inline no-control-regex disables on the three deliberate control-char patterns (same pattern as lib/ansi.ts).
Skip argv ownership verification after the remote PID is already proven dead, then remove only the validated lock/log metadata and continue with a fresh spawn.