Commit Graph

16 Commits

Author SHA1 Message Date
Teknium 7d28a33bd9 refactor(web): simplify dashboard_procs/dashboard_register (dedupe home/profile helpers, drop _HEX16 alias, compact docs) 2026-09-02 13:33:23 -07:00
joaomarcos 2d70d67b15 fix(update): stop the managed dashboard restart from skipping serve backends
`_kill_stale_dashboard_processes(restart_managed=True)` returned as soon as
`_restart_managed_dashboard_service()` handled `hermes-dashboard.service`.
On a host that runs both that unit and `hermes-serve.service` -- the exact
unit set in #92145 -- the serve backend hosting `tui_gateway` was never
scanned, never stopped and never restarted, so it kept its pre-update
`sys.modules` after the checkout advanced.

The early return exists so the dashboard's own PID is not raw-killed
(systemd reads our SIGTERM as a clean stop). That only requires excluding
the unit, which the `already_restarted_units` filter below already does.
Record the unit as handled and continue the pass instead of ending it.
2026-09-01 07:00:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
liuhao1024 57309c0cbb fix(update): never respawn backends from a foreign HERMES_HOME (#94030)
The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.

Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
2026-08-26 08:39:04 -07:00
fangliquan 858916acc4 fix(update): preserve SSH ownership only during updates 2026-08-26 08:39:04 -07:00
fangliquan 1676c614b3 fix(update): preserve SSH-owned backends during cleanup 2026-08-26 08:39:04 -07:00
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
Teknium 3c44cd0c67 docs: record why the reap grace window exists in the reaper docstring
Follow-up for salvaged PR #91994 -- without this note a future cleanup
pass could read the age gate as dead weight.
2026-08-23 00:52:50 -07:00
Ishman82 b44c2bdab2 fix(desktop): spare concurrently starting backends
Desktop writes backend.lock.json only after a remote profile backend announces readiness. Concurrent Bot Mode profile starts could therefore see their young siblings as unowned PPID-1 processes and mutually reap them, causing SSH reconnect storms and stale turn leases.\n\nProtect unregistered backends for a bounded startup grace period, fail closed when age cannot be read, and cover young, unknown-age, boundary, lock-owned, and old-orphan behavior.
2026-08-23 00:52:50 -07:00
chelsealong 0b42cae068 fix(update): tighten hermes-serve unit gate, dedupe fleet/cleanup restarts
Review on #83595 flagged two service-lifecycle gaps in the hermes-serve
restart support:

- The unit-name gate accepted anything starting with "hermes-serve",
  which also matched the unrelated hermes-server.service. Require the
  exact base unit or the hyphenated profile family instead.
- The fleet-restart loop and _finish_dashboard_update_cleanup() could
  both restart the same hermes-serve unit — the loop restarts it
  directly, then cleanup's PID scan finds the fresh process and
  restarts its owning unit again. Thread the fleet loop's restarted
  unit names through to _kill_stale_dashboard_processes() so it skips
  units already handled.
2026-08-16 10:51:50 -07:00
kshitij 4e3de140c1 fix(cli): bound the Windows process-scan probes so a slow WMI scan cannot wedge hermes update (#87134)
subprocess.run(capture_output=True, timeout=N) is not hang-safe on
Windows: after the timeout fires, run()'s cleanup kills the direct child
and then joins the pipe reader threads with an UNBOUNDED communicate().
A descendant (conhost.exe under wmic/powershell) holding duplicated pipe
handles keeps the pipes from EOF and the join never returns.

_scan_gateway_pids() runs its wmic / Get-CimInstance Win32_Process scans
exactly that way, and on machines where the full process scan genuinely
exceeds its 10/15s budget (cold WMI on first boot, ARM VMs, heavy
Update/AV activity) hermes update wedged forever inside
_pause_windows_gateways_for_update() before printing a single line —
observed live on a fresh Windows 11 ARM64 VM with a faulthandler stack
pinning the main thread in subprocess._communicate and only a conhost.exe
child surviving. The single-flight update lock then blocks retries until
the wedged process is killed by hand.

This is the same deadlock class bounded_git_probe already fixed for git
probes (#68609 / #66037). Generalize that proven pattern into a shared
bounded_probe_run() — explicit communicate(timeout), kill_process_tree on
failure, bounded 1s drain, then abandon the daemonic readers — and
migrate the whole call-site class onto it:

- hermes_cli/gateway.py _scan_gateway_pids (the site that hung; reached
  from hermes update, cron, gateway restart/status, dashboard)
- hermes_cli/dashboard_procs.py wmic scan (same shape, reached on update)
- hermes_cli/claw.py tasklist + PowerShell probes (same shape; its
  try/except cannot catch a hang because a hang raises nothing)
- bounded_git_probe now delegates to bounded_probe_run (identical
  contract, one copy of the cleanup logic)

Unlike bounded_git_probe, bounded_probe_run returns the CompletedProcess
(or None) rather than collapsing to stdout, because the gateway scan
branches on returncode to trip its wmic -> powershell fallback.

Tests: tests/hermes_cli/test_bounded_probe_run.py covers success,
nonzero-exit passthrough, spawn failure, bounded timeout (fails against
the old unbounded semantics — verified by sabotage), errors= decoding,
DEVNULL stdin, POSIX process-group placement, and the bounded_git_probe
delegation contract. Existing test_git_probe_tree_kill.py passes
unchanged against the delegated implementation.

Closes #87134
2026-08-16 06:32:08 -07:00
Jeremy 1f6f86119f fix(cli): stop hermes update from respawning orphan serve --port 0 (#78821)
Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
2026-08-14 21:46:32 -07:00
Teknium 9b1a2a14ca fix: use psutil.pid_exists for orphan-reap liveness probe (Windows footgun lint)
os.kill(pid, 0) sends CTRL_C_EVENT on Windows (bpo-14484). The reap path
is POSIX-only, but the blocking lint rejects the pattern repo-wide and
psutil is a core dependency.
2026-08-10 17:02:56 -07:00
Leon Phull 888624ae61 fix(cli): never reap serve processes owned by a valid backend.lock.json
Production incident: the orphan reap killed a legitimate SSH remote backend
started by another client machine. Its process sat at ppid 1 with the same
cmdline shape as a genuine orphan, and the exclusion list only covered THIS
app instance's children — ownership by OTHER clients was invisible.

The reap now treats every backend.lock.json under ~/.hermes/desktop-ssh/*/
as an ownership claim: lock payloads are schema-validated (mirroring
remote-lifecycle.ts) and their PIDs are excluded both before the scan and
re-checked after it (defense in depth against a lock written mid-scan).

Regression tests cover the exact incident shape: a lock-owned PID and a
genuine orphan with identical process shapes — only the orphan is reaped.

Also: fold the new single-field `runtime` config category into `agent`
(_CATEGORY_MERGE) and fix an env leak in the serve-startup test
(HERMES_SERVE_HEADLESS restored via monkeypatch) so the combined suites
run green in any order.
2026-08-10 17:02:56 -07:00
Leon Phull 6386c75306 fix(desktop): reap orphaned local serve backends on desktop boot
When Desktop exits uncleanly, leftover `hermes serve --host 127.0.0.1 --port 0`
processes can be reparented to pid 1 and keep full MCP trees alive. The next
boot then stacks another backend on top of the corpses until EMFILE kills
sidebar/session APIs and tabs disappear.

- Detect Desktop-local serve shape (loopback + ephemeral port 0)
- Only reap processes whose ppid is 0/1 (true orphans)
- Spare fixed-port remote serves (e.g. --port 9119) and HERMES_DESKTOP_CHILD_PID
- Run at Desktop backend start (HERMES_DESKTOP=1) before parent-death watchdog

Complements parent-death watchdog (prevents future orphans) and configurable
nofile soft limit (capacity floor). Together these stop the multi-backend
pile-up cascade observed on macOS Desktop SSH/local installs.
2026-08-10 17:02:56 -07:00
teknium1 c64a4d75e5 refactor: extract dashboard process-hygiene helpers to dashboard_procs.py 2026-07-29 10:59:54 -07:00