Commit Graph

118 Commits

Author SHA1 Message Date
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
Teknium 2f9e187001 revert(macos): remove the TCC interpreter anchor — anchored copies could not load libpython
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.

Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).

Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.

Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
2026-08-26 06:50:53 -07:00
Teknium 7125e839d3 fix(state): fail closed when a live process still holds state.db during destructive restore (#90950)
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:

- restore_quick_snapshot's unlink+move fallback replaced the inode and
  deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
  the live writer's next checkpoint writes wrong-offset pages (the
  page-1 compression_locks clobber from the report).

Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
2026-08-26 06:28:43 -07:00
briandevans 33dce0eb7e refactor(update): fold the auto-restore sequence into a shared helper
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.

Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.

The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.

Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
2026-08-26 06:28:43 -07:00
briandevans 86d719067b fix(update): clear stale SQLite sidecars before auto-restoring state.db
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.

The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.

copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.

A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.

Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.

Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
2026-08-26 06:28:43 -07:00
ethernet 34c5fcb2a2 fix: update on macos referenced nonexisting variable 2026-08-26 00:24:26 -07:00
webtecnica cc5ff96f9e fix(macos): stable TCC anchor for uv-managed python interpreter (#85345) 2026-08-25 22:10:12 -07:00
David Metcalfe 2d37ed056e fix(macos): harden TCC check against codesign timeouts, clarify scope
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
  so a hanging codesign degrades to the unreadable-DR warning, never crashing
  the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
  matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
  com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
2026-08-25 21:59:36 -07:00
David Metcalfe 36c1755065 fix(macos): detect stale TCC grants and guide one-time re-grant
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.

- hermes doctor: new check_macos_tcc_grants() reports the desktop
  bundle's DR class (cdhash-pinned → grants reset on every update;
  identifier-pinned → stable) and prints the exact stale-grant repair
  (tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
  installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
  documents the one-time re-grant for pre-fix grants.

Closes #86385
2026-08-25 21:59:36 -07:00
Royalaid 0c23bf19af fix(update): defer interactive CUA installs on Windows 2026-08-25 16:53:01 -07:00
Krzysztof Radzikowski 36f1423411 fix(update): gateway-only concurrent instances no longer abort hermes update (#37039)
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.

The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.

Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.

Co-authored-by: Hermes <hermes@nousresearch.com>
2026-08-25 14:48:51 -07:00
Jeff Mettel bee489d563 fix(update): restart a booted-out launchd gateway instead of silently skipping it (#74973, salvage #75021)
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.

- _restart_launchd_gateway_after_update() (his extraction, adapted):
  plist-exists is the ONLY gate; launchd_restart() owns the
  bootout/bootstrap/kickstart ladder for every plist-present state;
  every failure path is loud and names the manual recovery command.
  The gate-error 'except: pass' (the second silent variant) now counts
  the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
  supervised PID), composing his fix with the returned-is-not-supervised
  guard.
- His regression suite adapted to the (restarted, failed) contract; the
  old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
  bug.

A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
2026-08-25 14:10:10 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
RelaxJonh d74bbb9bd9 fix(cli): treat empty fleet probe as incomplete when gateways were restarted (#93406)
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list.  print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.

After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy.  Treat it as incomplete so
the receipt records "partial" and the exit code is 1.

Fixes #93406
2026-08-24 03:21:18 -07:00
Teknium 706f33d424 feat(update): sibling profiles' configs migrate with the fleet — no more silent version drift (#20438/#54926/#79048)
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).

_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.

Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
2026-08-23 05:12:59 -07:00
Teknium 18b7fc82b6 feat(update): the plan is now the restart worklist — every planned runtime must be accounted for (#91277 Phase 2)
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).

- update_inventory: restart_via becomes a machine-readable mechanism id
  (systemd|launchd|desktop|manual) — THE policy table as data; display
  derived via describe_restart_mechanism. match_runtime_outcomes()
  reconciles every planned runtime against the restart phase's
  bookkeeping (restarted/stopped/failed/unaccounted);
  report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
  land in the receipt (runtime_outcomes); any unaccounted runtime
  escalates exactly like a STALE/DOWN fleet row (exit 1).

Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
2026-08-23 04:45:29 -07:00
Jack Lau 1bf93660f1 fix(update): verify launchd is supervising the gateway after a restart
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).

_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.

That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.

Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.

A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.

Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.

tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
2026-08-23 04:45:29 -07:00
Jack Lau dfcef70061 fix(update): stop a gateway we cannot relaunch instead of leaving it on stale code
Fixes #88654.

After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:

    restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
    if restart_mode is None:
        continue

A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:

    cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'

with no operator signal of any kind.

Two changes.

_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.

When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
2026-08-23 04:25:18 -07:00
Teknium 0c14f060db fix: import managed_python_env at the git-path site; assert the managed-env contract in the repair test
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
2026-08-23 03:55:14 -07:00
Teknium fbfdb9312b fix(update): widen UV-env isolation to the sibling dependency-sync sites
The salvaged fix covered the git-path sync; the same raw-os.environ
construction existed at the main update path and the interrupted-install
recovery path. All three now build their uv env via managed_python_env()
(#83914 class — same bug, all sites).

A/B-proven with real uv: poisoned UV_PYTHON/UV_SYSTEM_PYTHON steers the
merge-base construction into the hijacker's interpreter (VERDICT:
HIJACKED); the managed construction installs into the install's venv
(VERDICT: ISOLATED). Compose-checked with #92824's stale-VIRTUAL_ENV pin:
isolation + pin together install into the running interpreter on the
site-packages shape.
2026-08-23 03:55:14 -07:00
suntech-wang 6ce145f38f test(update): lock managed uv-env isolation regression
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
  UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
  UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
  the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
2026-08-23 03:55:14 -07:00
suntech-wang 08f5a0a98b fix(update): isolate pip install from third-party UV env vars
uv respects UV_PYTHON_INSTALL_DIR from the process environment. When a
third-party app (e.g. WorkBuddy) sets a User-level UV_PYTHON_INSTALL_DIR,
the update's uv pip install can target the wrong interpreter and fail
installing extras, leaving the venv entry-point shims missing. Use the
official managed_python_env() isolation (drops VIRTUAL_ENV/PYTHONPATH/
UV_PYTHON, forces UV_PYTHON_INSTALL_DIR to .hermes-runtime/python,
UV_NO_CONFIG=1) and then point VIRTUAL_ENV at this install's venv.
2026-08-23 03:55:14 -07:00
Franci Penov e366df6889 fix(cli): treat a fork's upstream sync as an update
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.

Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.

Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.

steps still being skipped afterwards.

Refs #73108
2026-08-23 00:19:46 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Teknium f4067774aa fix(update): token-based control-plane classifier + live E2E for the Desktop-lifecycle cold-start skip (#76129 salvage follow-up)
On top of @686f6c61's premise-corrected #76745:

- _looks_like_desktop_control_plane now uses the parser-derived
  _hermes_holder_subcommand instead of substring matching — the
  #90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
  argv no longer read as control planes). Regression test added,
  sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
  supervised serve owns lifecycle; killed spawner (orphan) does not;
  dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
  entry suppresses the actual cold-start plan; dead serve restores it;
  holder-scan fallback rung proves the token classifier live.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-08-22 21:36:44 -07:00
686f6c61 4ccc4b6931 fix(update): skip Windows gateway cold-start when Desktop owns lifecycle
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
2026-08-22 21:36:44 -07:00
Teknium c9c44d0df9 Merge pull request #92636 from NousResearch/fix/windows-launcher-managed-bin
fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
2026-08-22 20:58:00 -07:00
Teknium 0c435f4601 fix(update): reword refusal message — footgun linter matched prose 'venv open (' as bare open() 2026-08-22 19:30:10 -07:00
Teknium 83864c0b5d fix(update): a contended venv is never mutated — failed shim quarantine now refuses instead of warning (#87331)
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.

- _run_quarantined_install gains strict_quarantine: any shim whose
  rename failed every retry aborts BEFORE the install command runs
  (successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
  boundary turns the error into a refusal: defer via the
  update-incomplete marker, exit 2 (recorded as refused by the receipt
  net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
  (their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
  unconditionally: marker survives, next launch retries after the
  holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
  without FILE_SHARE_DELETE (the exact field lock shape), strict path
  refuses with zero installer invocations, releases roll back, and the
  same path proceeds once the holder exits.

Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
2026-08-22 19:30:10 -07:00
emozilla fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
Gille 9782275b2a fix(windows): restore dedicated CLI launchers on update 2026-08-22 00:05:51 -07:00
JonthanaHanh 01c14ad7f3 fix(update): ZIP swap preserves the built desktop app (apps/desktop/release)
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).

Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-08-21 22:08:16 -07:00
kshitijk4poor eac3f645ef fix(update): don't ZIP-fallback on dependency failures or dirty trees
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.

- ZIP fallback now keys on git ACTUALLY having failed
  (_should_zip_fallback_on_update_error): a dependency-install failure
  after a successful pull can't be fixed by re-downloading source and
  would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
  (-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
  re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
  'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
2026-08-21 22:08:16 -07:00
Teknium 7d6db4efb8 fix(update): holder classifier derives value-flags from the real parser; de-flake goal-resume fixture
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.

De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
2026-08-21 19:11:55 -07:00
Teknium c02cac00ce fix(update): venv-holder labels parse the real subcommand; gateway ancestors stay visible to the scan
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.

#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.

15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
2026-08-21 19:11:55 -07:00
Teknium 04acfb9673 fix: remove function-level 'import time as _time' that shadowed the module import
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
2026-08-21 15:23:49 -07:00
Willian Santos 4fad27a101 feat(update): --switch-branch opts an unmerged branch out of the in-place merge
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.

--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.

Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.

Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Willian Fernandes Santos 91096bb2f0 feat(update): update branches carrying unmerged commits in place instead of skipping
The parked-branch guard (8ce8ffd429) distinguishes checkouts by what the
branch carries, then treats both non-clean cases the same: a stale
fully-merged leftover is switched back to the target (correct), but a
branch with unmerged commits — a branch someone is actually working on —
gets CODE UPDATE SKIPPED and exit 1. For anyone running a maintained
custom branch on top of main, every update now refuses, and the guidance
('checkout main') abandons their branch.

The guard's own reason codes already separate the cases, so use them:

- fully merged      -> switch back to the target (unchanged)
- unmerged:N        -> update the branch IN PLACE: fetch, then bring
                       origin/<target> into the checkout. Fast-forward
                       when possible; on divergence, a true merge behind
                       a pre-update safety tag, stopping cleanly on
                       conflict. The checkout never moves; local commits
                       survive; the running code advances.
- dirty/unverifiable/opted out -> skip loudly (unchanged)

The post-pull success gate learns that an in-place update legitimately
ends on a non-target branch: origin/<target> was merged INTO the checkout,
so refusing to claim success there would fail every update that did
exactly the right thing.

Guard tests updated: the unmerged case now asserts the in-place outcome —
target code arrives (b.txt from c3), the branch's own commit survives, and
HEAD never moves. 18/18 guard tests, 20/20 with the diverged-update suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 15:23:49 -07:00
Teknium bbbc50acc2 fix: hermes update no longer strands non-interactive updates on a parked branch with unmerged commits
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.

Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
2026-08-21 15:23:49 -07:00
Teknium 1575116629 fix(update): pre-update snapshots now cover every profile, not just the invoking one (#66140)
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.

- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
  snapshot set, per-file 1GiB cap, and keep policy as the invoking
  profile (no partial tier, no new restore-coherence class), each into
  the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
  runs the #34600 cron-loss safety net per profile against its OWN
  snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
  profile's (best-effort, receipt-recorded); post-update cron restore
  extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
  behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
  jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
2026-08-21 13:01:35 -07:00
Teknium 0aecadc17c feat(update): hermes update --plan — read-only fleet inventory + plan phase in every update
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.

- hermes_cli/update_inventory.py (new): side-effect-free runtime
  inventory — install kind via detect_install_method (git / docker / nix
  / apt, updatable-in-place or not, with the correct external update
  command for image/package-managed installs), all profiles, every live
  gateway with its supervisor (systemd / launchd / manual via the
  fleet-wide _get_service_pids), running code_sha/code_version from the
  #91283 gateway_state.json stamps, and the restart mechanism each
  runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
  docker/nix refusal gates so image-managed installs get a useful
  'not updatable in place + right command' report instead of a bare
  refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
  ('plan' key) and prints a one-line fleet summary, so post-mortems can
  compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
  cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
  dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
  never-raises, JSON round-trip for the receipt, print output shapes,
  receipt integration.
2026-08-21 04:23:13 -07:00
PT f29ee96dd3 fix(update): restart all macOS launchd gateways on hermes update
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:

- _restart_macos_launchd_gateways(): the invoking profile keeps the
  existing launchd_restart() path; every other gateway of this install
  is drained via SIGUSR1 (same as systemd siblings), then hard-
  kickstarted unless KeepAlive already respawned it, then verified on a
  fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
  counts toward failed_or_stale_units — including timeouts during
  liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
  derives labels from THIS install's profiles (get_default_hermes_root),
  not by globbing the shared per-user ~/Library/LaunchAgents — a
  sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
  installs) must never enumerate, let alone restart, another install's
  fleet. This also keeps the hermetic test suite blind to a dev
  machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
  liveness, kickstart, and fresh-PID verification all use the domain the
  service was actually located in (gui/<uid> vs user/<uid> probed per
  label via `launchctl print`). This addresses the #41403 review defect:
  the process-wide _launchd_domain() cache resolves the current profile's
  domain and must never be reused for a sibling. _launchd_domain() itself
  becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
  sweep excludes every gateway service PID (mirror of the systemd
  hermes-gateway* pattern) so it cannot mistake a freshly respawned
  sibling service for a stale manual gateway. Default-scope callers
  (gateway status, cron checks, stop_profile_gateway's orphan reaper —
  which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
  hints for launchd labels alongside the systemctl ones.

Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).

Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 03:51:03 -07:00
Tranquil-Flow d8047c303b fix(gateway): BSD-compatible ps flags and all-profile launchd pid discovery (#74075)
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
  on macOS/BSD ps, making the fallback silently return [] on every macOS
  machine. The matcher only needs argv (not env vars), so e is
  unnecessary. -ww keeps unlimited-width output on both BSD and
  procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
  on macOS, enumerate every ai.hermes.gateway* launchd agent across
  profiles via bare launchctl list instead of only the current
  profile's label. This prevents the update sweep from misclassifying
  sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
  _get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
  pass all_profiles=True so the update fleet sweep excludes every
  service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
  not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
  launchctl list <label>, all_profiles uses bare launchctl list
  with prefix filtering, handles empty/broken output gracefully, and
  preserves systemd behavior.

Tranquil-Flow
2026-08-20 23:41:00 -07:00
Teknium 1acbeed146 fix(update): resolve code identity by reading .git directly — no subprocess
get_code_identity() shelled 'git rev-parse HEAD', which broke two tightly
mocked test suites (sequenced subprocess.run side effects in the
head-moved gate, call-count asserts in the Windows taskkill test) and
added process-spawn cost to gateway runtime-status writes.
_resolve_git_head_sha() now reads HEAD/refs/packed-refs directly,
handling regular checkouts and worktree/submodule pointer files.
Also: skip the 2s fleet settle wait when the restart phase touched no
gateways, and hoist killed_pids init outside the restart try-block.
2026-08-20 21:54:25 -07:00
Teknium 1d74833d8d feat(update): structured update receipts + post-update fleet version verification
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.

- hermes_cli/build_info.py: get_code_identity() — process-cached code
  identity (git sha for source installs, baked .hermes_build_sha for
  Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
  code_sha/code_version into gateway_state.json, so a running gateway's
  actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
  update run (steps, skips with reasons, gateway restart outcome, fleet
  snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
  pointer for the dashboard/desktop; plus collect_fleet_versions() /
  print_fleet_version_matrix() comparing every live profile gateway
  against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
  git, ZIP, and hard-failure paths; after the restart phase, prints the
  fleet version matrix and escalates provably-stale gateways into the
  existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
  gateways report 'unknown' and never fail the update (no false
  positives during rollout).

Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
2026-08-20 21:54:25 -07:00
brooklyn! 0723cb6c06 fix(update): don't reinstall the editable package when the pull can't affect it (#90967)
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.

Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.

The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.

This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.
2026-08-20 12:41:33 -05:00
Teknium 044acf2bf7 fix(update): gateway auto-restart no longer dies on stale cached modules after the pull
`hermes update` runs in the pre-pull interpreter. The auto-restart phase
imports freshly-pulled gateway source, which resolves sibling imports
against the OLD sys.modules cache — so any update where an already-cached
module gained a new export ImportErrored the whole phase and left the
gateway serving pre-update code (2026-08-20 field failure: new gateway.py
needs cli_output.line_input, cached cli_output predates it).

Class fix replacing the per-symptom _UPDATE_RUNTIME_RELOAD_MODULES
approach: _purge_stale_hermes_modules() evicts every cached module under
the Hermes package prefixes (hermes_cli/gateway/tools/tui_gateway/agent)
right before the restart phase, so later lazy imports rebuild a
self-consistent module graph from the updated checkout. The updater's own
executing modules are exempt (purging them buys nothing; reload-in-place
is the unsafe op, and we never reload). Root-segment check spares
prefix-lookalike packages. Best-effort, never raises.

5 new tests incl. an end-to-end repro of the field failure shape
(stale module missing symbol -> ImportError -> purge -> import resolves).
2026-08-20 05:02:36 -07:00
Teknium 5dd221d442 feat: desktop updates no longer re-apply local source edits (--keep-stash)
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').

New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.

Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
2026-08-20 04:54:14 -07:00
Teknium 95fa814269 feat(process): positive process identity — spawn tags, machine spawn ledger, Windows job-object self-attach
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:

- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
  spawn-ledger.json self-registration keyed on (pid, create_time) — PID
  reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
  and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
  the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
  themselves at startup and attach to the job; Desktop legacy
  HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
  works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
  _ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
  backends (purpose reapable + recorded spawner provably dead) in ANY update
  context. Ledger-unknown holders fall through to the existing rungs.

22 new tests, sabotage-verified.
2026-08-20 04:47:38 -07:00