Commit Graph

189 Commits

Author SHA1 Message Date
Teknium d66a62469b refactor(update): lift parked-branch guard out of _prepare_checkout_for_update (161 -> 101 LOC) 2026-09-02 17:03:42 -07:00
Teknium 73ca3efea4 refactor(update): lift divergence reconcile and syntax-guard rollback out of _pull_updates (153 -> 47 LOC) 2026-09-02 17:02:23 -07:00
Teknium 2d7fcf636e refactor(update): compact re-export import blocks (AST-identical, -117 lines) 2026-09-02 17:00:07 -07:00
Teknium 60041b787a refactor(update): join short multi-line statements onto one line (AST-identical, -416 lines) 2026-09-02 16:52:38 -07:00
Teknium 1aa9285312 refactor(update): hand-compact comments/docstrings in update_cmd.py, update_cmd_fleet.py, update_cmd_zip.py (AST-identical); rewrite stale module docstring 2026-09-02 16:45:44 -07:00
Teknium d46eb1f964 refactor(update): collapse 90 try/except-pass and try/except-logger.debug blocks into suppress()/_best_effort() (new leaf update_cmd_common.py) 2026-09-02 16:34:55 -07:00
Teknium ebf2875049 refactor(update): lift already-up-to-date finish out of _cmd_update_impl (290 -> 257 LOC) 2026-09-02 16:31:17 -07:00
Teknium bff23bea2c refactor(update): lift option resolution, receipt/plan, git prep, HEAD verification and CalledProcessError handling out of _cmd_update_impl (521 -> 290 LOC) 2026-09-02 16:29:04 -07:00
Teknium a59680217f refactor(update): hand-compact comments/docstrings in split modules (AST-identical); re-export get_hermes_home/shutil; repoint source-inspection tests 2026-09-02 16:11:02 -07:00
Teknium 4674f8b254 refactor(update): split zip/stash/config/deps/git/maint clusters out of update_cmd.py 2026-09-02 16:00:26 -07:00
Teknium 096826bf7d refactor(update): split gateway fleet restart/verify into update_cmd_fleet.py 2026-09-02 15:55:50 -07:00
Teknium 5206a504ca refactor(update): split Windows gateway lifecycle into update_cmd_windows.py 2026-09-02 15:40:14 -07:00
Teknium e2d7a6f8b8 refactor(update): drop dead _gateway_restart_recovery_profiles; write sibling-snapshot cache via module attribute 2026-09-02 15:27:33 -07:00
Teknium 401efb0b2b refactor(cli): decompose _cmd_update_impl into phase helpers, unify duplicated update helpers, compact comments
hermes_cli/update_cmd.py 11217 -> 10204 LOC; _cmd_update_impl 2797 -> 521 LOC.

Decomposition (behavior-neutral, AST free-name verified) — _cmd_update_impl is
now a thin orchestrator calling, in order:
  _clear_windows_venv_holders_or_exit -> _prepare_checkout_for_update
  (-> _CheckoutPlan) -> _repair_current_checkout | _pull_updates ->
  _sync_python_dependencies_after_pull -> _run_post_update_maintenance ->
  _restart_gateway_fleet_after_update (-> _GatewayRestartOutcome;
  _restart_systemd_gateway_units) -> _resume_windows_gateways_and_merge_outcome
  -> _verify_fleet_after_update. _resolve_manage_cmd hoisted to module level.

Unified helpers: _git_run (49 captured git subprocess.run sites),
_systemctl / _systemctl_reset_and_restart (17 sites + 2 pairs),
_sweep_bytecode_after_update (3x), _print_bundled_skills_sync_report (ZIP+git),
_ensure_venv_pip (ZIP+git), _self_and_non_gateway_ancestor_pids (2x),
_record_update_step (4x), _write_gateway_update_exit_code for the 3 inline
".update_exit_code" writes. Dropped duplicate nested copies of
_wait_for_service_active/_service_restart_sec/_print_items, dead
upstream_exists, a dead if/pass branch and unused imports.

Comments/docstrings hand-compacted (236 blocks, AST-identical with docstrings
normalized); rationale/invariant sentences kept.

Tests: source-inspection guards repointed to the helper that now owns the code
(test_update_self_lock, test_update_fleet_check_fail_closed,
test_update_apply_shallow_count); new regression test drives the real Windows
resume/merge helper.
2026-09-02 13:30:54 -07:00
Teknium 10088c569e fix(update): hermes update no longer hangs on a GitHub username prompt
GitHub answers anonymous fetches with HTTP 401 during outages (and for
renamed/private repos). git then prompts `Username for 'https://github.com':`
on the inherited terminal and `hermes update` sits there — users read it as
Hermes demanding a GitHub login.

Every network git call in the updater (fetch/pull/push, apply + --check +
fork sync) now runs with GIT_TERMINAL_PROMPT=0 / stdin=DEVNULL, so the 401
fails fast into the fetch-failure classifier, which now reports it as a
GitHub-side rejection (likely outage) rather than blaming the user's
credentials. Credential helpers/askpass are left configured so private-fork
origins still authenticate.

Live repro: PTY-attached update --check against a 401 origin hung 15s+ on
the prompt before; exits rc=1 in 0.2s with the diagnosis after.

Same class as #73751 (@Frowtek, pre-main.py decomposition); passive banner
half salvaged from #101421 (@RobbertC5).
2026-09-02 12:34:26 -07:00
Teknium f9bca5a0d0 fix(update): reconcile serve/dashboard runtimes in their own vocabulary and escalate survivors (#100479)
Widen the two salvaged fixes (#100490, #100493) to the whole class:

- match_runtime_outcomes: serve/dashboard rows never borrow gateway
  bookkeeping at ANY site — not just the bare hermes-gateway unit name
  (#100490) but also relaunched_profiles / externally_supervised_profiles
  and the profile-substring unit match (hermes-gateway-work credited the
  'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
  units (exact names, scope prefix tolerated) or, when the caller passes
  the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
  feed the Phase-2 reconciliation, so a surviving unmanaged serve is
  'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
  remedy instead of 'hermes gateway restart', which cannot reach it.

Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
2026-09-02 00:42:53 -07:00
TwotNguyenVN 5733d55f77 fix(update): warn surviving pre-update serve and dashboard runtimes on success (#100479) 2026-09-02 00:42:53 -07:00
kshitijk4poor d4611ac837 chore: review follow-ups for salvaged #99954
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
2026-09-02 01:52:11 +05:30
Sahilvishnaliya d0f0afb009 fix(update): post-update state.db guard covers every profile, not just the root
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).

Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.

Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes #97994.
2026-09-02 01:52:11 +05:30
kshitijk4poor 6b46725a17 fix(update): surface leftover update autostashes older than 7 days (#63717)
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
2026-09-02 01:50:21 +05:30
Teknium ab9866bc64 fix(gateway): survive Windows Job-Object teardown across gateway restarts (#48820)
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.

Three surgical changes:

1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
   The inlined watcher now routes the respawned gateway's stray
   stdout/stderr to the same sidecar log gateway_windows._spawn_detached
   uses (DEVNULL only as fallback), so a gateway killed moments after
   respawn leaves a trace. Direct implementation of the 4th repro's
   hardening suggestion (1).

2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
   canonical _spawn_detached, so the respawned gateway's exit-diag /
   lifecycle records show whether it escaped the parent Job Object — a
   job-teardown kill is no longer indistinguishable from any other silent
   death.

3. Post-update resume verifies liveness before vouching
   (hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
   runs the same provisional-hit + 2s-confirmation liveness poll every
   other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
   with all_profiles= for the fleet) before printing ✓, writes the #91675
   start attestation for the verified PIDs, and fails the resume with a
   "restart could not be verified" warning + recovery hint when no stable
   gateway appears. Suggestion (2) of the 4th repro; closes the last
   silent-success hole in the family (#84185 fixed the cold-start leg,
   #91675 the direct-start leg; this is the relaunch leg).

Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.

Fixes the Bug-1 relaunch-trust leg of #48820.
2026-09-01 11:43:36 -07:00
Teknium e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
Teknium 045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
Teknium ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium 9f377264aa refactor(update): consolidate the gateway drain triage into one shared helper
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.

No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
2026-09-01 08:34:51 -07:00
salch-cred 49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Justin Wilson fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
joaomarcos 27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium 0f16e413e6 fix(update): run pending fleet-restart catchup before the runtime-verification exit gate
The #95294/#91277 fleet contract requires the pending-restart check to
execute on every already-up-to-date pass; the verification exit(1) now
fires only after the catchup runs, so a vulnerable runtime demotes the
outcome to partial without stranding the fleet on stale code.
2026-08-31 12:58:14 -07:00
Teknium d6bf89de40 fix(update): treat an unprobeable post-update interpreter as non-blocking
Verification-features-must-not-false-positive-on-rollout rule: only a
POSITIVE vulnerable SQLite probe demotes success to partial. A dev
checkout without a venv (or a failed probe subprocess) keeps the
success banner and the fleet-restart catchup path — the CI-red
test_update_fleet_restart_pending already-up-to-date trio pinned this.
2026-08-31 12:58:14 -07:00
fangliquanflq 9caf31394f fix(update): verify repaired current-checkout runtime 2026-08-31 12:58:14 -07:00
fangliquanflq 4f19808915 fix(update): fail current-checkout runtime verification 2026-08-31 12:58:14 -07:00
fangliquanflq 7e22725bb4 fix(update): verify SQLite runtime remediation 2026-08-31 12:58:14 -07:00
SayHell0W0rld 6cccc2ef7e fix(update): locate PortableGit under the shared root, not profile home
Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.

Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
2026-08-31 12:21:46 -07:00
SayHell0W0rld 0bab9ff9e7 fix(update): self-heal broken Git-for-Windows trampoline on Windows
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).

Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
2026-08-31 12:21:46 -07:00
andyst-dev 4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00
Teknium ce942ab125 fix(update): fingerprint orphan backends from the classification psutil handle
The orphan-backend classifier fingerprinted candidates via
gateway.status.get_process_start_time, which prefers /proc/<pid>/stat —
the HOST process table, in clock ticks. Under the fake-psutil test harness
(and any containerized run where the PID number happens to exist on the
host) that returns the WRONG process's fingerprint in the WRONG units,
while pid_is_hermes verifies via psutil centiseconds at kill time: the
guard would then refuse every legitimate reap. Read create_time() from the
same psutil handle used for classification, quantized exactly like
gateway.status does on Windows, so the fingerprint round-trips.

Also covers the Windows-lane sibling: test_uses_netstat_and_taskkill_on_windows
now pins the guarded call path, plus a new refusal test for a non-bridge
listener PID (#89614 class).
2026-08-31 10:41:54 -07:00
Ayush Nangia c923b53913 fix(update): refuse gateway ancestor tree-kill on Windows 2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
fangliquanflq 915ec169a0 fix(update): verify failed restore cleanup 2026-08-31 10:04:56 -07:00
fangliquanflq 53cf38c3b8 fix(update): fail closed on incomplete restore checks 2026-08-31 10:04:56 -07:00
fangliquanflq 3e4dc5f2c3 fix(update): preserve unknown restore cleanup state 2026-08-31 10:04:56 -07:00
fangliquanflq 8236b51878 fix(update): authenticate import health markers 2026-08-31 10:04:56 -07:00
fangliquanflq 6c608e2f59 fix(update): reject terminated import probes 2026-08-31 10:04:56 -07:00
fangliquanflq 68f9681acb fix(update): capture terminating restored imports 2026-08-31 10:04:56 -07:00
fangliquanflq f6c3942957 fix(update): compare every restored module failure 2026-08-31 10:04:56 -07:00
fangliquanflq 7b958b3575 fix(update): detect restored import-time failures 2026-08-31 10:04:56 -07:00