taskkill /T /F returns when termination is INITIATED, not completed, and
the pre-handoff unlock gate only probed the venv hermes.exe shim — which
the 'python.exe -m hermes_cli.main serve' backend need not hold at all.
The gate could therefore pass on its first iteration with zero dwell
while the killed pythons were still unmapping .pyd files; the
venv-blocker scan (no liveness filter) then reported those dying
processes as holders and aborted the hand-off. Every first update
attempt from the footbar failed; the manual retry succeeded because the
process table had settled by then.
The unlock gate now lives in backend-release-gate.ts (dependency-free,
backend-child.ts pattern) and requires BOTH the shim unlocked AND every
signalled PID to have actually left the process table; stragglers
collected per-pass are killed and join the watch set. On deadline the
old shim-only criterion survives as the escape hatch — lingering PIDs
past 15s are the venv-blocker re-scan's job. applyUpdates additionally
re-scans up to 2x with a 1.5s settle before aborting on 'blocked', so
untracked grandchildren an AV driver holds in teardown stop failing the
update while a REAL holder still aborts on the third scan.
Surgical reapply of PR #78037 fix 1 by @3x3xX3N0N onto the post-#87599
code shape (stopBackendTreesForUpdate extraction, stopSafeBlockers
re-scan path). The re-scan settle idea was first submitted by
@MaheshBhushan (#74831); the killed-PID tracking seam matches
@webtecnica's #74956.
Co-authored-by: MaheshBhushan <128616744+MaheshBhushan@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Hermes <hermes@nousresearch.com>