Sixth review - the first against the interval architecture - verdict:
the architecture holds (idempotence, order-independence, 4-process
concurrent-writer safety, rollback immunity, format consistency, and a
120-permutation order sweep all verified), with ONE high finding, which
I had independently reproduced while the review ran: the FORWARD clock
adversary was unhandled, and unlike every other failure mode in this
subsystem it failed OPEN.
The 'obs' mark is a MAX-upsert - monotonic in the leak direction. One
glitched-forward sample (NTP flap reading 2099) while consented dragged
last_confirmed_at to 2099; a later revoke stamped closed_at = 2099; the
closed window then CONTAINED every refused period that followed. Both
the reviewer and I reproduced refused packages becoming gate-eligible.
The rollback twin was mutation-tested since round 5; nobody had asked
whether the mirror image existed.
Two clamps, each covering what the other cannot:
- The obs mark advances at most MAX_OBS_ADVANCE_SECONDS (30 days) per
call. Honest heartbeats never bind it; a machine off for months
catches up in a few hook fires (fail-closed latency only); one insane
sample moves the horizon by a bounded step that real time overtakes.
- A close is MIN(last_confirmed_at, closing observation's raw stamp).
Confirmed-time keeps unobserved gaps out of windows (v1's leak); the
raw stamp lets an honest clock at revoke time pull a poisoned horizon
back to the true revoke moment. A rolled-back clock at close time
only closes earlier - fail-closed.
Also from the review:
- D2: the data-mark advance in the REAL package writer had no coverage
(the harness re-implemented the insert; deleting the production line
survived 314 tests). Now driven through create_and_export_package_if_due.
- D3: the "don't create ~/.hermes/telemetry for fully-disabled users"
skip was dead code - the store constructor creates the directory
before the exists() check ran. The probe now checks the default path
without constructing; verified empirically on a fresh HERMES_HOME.
- Upgrade note in A.4: pre-interval backlog is never transmitted after
upgrade (fail-closed; deliberate).
New harness scenarios: forward-poison-then-revoke (the leak), and
forward-poison-cannot-wedge (the cap). Mutation check: unclamping the
close, removing the cap, and removing the real writer's data-mark
advance each fail the suite.
273 tests pass; ruff and windows-footguns clean; staging E2E 202.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.
- close_browser_holding_profile: graceful terminate → kill → poll until the
cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
unrelated same-name process on a different dir).
- Config key + docs admonition.
Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.
Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
Live windows-latest proof settled it: a running Chrome opens its cookie DB
deny-all (even CreateFile with FILE_SHARE_READ|WRITE|DELETE fails; sqlite
mode=ro/immutable/nolock all 'unable to open'), so copy-while-running is
impossible on Windows without VSS/admin — and the prior code HUNG ~24min on the
locked file.
Fix: a fast up-front lock probe (_profile_is_locked: one open() of the active
profile's cookie DB; PermissionError = locked) runs BEFORE any copy in
snapshot_real_profile. If locked, bail immediately with 'fully quit the browser
(incl. background/tray) and retry, or turn browser.use_real_profile off'. Never
hangs, never a silent signed-out copy. POSIX has no mandatory locking so the
probe never trips there — copy-while-running still works on macOS/Linux.
Docs: admonition stating Windows needs the browser fully closed (background
apps included); the live-drive-while-running path is #95669.
Tests: lock-probe unit coverage (readable/no-db/PermissionError), snapshot
fails-fast-no-copytree when locked. Windows live E2E asserts the fast-fail
contract (returns <30s with the quit message) + the read-strategy diagnostic.
On Windows a running Chrome holds Cookies / Login Data / Web Data with an
exclusive lock, so the raw file copy the snapshot used raised WinError 32
('being used by another process') and the best-effort skip left a signed-out
copy — the reported profile-cloning failure.
Fix: copy the SQLite auth DBs via SQLite's online-backup API (read-only
connection + Connection.backup()), which reads a consistent COMMITTED snapshot
while the writer holds the lock. Non-DB files (Preferences, Local State) stay a
plain copy. If even the online-backup can't read a DB, snapshot_real_profile now
FAILS CLOSED with an actionable 'close <browser> and retry' message instead of
launching a silently signed-out session.
- _copy_auth_file: sqlite-backup for Cookies/Login Data/Web Data, raw copy
otherwise, raw-copy fallback if backup fails.
- Drop -journal/-wal/-shm sidecars from the auth set + snapshot ignore: the
backed-up DB is self-contained; a stale sidecar next to it corrupts it.
- Fresh copytree excludes the auth DBs (raw copytree of a locked file raises on
Windows); they're always mirrored lock-aware afterward.
- _mirror_profile_auth returns the count of DBs it could not copy so the caller
can fail closed.
Tests: locked-DB copied-via-backup (open write txn = live-lock analog, 42
committed rows, uncommitted excluded, no journal sidecar), _copy_auth_file DB vs
plain, fail-closed when unreadable. 192 browser tests pass. Live: 68 real
cookies copied through the backup path and the session launches.
Addresses the round-3 findings from @Adolanium + @kshitijk4poor on #95620:
1. Overlay-before-reuse race (blocker): _real_profile_cdp ran snapshot_real_profile
BEFORE the session-reuse check, so a cold resolve that ends in reuse rewrote
Cookies/Login Data under a live Chromium holding the user-data-dir open (torn
DBs, locked txns, phantom logouts). Now: resolve copy dir as a PATH, probe
reuse first, return early on a hit; snapshot/overlay only on the relaunch
path when no live browser owns the dir.
2. Torn first copy poisoned freshness forever: freshness keyed on isdir(Default),
so a half-written copy (disk full / Ctrl+C) was treated as populated and only
ever got auth overlays. Now gated on a .hermes-snapshot-complete marker
written only after a full copy succeeds; a torn copy is rebuilt from scratch.
3. Consent revocation left copied credentials on disk: turning use_real_profile
off now deletes ~/.hermes/browser-profile/ on next browser use
(cleanup_real_profile_snapshots), so cookies/logins don't outlive consent.
4. Stale non-active profile copies: only the ACTIVE profile (last_used) is copied
into the copy's Default now — other Chrome profiles are never snapshotted
(smaller copy, no stale credential dirs lingering).
5. Docs/config/desktop wording aligned to actual behavior (active-profile only,
refresh on fresh session, consent-off cleanup).
Tests: overlay-skipped-on-reuse + overlay-runs-on-relaunch, done-marker gating +
torn-copy rebuild, active-only copy, consent-off cleanup (removes store +
idempotent + triggered from _real_profile_cdp). 206 browser + 222
backup/file_safety pass. Live: reuse skips re-snapshot; direct launch on the
active-only copy loads the real signed-in Gmail inbox.
Addresses the five findings from @kshitijk4poor + @GottZ on #95620:
1. macOS 26 LSHandlers parser returned a version number ('7559.97') from the
nested LSHandlerPreferredVersions block instead of the bundle id — detection
returned None on a machine whose default IS Chrome. Strip the nested block
before the role regex.
2. Wrong profile launched (the LinkedIn/Gmail 'logged out' bug): Chrome opens
Default, but the session lives in Local State profile.last_used (e.g.
'Profile 6'). Resolve last_used and mirror its auth files into the copy's
Default on both fresh and refresh paths, so the launched browser is signed in.
3. Private-URL sidecar carried the real cookie jar to arbitrary LAN hosts:
_create_local_session gains allow_real_profile (default True); the
force_local sidecar passes False → always a throwaway profile, and a
real-profile resolve failure no longer breaks private-URL routing.
4. Snapshot permissions were set once (fresh only): now secure the snapshot dir
AND its browser-profile parent on every consented launch.
5. browser.engine=lightpanda + consent gave an unactionable error: guard with
_using_lightpanda_engine() before detection, naming the setting and the fix.
Tests: last_used mirroring (fresh+refresh+fallback), sidecar throwaway + error
isolation, macOS26 parser + detect, perms-on-refresh, lightpanda guard. 198
browser tests pass. Live: Profile-6 cookie DB lands in copy Default (file-level);
real Gmail (Default profile) still signed in.
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:
Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
(_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
(honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod
Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots
Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.
- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
real_profile_data_dir hard-wired Linux to $XDG_CONFIG_HOME/<name>, and the
xdg fragment map only knew the native package names. Ubuntu's default snap
Chromium (xdg reports chromium_chromium.desktop, profile under
~/snap/chromium/common/chromium) and Flatpak builds (~/.var/app/<id>/config/…)
therefore ended in 'profile directory was not found' for a browser the user
runs every day, and Flatpak Chrome (com.google.Chrome.desktop) was reported as
'not a supported Chromium browser'.
Try the native, snap and Flatpak locations and return the first that exists;
fall back to the native path so the error message still names a concrete
directory. Map the Flatpak application ids in the xdg lookup.
Tests cover the xdg names for all four browsers in native and Flatpak form,
and the directory preference order with a temp HOME.
_detect_default_darwin matched a Chromium bundle id and the literal 'https'
anywhere in the whole LSHandlers dump, so a browser registered for ftp or a
content type was reported as the https default, and map order decided ties.
When nothing matched it fell back to the first installed Chromium app — with
Safari or Firefox as the actual default that drove a browser the user never
consented to, contradicting the docstring, the config comment and the desktop
copy ('a non-Chromium default fails with a clear message').
Parse the dump entry by entry, take the LSHandlerRoleAll/Viewer of the entry
whose LSHandlerURLScheme is https, and fail closed on anything else — an empty
handler list is what macOS stores while Safari is still the implicit default.
Tests feed real 'defaults read' output shapes instead of patching the detector
(reviewer fixture from the PR discussion: Safari on https, Chrome on ftp).
The two default-browser detectors call subprocess.run(text=True) without an
explicit encoding, which the Windows-footgun linter (and its full-repo test,
tests/scripts/test_footgun_subprocess_encoding.py) rejects. Pass
encoding='utf-8', errors='replace' like the rest of the tree.
Two existing tests replace browser_navigate / _navigation_session_key with
positional-only lambdas; both callables now receive local_browser= from the
registry handler and browser_navigate, so the spies raised TypeError. Accept
the keyword with its default.
Fixes the three CI failures on the PR head (Windows footguns lint,
test_browser_extension_router_wiring x2, test_browser_open_timeout).
Fixes#93406 (residual). _fleet_probe_expected_runtimes counted the
_windows_gateway_resume pause/resume token (profiles/unmapped entries)
as an 'expected fleet rows' signal. The token is pause/resume
bookkeeping, not a runtime inventory, and its entries have no rows
collect_fleet_versions() can return: unmapped Scheduled-Task gateways
never publish gateway_state.json, and a resumed profile gateway
relaunches detached and may not republish within the probe window. So
every Windows update that paused a gateway set _fleet_rows_expected,
the verification loop silently waited out its polling window (~14 min
wall clock with the retry loop on user reports), printed 'Fleet version
check returned no rows', and exited 1 for an update that succeeded.
Expected-runtimes now keys only on row-capable signals: restart-phase
bookkeeping, the pre-restart PID snapshot, and the pre-update plan
inventory -- which already cover any genuinely live pre-update Windows
gateway.
Counterfactual proof: tests/hermes_cli/test_update_fleet_probe_resume_token.py
fails on the pre-fix predicate (token-only => True) and passes with the
fix; the row-capable signals are pinned unchanged.
Stdio MCP helper subprocesses (npx/binary servers) never import Hermes
code, so they could not self-register in the machine spawn ledger and an
unclean parent exit left them running invisibly forever.
- process_identity.register_child(pid, purpose): ledger mirror of
register_self for spawned children — records the CHILD (pid,
create_time) with this process as spawner. Refuses pid-only entries a
PID reuse could forge. Writes go through the single _append_entry
path under _LEDGER_LOCK (prune + atomic tmp/replace unchanged).
- 'mcp-helper' added to REAPABLE_PURPOSES so the updater's
_ledger_reapable_backend_pids rung flows helpers through its existing
spawner_is_dead gate (live spawner => never reaped).
- tools/mcp_tool.py: best-effort register_child(pid, 'mcp-helper') at
the post-spawn PID capture; never breaks MCP startup.
- reap_orphaned_mcp_helpers(): startup sweep mirroring
_reap_orphaned_desktop_local_serves but ledger-driven — kills only
helpers whose recorded spawner is PROVABLY dead, with a create_time
re-check at kill time. Wired next to the desktop serve reap in
web_server.py.
A failed update attempt can pull fresh code onto disk and then die before
the config-migration block (e.g. a PyPI timeout during the dependency
sync). The desktop hand-off retries; the retry takes the commit_count == 0
branch, repairs deps, prints 'Already up to date!' and returns early -
skipping _run_config_check_fresh / migrate_config entirely. The fresh
code (requiring a newer _config_version) then refuses to start against
the old config until 'hermes doctor --fix' is run.
Fix: _maybe_migrate_config_on_current() mirrors the version_bump_only
handling (silent, non-interactive) and is called on both repair-path
completion points before claiming success.
Also: scripts/desktop-update/posix.sh no longer retries when the update
was deliberately SKIPPED (checkout parked on a non-target branch) -, the
retry is deterministic and only wastes time. Uses a dedicated non-
colliding exit code (8) and an honest message instead of 'Update failed'.
New tests: tests/hermes_cli/test_update_config_migration_on_current.py
(5 cases: migrate-when-behind, noop-current, noop-ahead, warning re-
surface, silent check failure).
When an update was interrupted or failed mid-install (e.g. dependency install
timeout) after pulling new code, the subsequent update run takes the
'commit_count == 0' path and early-returned without checking or migrating
the configuration. Fresh code requiring a newer config version would fail to
boot on the next run.
Extract _check_and_apply_config_migration and invoke it across all update
completion paths (normal update, current checkout / node repair, and python
dependency repair).
Structural fix after five review rounds put four blockers in the same
subsystem. The root cause was representational: consent history is a
sequence of on/off intervals, but it was stored as ONE moving day-stamp
plus a revoked flag. Every fix had to mutate that scalar at exactly the
right moment from exactly the right place, and each round the mutation
was missing from some reachable path (write-once stamp in R3; recorded
inside a loop that never runs when sending is off in R4; dead code
whenever collection was off in R5).
Consent is now recorded as explicit intervals (send_consent_windows) and
eligibility is a pure derivation: a package is sent only when its whole
period falls inside a recorded window. One writer -
reconcile_send_consent - derives window state from an observation of
(config, now). It is idempotent and order-independent, so the wizard,
the relay, and the mid-pass check all call the same function and cannot
disagree; there are no edges to detect and no ordering between writers
to get wrong. The relay reconciles once per process BEFORE the
collection gate, which fixes round-5 D1 (enabled:false made the only
idle-path observer unreachable). The claim reads the table and never
writes it, removing the read-path mutation (D2's rewrite vector).
Timestamp discipline, each rule load-bearing and mutation-tested:
- 'obs' high-water mark: monotonic, advanced only by observations;
confirms an open window forward (last_confirmed_at).
- 'data' high-water mark: advanced only by stored package period_end;
clamps window OPENS so a rolled-back clock cannot slide a window
under refused packages already on disk (round-5 D2).
- A close stamps last_confirmed_at, never "now": consent is asserted
only for observed time, so a hand-edited config with no process
running for 90 days fails closed (round-5 D1 strongest form).
- The gate requires period containment, not period_start >=, so an
intra-day revoke/re-enable holds back the day package (round-5 D3).
- Unlike the day-stamp, a revoke/re-enable cycle no longer destroys the
undelivered backlog from the earlier consented window (round-5 D4).
The redesign was validated BEFORE implementation against all 13
reproduced defect scenarios on a real store; the first two drafts each
failed scenarios in that harness (v1 leaked the unobserved-gap case by
closing at "now"; v2 leaked refused windows by letting data stamps
confirm consent). The harness ships as
tests/hermes_cli/test_shared_metrics_consent_windows.py.
Deleted: OPT_IN_PERIOD_KEY, SEND_REVOKED_KEY, LAST_SEEN_SEND_KEY,
opt_in_period(), record_revoked(), the relay edge detector body, and the
setup wizard's key bookkeeping (~170 lines of transition machinery).
Schema: two additive tables, version deliberately unchanged; verified
against a copy of the real production DB (13 rows intact, reopen no-op).
Also kills round-5's M8 survivor: the seen-exclusion mutation now fails
the suite. New mutation sweep: 8/8 killed, including one vacuous test of
my own this round (obs-mark monotonicity was covered only by
coincidence of the data mark; now pinned directly).
Documented cost: a fresh package waits at most one process start after
its period completes before release (fail-closed direction).
270 tests pass; ruff and windows-footguns clean. Staging E2E re-run
through the interval gate: both packages 202.
Fourth independent review. Two more consent leaks, both reproduced through
the real relay entry point before and after the fix. Both are failures of
my own round-3 fix, which recorded revocation in the wrong place.
BLOCKER 1 - revoking while idle recorded nothing. _record_revocation lived
inside send_pending's loop, but _send_exported_packages returns early when
send is false, before a sender is ever constructed. The dominant case is a
user turning sending off while no pass is running, so the loop that was
meant to observe the revocation could never run. Reproduced: 6 periods
collected during a refused window were transmitted on re-enable.
The window now closes on the observed config EDGE, before the early return.
Last-seen send state is persisted because each hook fires in a fresh
process, so a true->false transition is only visible by comparison. The
rising edge also opens the window explicitly: the sender only runs when
there is something to send, so a user who opts in and out before any
package exists would otherwise have no window for record_revoked to close.
BLOCKER 2 - turning COLLECTION off never recorded revocation. The
not-enabled branch in setup.py force-set send=false and returned without
calling _record_send_consent_change, so `hermes tools` -> disable shared
metrics silently dropped consent while leaving the window open. Same
retroactive release on re-enable. Both consent surfaces now record, and
setup keeps the relay's edge detector in step.
Also, from the same review's mutation sweep:
- the scheme check is now pinned as an allowlist. Replacing the http test
with `if True` survived the entire suite, because every non-http case
targeted a REMOTE host where the loopback branch rejects anyway. Only a
non-http scheme on loopback distinguishes the two. Shipped behaviour was
already correct; nothing guarded it.
- A.3 no longer claims rotation bounds long-term linkability outright.
Measured against 11 real packages: resource is a stable low-entropy
tuple and periods are contiguous across a rotation, so for a RARE
configuration those can bridge windows. The honest claim is that
rotation raises the cost, not that it makes correlation impossible.
Two mutants are documented as unkillable rather than papered over with
tests that only appear to cover them: the _defer clamp is unreachable from
any current caller, and widening the falling-edge check to an
unconditional else is behaviourally equivalent because record_revoked is
idempotent and no-ops without an open window.
An earlier version of the anti-spurious-revocation test could not fail
either - it used a never-consented store, where record_revoked no-ops
regardless. Rewritten to opt in, revoke, re-enable, and then assert that a
steady enabled state does not re-close the reopened window.
259 tests pass. Staging E2E re-run: both packages 202.
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.
This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:
- gateway/status.py: expose service-ownership discovery for gateway
runtimes (find_windows_gateway_services maps validated gateway PIDs
through process ancestry to running SCM service PIDs, with
create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
venv mutation and restart them afterward. Stops wait for a stable
SCM 'stopped' state AND for the original descendant processes to
exit (service 'Stopped' is not proof the child released its
handles). Failure to prove ownership, stop a service, or restart it
fails closed; rollback restores attempted services, and rollback
failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
or a service that will not reach a stable state abort the update
before any file mutation.
Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.
Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Salvage adjustments to PR #94392 per review:
- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
recovery child now probes 'systemctl --user is-active' after each relaunch;
only an observed-active systemd unit is reported 'verified'. A relaunch that
merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
update_inventory serve collector) are no longer silently skipped: the
recovery pass records them (and manual gateways) as skipped-with-reason in
the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
(requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
fresh interpreter (sitecustomize shim intercepts the grandchild
'gateway restart' and systemctl probes).
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.
clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).
Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).
Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.
Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).
_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.
Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
Third independent review. Both blockers reproduced against a real store
before and after the fix.
BLOCKER 1 — head-of-line starvation. The claim query is LIMIT 1, and a
package already handled this pass was rejected AFTER the fetch, so
_claim_next returned None and send_pending read that as 'queue empty'.
Any row that sorts first and becomes eligible again mid-pass therefore
terminated the pass. This is reachable normally: a 429 with a short
Retry-After, or a pass outliving the 15-minute failure backoff (a legal
pass runs ~1900s). Measured: 10 of 19 healthy packages silently dropped.
The seen-set is now excluded IN SQL, so None genuinely means no eligible work.
Same scenario now delivers 19 of 19.
BLOCKER 2 — revoking consent leaked once it was re-granted. opt_in_period
was write-once, so packages collected while the user had send: false
still had period_start >= the ORIGINAL opt-in day; re-enabling released
the whole refused window. Reproduced: 5 packages from a 5-day opted-out
window transmitted on re-enable. Turning sending off now closes the
consent window, and the next enabled pass opens a new one from that day.
Recorded both in the setup wizard and in the sender itself, because
config.yaml can be hand-edited where the wizard never sees it.
Also: a send_attempts ceiling (a poisoned head row burned ~160 requests
over 30 days, unbounded), _defer clamps to >= 1s so it cannot write a
past deadline, and the dead skipped_not_due field is removed.
Test-quality fixes, since vacuous tests have been the recurring problem:
- the lease test asserted only 'in the future', passing for a 1s lease;
it now requires the lease to outlast one package's worst legal case
- test_shutdown_joins_the_send_thread grepped getsource for a method
name — a change-detector AGENTS.md rejects — and is now behavioural
- gzip determinism was unguarded: both retries in one pass compress in
the same second, so removing mtime=0 was caught by nothing. Now
compares output across a real second boundary.
All five new regressions are mutation-verified: reintroducing each bug
fails its test. The first attempt-ceiling test SURVIVED its mutation
(the seeded row was excluded by another predicate) and was rewritten to
drive the real loop.
251 tests pass. Staging E2E re-run: both packages 202.
The Electron shell boots by fetching / and extracting
window.__HERMES_SESSION_TOKEN__ to authenticate /api/ws
(dashboard-token.ts adoptServedDashboardToken). Headless serve 404'd
every path, so when the renderer's spawn token drifted from the
backend's live token — e.g. hermes update replaced the backend and the
env pin no longer matched — the renderer had no way to adopt the served
token, the WebSocket handshake failed, and the primary window
white-screened (#95575).
Serve a minimal token-only HTML page at the exact root path in
mount_spa()'s headless branch, matching the renderer's extraction regex.
Gate it on app.state.auth_required read at request time: a gated
(non-loopback / remote public_url) serve keeps returning the 404 JSON so
the session token never leaks past the loopback boundary. Every other
path stays 404 JSON — the SPA remains unserved.
Regression tests: TestHeadlessServeTokenPage (3 cases) — verified to
fail against the pre-fix headless branch.
hermes_cli/memory_setup.py::_write_env_vars() wrote provider-controlled
.env entries with a direct Path.write_text() + post-hoc chmod, bypassing
the denylist/regex/CRLF-stripping/atomic-replace validation that
hermes_cli/config.py::save_env_value() already provides for every other
.env writer in the codebase. A malicious or buggy memory-provider plugin
declaring a crafted env-var name/value in its setup schema could inject
arbitrary lines into .env.
Routes memory-provider env writes through save_env_value(), and fixes a
regression this surfaced in plugins/memory/supermemory/__init__.py::
post_setup(), which called the old two-parameter _write_env_vars(env_path,
values) signature — restores the caller via context-local
hermes_constants.set_hermes_home_override()/reset_hermes_home_override()
instead of a removed env_path parameter, so explicit HERMES_HOME overrides
during setup still resolve correctly.
Adds test_env_file_created_with_secure_permissions, guarded on Windows
(POSIX mode bits aren't enforced there, mirroring the existing skip in
test_openviking_provider.py / test_supermemory_provider.py) since
save_env_value's atomic-replace path creates the temp file at 0o600 before
writing content, closing the TOCTOU window the old direct-write + chmod
implementation had.
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".
Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.
A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.
Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
Cherry-picked core of #92545: the image build writes a versioned,
non-secret marker (/etc/hermes/image-provenance.json) outside both the
bind-mountable checkout and the HERMES_HOME volume, and
hermes_cli/image_provenance.py reads it fail-closed — absence means
'not image-managed', any present-but-malformed marker still means
image-managed (an integrity defect is never permission to mutate the
image in place).
(#91277 Phase 3; salvaged from #92545 by @andrexibiza — marker bake +
reader only, the scoped carve-out.)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.
- gateway/run.py: pause-for-update verb handler registered on the
existing control server; marshals onto the loop thread, ACKs with
{pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
no-answer (older gateway / no socket), so every caller keeps the
legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
per mapped profile gateway before the drain wait; positive ACKs extend
the wait to the gateway's own declared drain budget (+ teardown grace)
so a mid-turn gateway isn't force-killed at the end of a too-short
local default. Marker write + force-kill ladder retained verbatim as
the fallback.
Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.
Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.
Built on the spawn ledger (positive identity, never argv guessing):
- process_identity.py: LedgerEntry gains structured host/port/profile
(backward-compatible — readers .get()); register_self accepts detail=;
argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
manual backends inventory as supervisor=manual-serve with
restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
are stopped for the update and relaunched via an idempotent atexit
token built from structured identity (same contract as the gateway
pause/resume); receipts record serve_pause/serve_relaunch.
Desktop-owned backends keep the refusal (the app respawns what we
kill).
- dashboard_procs.py: the process scan is augmented with live ledger
rows, so profiled launches (`hermes --profile p serve ...`) that match
no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
— closing the #81564 status/stop asymmetry.
Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Slots below glm-5.3, above glm-5.2 in both curated lists; regenerates
model-catalog.json. No new metadata entries needed: context resolves via
the existing glm-5.3 fuzzy key (1,048,576 — matches OpenRouter live), and
both routes bill via official_models_api (live pricing).
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.
Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).
Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.
Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:
- restore_quick_snapshot's unlink+move fallback replaced the inode and
deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
the live writer's next checkpoint writes wrong-offset pages (the
page-1 compression_locks clobber from the report).
Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.