Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
De-flakes tests/hermes_cli/test_ssh_ownership_endpoint.py, which failed CI
twice on PR #95563 with teardown-time daemon-thread excepthook crashes — a
different test in the file each attempt, always green in isolation. Root
cause: three PROCESS-GLOBAL monkeypatches leaked into every other thread
sharing the per-file worker:
- monkeypatch.setattr(web_server.os, 'stat', ...) — web_server.os IS the os
module; any daemon thread from an earlier test that stat()ed during the
patch window got the fake 2-field stat and died in its excepthook, which
fired at interpreter teardown.
- monkeypatch.setattr('builtins.open', ...) — same class, worse blast radius.
- monkeypatch.setattr(web_server.sysconfig, 'get_paths', ...) — sysconfig is
process-global too.
Fixes, none of which weaken coverage:
- replaced-runtime test: a REAL tmp_path purelib whose recorded inode
deliberately mismatches (st_ino + 1) — real os.stat, same code path.
- readonly-purelib test: chmod 0o555 on the real directory instead of an
open() interceptor — exercises the genuine OSError branch (root-skipped,
where mode bits aren't enforced).
- sysconfig patches swapped for a SimpleNamespace on the web_server module
attribute — module-scoped, invisible to other threads.
Verified: 14 consecutive full-file runs green; sabotaging
_ssh_runtime_intact to always-True still fails 2 tests (coverage intact).
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.
Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).
Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.
Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:
- restore_quick_snapshot's unlink+move fallback replaced the inode and
deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
the live writer's next checkpoint writes wrong-offset pages (the
page-1 compression_locks clobber from the report).
Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.
Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.
The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.
Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.
The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.
copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.
A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.
Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.
Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
`restore_quick_snapshot` swaps only the main `.db` file: copy to a temp name,
`unlink()` the destination, `move()` the temp into place. The destination's
`state.db-wal` / `state.db-shm` are left untouched.
The snapshot is a checkpointed `sqlite3.backup()` image (`_safe_copy_db`) that
owns no WAL, so a leftover `-wal` describes the database that was just
unlinked. An ungracefully killed gateway (SIGKILL / OOM / power loss /
container stop) strands exactly those sidecars -- and that is precisely when an
operator reaches for a snapshot restore. SQLite then replays that foreign WAL
over the restored file on the next open.
The module already documents this hazard for the backup archive:
`_EXCLUDED_SUFFIXES` says shipping sidecars next to a `backup()` copy "would
pair a fresh snapshot with stale sidecar state and produce a torn restore on
the next open." The same reasoning was never applied to the restore
destination.
Reproduced against the real `create_quick_snapshot` / `restore_quick_snapshot`:
origin/main restore_returned=True sidecars_left=['-wal','-shm']
integrity=*** in database main *** Tree 2 page 295:
btreeInitPage() returns error code 11 rows=DatabaseError
patched restore_returned=True sidecars_left=none
integrity=ok rows=2000
The restore reports success and returns True while leaving the database
malformed. `SessionDB` then fails to open on every subsequent start, and
`repair_state_db_schema` only rebuilds `sqlite_master`/FTS -- the damage is in
the `sessions`/`messages` b-trees, so it re-raises. The pre-restore `state.db`
is already unlinked, and because the stale `-wal` survives the restore,
re-running it corrupts the file again identically: the documented last-resort
recovery is wedged.
Every other DB-move site in the repo already handles sidecars --
`hermes_state.py:742` and `:1698`, `hermes_cli/kanban_db.py:1759`,
`hermes_cli/session_recovery.py:60`. `restore_quick_snapshot` was the outlier.
The two `hermes update` auto-restore sites (`hermes_cli/main.py`) have the same
shape and run unattended: they `copy2` the snapshot over `state.db` with the
sidecars still present, then `verify_sqlite_integrity` fails and prints
"Auto-restore FAILED -- restored copy also failed integrity". Same fix.
Older builds persisted group members with a FRIENDLY name as the
descriptor's `name` (e.g. '大司命' for slug 'taiyi'), some predating
connection scoping entirely (no connectionId). Key matching alone seated
those descriptors as ghosts NEXT TO their own live rows ('4 bots' in a
2-bot room, reproduced live), and any path passing ghost identity onward
targeted a profile that does not exist on disk.
groupChatMemberBots now normalizes stored descriptors before seating:
an unmatched descriptor re-tries by case-drifted slug or friendly name
(botFriendlyNames precedence) against rows on its own connection —
connectionless pre-scoping descriptors match local rows only. The next
persistence pass rewrites storage to slugs, so the repair self-heals.
Unresolvable descriptors still seat as degraded ghosts and are never
used as profile targets.
The last piece of the macOS permissions campaign (#52010 follow-up): macOS
prompts per-folder (Desktop, then Downloads, then Documents, ...) as the
agent touches each one — a drip-feed of dialogs on first use. ONE Full Disk
Access grant covers all of them permanently, and with the stable signing
identities merged this week it survives every update. Nothing in Hermes
taught users that.
- hermes doctor: check_macos_full_disk_access() — prompt-free probe (the
FDA-gated TCC db dir returns EPERM without a dialog; TCC only prompts on
protected-CATEGORY paths), reports granted state or prints the one-switch
setup with the Privacy_AllFiles deep link.
- hermes setup: same probe at the end of onboarding — the moment users are
primed to do system setup — silent when already granted, indeterminate,
or non-macOS.
- docs: desktop.md TCC section now leads with the one-switch guidance.
- 7 tests (granted / denied / indeterminate / non-macOS, both surfaces).
Two fixups the #75785 review required before landing:
- grep fallback no longer uses --exclude-dir for protected dirs: grep
matches exclude-dir globs against BASENAMES anywhere in the tree, so
--exclude-dir=Downloads silently skipped every nested directory named
Downloads (a repo's own Downloads/ included). Protected-dir searches now
route through find's path-scoped -prune (same traversal-prevention the
find backend uses) feeding grep via -exec. Regression test proves a
nested work/repo/Downloads/notes.txt is still found while ~/Downloads is
not (live filesystem, real find+grep).
- exclusions gated on env.is_local (new BaseEnvironment flag, True on
LocalEnvironment): sys.platform/Path.home() describe the controller, not
the execution host — a macOS controller driving a Linux SSH/container
backend must not prune the remote's unprotected Downloads. Environments
without the flag default to local semantics (warning-carrying skip,
never data loss).
Both sabotage-verified: restoring basename --exclude-dir fails 2 tests.
Follow-up on the cherry-picked #82644: the (st_dev, st_ino) snapshot of
site-packages does not survive contact with ext4 — a recreated directory
routinely REUSES the freed inode, so the exact reported repro
(rm -rf venv && uv venv) passed the intact check undetected. Proven live
during salvage: the E2E's replaced venv came back with the identical
inode and runtimeIntact stayed true.
Primary identity is now a marker file written into site-packages when
the SSH owner nonce activates: it deterministically dies with the old
tree on ANY replacement (same or different Python version) and survives
in-place pip/uv installs (no false stales). The stat snapshot remains as
the fallback for read-only site-packages, where it still catches
cross-device moves and version-bump path changes. Client classifier
semantics unchanged: only an explicit runtimeIntact:false rejects, so
older remotes stay compatible.
Three new tests: recreated-venv-with-reused-inode (the live-proven
case), in-place-install stays intact, read-only fallback arms the stat
tier.
Folds review findings: surface failed_deletes in CLI and gateway
/rollback output (new gateway.rollback.failed_deletes locale key, 17
locales), emit skipped_oversize on the nothing-to-restore early return
too, document all three report keys in the restore() docstring, and pin
the failed_deletes contract from both sides in tests.
Follow-up to #95491. The restore result dict had two inconsistent
reporting surfaces: skipped_oversize was only present when non-empty
(unlike skipped_user_edits), and failed_deletes was filtered from
restored_files but never surfaced to the user at all (debug-level log
only). Both are the same silent-omission class #95491 fixed for
oversize files; this completes the cleanup.
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.
New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
async completion test)
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
describes the state at timeout; a recycling cleanup never poisons the
healed replacement)
- the async flavor offloads the mark off the event loop
(asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.
Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.
Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.
_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.
Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.
Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).
Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.
The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.
New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.
attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).
Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.
Tests:
- electron/hardening.test.ts pins both broadcast branches in
saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
(/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
'saved' push arrives, menu after — without reload — shows it.
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.
SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.
The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
Two review-thread deltas on the salvaged #94950 latch:
- Match the gateway's structured 4001 code, not a message substring, when
the rejection carries one (JsonRpcGatewayError). A coded error that
merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
failure) must not latch the guard and freeze the status stack on a
healthy session. The substring fallback survives only for codeless
legacy errors.
- Reset the latch on runtime re-mint, not only on status-stack rebind:
wire resetBackgroundPollingGuard() at both reconnect seams that already
drop stale runtime bindings (use-gateway-boot's post-reconnect
resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
dead runtime 4001'd resume polling once a respawned backend re-mints
them.
Tests: code-specific match both directions; full-reset resumes every
latched session.
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".
`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.
The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.
Distinguish the two failure classes:
- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
session may well still be alive. Misclassifying that direction would
silently freeze the status stack on a healthy session.
The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.
Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.
Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.
Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.
If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch latches
true forever since the try block's finally never runs.
Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.
The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.
Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".
- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.
Closes#82614
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.
Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.
Two fixes:
1. _check_via_local_git now detects fetch failure (returncode != 0 or
exception) and returns None instead of falling through to stale
refs. The caller treats None as 'check could not run' rather than
'up to date'.
2. check_for_updates no longer caches None results. Previously, a
None from a failed check was cached for 6 hours, suppressing
retries until the cache expired. Now only conclusive results
(0 or >=1) are cached, so the next check attempt runs immediately
on the next call.
Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
Follow-up to the salvaged #95207 fix, completing the misreport bug class:
- restore() now also drops delete_targets whose unlink failed (OSError
swallowed) from restored_files — the sibling of the kept-oversize
misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
(slash_commands + gateway.rollback.kept_oversize locale key in all 17
catalogs) now tells the user which files were kept because the size
cap excluded them from every checkpoint; previously the file was
correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.
The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.
Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.
No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.
Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.
Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.
Reproduced on main with shipped defaults:
corpus.jsonl (2 MB), agent appends to it, then /rollback 1
safe_restore_plan restore=['corpus.jsonl', 'notes.py']
restore ok=True restored_files=['corpus.jsonl', 'notes.py']
notes.py exists=True content="v1 = 'original source'"
corpus.jsonl exists=False <- deleted, was in no checkpoint
Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.
The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.
The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.
The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.
Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.
Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.