Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 8e1bc8d342 chore: remove stealth/ox-alpha from OpenRouter and Nous Portal model catalogs
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
2026-08-26 07:03:53 -07:00
Teknium 84b91a1dc5 test(ssh-ownership): remove process-global patches that crashed sibling threads
De-flakes tests/hermes_cli/test_ssh_ownership_endpoint.py, which failed CI
twice on PR #95563 with teardown-time daemon-thread excepthook crashes — a
different test in the file each attempt, always green in isolation. Root
cause: three PROCESS-GLOBAL monkeypatches leaked into every other thread
sharing the per-file worker:
- monkeypatch.setattr(web_server.os, 'stat', ...) — web_server.os IS the os
  module; any daemon thread from an earlier test that stat()ed during the
  patch window got the fake 2-field stat and died in its excepthook, which
  fired at interpreter teardown.
- monkeypatch.setattr('builtins.open', ...) — same class, worse blast radius.
- monkeypatch.setattr(web_server.sysconfig, 'get_paths', ...) — sysconfig is
  process-global too.

Fixes, none of which weaken coverage:
- replaced-runtime test: a REAL tmp_path purelib whose recorded inode
  deliberately mismatches (st_ino + 1) — real os.stat, same code path.
- readonly-purelib test: chmod 0o555 on the real directory instead of an
  open() interceptor — exercises the genuine OSError branch (root-skipped,
  where mode bits aren't enforced).
- sysconfig patches swapped for a SimpleNamespace on the web_server module
  attribute — module-scoped, invisible to other threads.

Verified: 14 consecutive full-file runs green; sabotaging
_ssh_runtime_intact to always-True still fails 2 tests (coverage intact).
2026-08-26 07:03:04 -07:00
Teknium 2f9e187001 revert(macos): remove the TCC interpreter anchor — anchored copies could not load libpython
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.

Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).

Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.

Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
2026-08-26 06:50:53 -07:00
Teknium 2b8b4542e9 docs: /snapshot restore live-safety behavior 2026-08-26 06:28:43 -07:00
Teknium 7125e839d3 fix(state): fail closed when a live process still holds state.db during destructive restore (#90950)
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:

- restore_quick_snapshot's unlink+move fallback replaced the inode and
  deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
  the live writer's next checkpoint writes wrong-offset pages (the
  page-1 compression_locks clobber from the report).

Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
2026-08-26 06:28:43 -07:00
briandevans 33dce0eb7e refactor(update): fold the auto-restore sequence into a shared helper
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.

Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.

The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.

Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
2026-08-26 06:28:43 -07:00
briandevans 86d719067b fix(update): clear stale SQLite sidecars before auto-restoring state.db
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.

The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.

copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.

A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.

Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.

Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
2026-08-26 06:28:43 -07:00
dsad a7988b830b fix(backup): clear stale SQLite sidecars on snapshot restore
`restore_quick_snapshot` swaps only the main `.db` file: copy to a temp name,
`unlink()` the destination, `move()` the temp into place. The destination's
`state.db-wal` / `state.db-shm` are left untouched.

The snapshot is a checkpointed `sqlite3.backup()` image (`_safe_copy_db`) that
owns no WAL, so a leftover `-wal` describes the database that was just
unlinked. An ungracefully killed gateway (SIGKILL / OOM / power loss /
container stop) strands exactly those sidecars -- and that is precisely when an
operator reaches for a snapshot restore. SQLite then replays that foreign WAL
over the restored file on the next open.

The module already documents this hazard for the backup archive:
`_EXCLUDED_SUFFIXES` says shipping sidecars next to a `backup()` copy "would
pair a fresh snapshot with stale sidecar state and produce a torn restore on
the next open." The same reasoning was never applied to the restore
destination.

Reproduced against the real `create_quick_snapshot` / `restore_quick_snapshot`:

  origin/main  restore_returned=True  sidecars_left=['-wal','-shm']
               integrity=*** in database main *** Tree 2 page 295:
               btreeInitPage() returns error code 11   rows=DatabaseError
  patched      restore_returned=True  sidecars_left=none
               integrity=ok   rows=2000

The restore reports success and returns True while leaving the database
malformed. `SessionDB` then fails to open on every subsequent start, and
`repair_state_db_schema` only rebuilds `sqlite_master`/FTS -- the damage is in
the `sessions`/`messages` b-trees, so it re-raises. The pre-restore `state.db`
is already unlinked, and because the stale `-wal` survives the restore,
re-running it corrupts the file again identically: the documented last-resort
recovery is wedged.

Every other DB-move site in the repo already handles sidecars --
`hermes_state.py:742` and `:1698`, `hermes_cli/kanban_db.py:1759`,
`hermes_cli/session_recovery.py:60`. `restore_quick_snapshot` was the outlier.

The two `hermes update` auto-restore sites (`hermes_cli/main.py`) have the same
shape and run unattended: they `copy2` the snapshot over `state.db` with the
sidecars still present, then `verify_sqlite_integrity` fails and prints
"Auto-restore FAILED -- restored copy also failed integrity". Same fix.
2026-08-26 06:28:43 -07:00
webtecnica 88d55f31e4 fix(backup): restore state.db through SQLite backup API so live connections see restored data (#65942) 2026-08-26 06:28:43 -07:00
Teknium bc21808e6a fix(desktop): legacy group members stored under display names seat their real bot once, not as ghosts (#92794)
Older builds persisted group members with a FRIENDLY name as the
descriptor's `name` (e.g. '大司命' for slug 'taiyi'), some predating
connection scoping entirely (no connectionId). Key matching alone seated
those descriptors as ghosts NEXT TO their own live rows ('4 bots' in a
2-bot room, reproduced live), and any path passing ghost identity onward
targeted a profile that does not exist on disk.

groupChatMemberBots now normalizes stored descriptors before seating:
an unmatched descriptor re-tries by case-drifted slug or friendly name
(botFriendlyNames precedence) against rows on its own connection —
connectionless pre-scoping descriptors match local rows only. The next
persistence pass rewrites storage to slugs, so the repair self-heals.
Unresolvable descriptors still seat as degraded ghosts and are never
used as profile targets.
2026-08-26 06:26:09 -07:00
Teknium be85903234 feat(macos): one-switch Full Disk Access guidance in doctor and setup
The last piece of the macOS permissions campaign (#52010 follow-up): macOS
prompts per-folder (Desktop, then Downloads, then Documents, ...) as the
agent touches each one — a drip-feed of dialogs on first use. ONE Full Disk
Access grant covers all of them permanently, and with the stable signing
identities merged this week it survives every update. Nothing in Hermes
taught users that.

- hermes doctor: check_macos_full_disk_access() — prompt-free probe (the
  FDA-gated TCC db dir returns EPERM without a dialog; TCC only prompts on
  protected-CATEGORY paths), reports granted state or prints the one-switch
  setup with the Privacy_AllFiles deep link.
- hermes setup: same probe at the end of onboarding — the moment users are
  primed to do system setup — silent when already granted, indeterminate,
  or non-macOS.
- docs: desktop.md TCC section now leads with the one-switch guidance.
- 7 tests (granted / denied / indeterminate / non-macOS, both surfaces).
2026-08-26 06:26:07 -07:00
Teknium 1145fcaed6 chore: map contributor email for deepeet-git 2026-08-26 06:26:02 -07:00
Teknium 979b7d14fd fix(search): path-scoped grep pruning + execution-backend gating for macOS TCC exclusions
Two fixups the #75785 review required before landing:
- grep fallback no longer uses --exclude-dir for protected dirs: grep
  matches exclude-dir globs against BASENAMES anywhere in the tree, so
  --exclude-dir=Downloads silently skipped every nested directory named
  Downloads (a repo's own Downloads/ included). Protected-dir searches now
  route through find's path-scoped -prune (same traversal-prevention the
  find backend uses) feeding grep via -exec. Regression test proves a
  nested work/repo/Downloads/notes.txt is still found while ~/Downloads is
  not (live filesystem, real find+grep).
- exclusions gated on env.is_local (new BaseEnvironment flag, True on
  LocalEnvironment): sys.platform/Path.home() describe the controller, not
  the execution host — a macOS controller driving a Linux SSH/container
  backend must not prune the remote's unprotected Downloads. Environments
  without the flag default to local semantics (warning-carrying skip,
  never data loss).

Both sabotage-verified: restoring basename --exclude-dir fails 2 tests.
2026-08-26 06:26:02 -07:00
takealook97 5fd6811dfe fix: avoid macOS privacy prompts during broad searches 2026-08-26 06:26:02 -07:00
Teknium cddb908aab fix(web_server): detect replaced venvs with a marker file — inode snapshots miss ext4 inode reuse
Follow-up on the cherry-picked #82644: the (st_dev, st_ino) snapshot of
site-packages does not survive contact with ext4 — a recreated directory
routinely REUSES the freed inode, so the exact reported repro
(rm -rf venv && uv venv) passed the intact check undetected. Proven live
during salvage: the E2E's replaced venv came back with the identical
inode and runtimeIntact stayed true.

Primary identity is now a marker file written into site-packages when
the SSH owner nonce activates: it deterministically dies with the old
tree on ANY replacement (same or different Python version) and survives
in-place pip/uv installs (no false stales). The stat snapshot remains as
the fallback for read-only site-packages, where it still catches
cross-device moves and version-bump path changes. Client classifier
semantics unchanged: only an explicit runtimeIntact:false rejects, so
older remotes stay compatible.

Three new tests: recreated-venv-with-reused-inode (the live-proven
case), in-place-install stays intact, read-only fallback arms the stat
tier.
2026-08-26 06:24:30 -07:00
toprakeker 8624c1e8f7 fix(desktop): reject SSH backends with replaced runtimes 2026-08-26 06:24:30 -07:00
kshitijk4poor d0351e3230 fix(checkpoints): display failed deletes to users and stabilize result keys
Folds review findings: surface failed_deletes in CLI and gateway
/rollback output (new gateway.rollback.failed_deletes locale key, 17
locales), emit skipped_oversize on the nothing-to-restore early return
too, document all three report keys in the restore() docstring, and pin
the failed_deletes contract from both sides in tests.
2026-08-26 18:02:17 +05:30
kshitijk4poor 37200847d1 fix(checkpoints): surface failed_deletes and make skipped_oversize unconditional
Follow-up to #95491. The restore result dict had two inconsistent
reporting surfaces: skipped_oversize was only present when non-empty
(unlike skipped_user_edits), and failed_deletes was filtered from
restored_files but never surfaced to the user at all (debug-level log
only). Both are the same silent-omission class #95491 fixed for
oversize files; this completes the cleanup.
2026-08-26 18:02:17 +05:30
kshitijk4poor dcfdc8deec fix(deadline): document the inline-mark contract; pin the ordering invariants (Phase 3a salvage round)
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.

New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
  invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
  async completion test)
2026-08-26 17:47:22 +05:30
Ayush Nangia 9ee2744097 fix(deadline): Phase 3a review round — mark ordering, loop offload, annotation
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
  describes the state at timeout; a recycling cleanup never poisons the
  healed replacement)
- the async flavor offloads the mark off the event loop
  (asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
2026-08-26 17:47:22 +05:30
Ayush Nangia 60b93eb521 test(deadline): Phase 3a poisoned-state coverage
- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
2026-08-26 17:47:22 +05:30
Ayush Nangia 7a3aaf0143 feat(deadline): SuspectableBackend protocol — mark timed-out backends suspect
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
2026-08-26 17:47:22 +05:30
hermes-seaeye[bot] 86ae906e88 fmt(js): npm run fix on merge (#95511)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-26 11:57:06 +00:00
miha 1bd5da3ac6 fix(desktop): skip macOS TCC-protected media dirs in git repo scan
The sidebar's home-dir repo crawl descends into ~/Pictures, ~/Music,
~/Movies and ~/Public — and into Photos/Music library packages — which
triggers Photos, Media Library and Files & Folders permission prompts
attributed to Hermes.app. Because the app is ad-hoc signed and re-signed
on every self-update (#49110), macOS drops all TCC grants after each
update, so these prompts re-fire every time.

Skip the media folders as direct children of a search root (nested dirs
like ~/dev/Music are ordinary and still scanned; an explicitly passed
root is still walked), and skip Apple library packages
(*.photoslibrary/*.musiclibrary/*.tvlibrary/*.aplibrary) at any depth.

Partial mitigation for #49110 / #52010: removes the Photos and Media
Library prompts entirely; the identity reset itself needs Developer ID
signed release artifacts (tracked in #49110).
2026-08-26 04:51:48 -07:00
Teknium 279726cc2f chore: map contributor email for wiseconnex 2026-08-26 04:51:41 -07:00
Dimar Anez 9eb13d07b6 fix(terminal): tolerate macOS TCC PermissionError in _safe_getcwd
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.

_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.

Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.

Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
2026-08-26 04:51:41 -07:00
Teknium bd134d0f30 test: loosen frozen bare-verdict dict in cua_0_9 sibling test to decision contract
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
2026-08-26 04:50:13 -07:00
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium 65605d4a7a fix(computer_use): pitch background-FIRST (not background-only) in schema, prompt block, and skill 2026-08-26 04:50:13 -07:00
Teknium bb3421bf25 fix(desktop): single-owner backend dial claim in Electron main (#90812)
reconnectGateway()'s in-flight lock lives at renderer module scope, so it
only dedupes reconnects inside ONE window. Two windows racing the same
wake both invoke the main-process backend ensure IPC, and for a pooled
SSH connection the loser of the pool-entry race could bootstrap a
duplicate remote backend (two tunnels, two remote serve processes).

Electron main is the single owner of backend lifecycles, so the claim
now lives there: BackendDialClaims keys in-flight dials by the pool
scope key from backendScopeKey(connectionId, profile) — the composite
identity seam wave-1 #93189 established for effective-identity reuse.
'hermes:connection' and 'hermes:connection:for' route through
backendDialClaims.run(), so concurrent renderer dials for one scope
coalesce onto one spawn and the second caller receives the first's
result. A claim exists only while its dial promise is unsettled: both
outcomes release it, a failed dial is never cached (fail closed, not
latched), and a synchronously-throwing dial rejects the claim instead
of escaping the seam.

The #93910 resume rebuild re-dials retired pool keys through the same
claim (redialPoolBackendAfterResume + new parseBackendScopeKey), so a
resume-driven rebuild and a concurrent renderer reconnect also coalesce
instead of racing.
2026-08-26 04:49:56 -07:00
Teknium 3123624c07 fix(desktop): revalidate pooled remote/SSH backends on power resume (#93910)
After macOS sleep/resume, pooled remote SSH descriptors kept serving dead
tunnels: a remote entry has no child 'exit' to clear it, the renderer
keepalive spares it from the idle reaper, and the wake-path nudges from
b90289b04/febed060a only re-drive the PRIMARY renderer socket. The
background failure-streak policy needs several probe rounds before it
drops a descriptor, so the Bots pane showed 'Gateway offline' long after
the network was back.

New revalidateSuspectPooledRemoteBackends(): on resume every pooled
remote is suspect — probe each once (bounded by
REMOTE_LIVENESS_TIMEOUT_MS), retire the dead ones immediately (pool
entry + SSH bootstrap + tunnel/master teardown) and rebuild them through
the caller's dial path, while healthy descriptors are left untouched. A
failed retire skips the rebuild (never dial on top of an installed
descriptor); a failed rebuild is logged and left to the renderer's
normal reconnect — the sweep never throws.

attachPowerResumeRemoteRevalidation() wires the sweep to the Electron
powerMonitor 'resume'/'unlock-screen' seam with a 15s holdoff so the
near-simultaneous macOS wake signals coalesce into one sweep and can
never form a hot loop; overlapping kicks additionally join the one
in-flight sweep via the existing RemoteRevalidationCoordinator.
2026-08-26 04:49:56 -07:00
Teknium 600d5166f0 test(gateway): prove delivered rows are never reclaimed by the reconnect sweep 2026-08-26 04:49:39 -07:00
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
Teknium b455abe0b3 fix(desktop): poll-guard reset is fire-and-forget off the redial path
composer-status imports $gateway from this module (cycle forces the
dynamic import), and awaiting the module load inside openSecondary sat on
the timed redial path — under CI load that pushed cold-start redials past
waitFor budgets in the lifecycle suite. The reset needs no ordering
guarantee relative to the dial; detach it.
2026-08-26 04:49:22 -07:00
Teknium 62534e2b5a fix(desktop): isolate the poll-guard reset import + sort-imports lint
The combined dynamic import meant a failed composer-status import (mocked
test graphs) silently skipped resetTileRuntimeBindings too — the exact
lifecycle regression CI caught. Separate best-effort trys per module.
2026-08-26 04:49:22 -07:00
Teknium 6bbae974d5 chore: map justinjohnson25600 and BrunoBza contributor emails 2026-08-26 04:49:22 -07:00
Teknium fe615a0099 fix(desktop): republish the connections registry to renderers after every successful save (#95393)
Live-confirmed on the Phase B build: hermesDesktop.connections.save()
succeeds and the registry on disk gains the row, but the switcher menu
(fed by the renderer $connectionsRegistry snapshot) keeps painting the
stale list until reload. remove() already broadcasts
hermes:connections:changed; save() only did so on the dial-material-edit
branch, so a brand-new connection or a label rename never reached the
switcher's onChanged re-pull (or any other window).

Fix at the publish seam only: saveRegistryConnection now broadcasts a new
'saved' reason for every successful save that isn't a dial-material edit.
'saved' is a pure registry-refresh signal — the use-gateway-boot listener
explicitly ignores it (nothing moved, so no dispose/redial/forget), while
the switcher's existing onChanged listener re-pulls the snapshot.

Tests:
- electron/hardening.test.ts pins both broadcast branches in
  saveRegistryConnection (source-assertion pattern; main.ts has no exports).
- connection-switcher.test.tsx mirrors the live repro scenario
  (/tmp/mg-ab/w2_95393.py): menu before save lacks the row, Electron's
  'saved' push arrives, menu after — without reload — shows it.
2026-08-26 04:49:22 -07:00
Bruno Bza 06be6cffbc fix(desktop): release reconnect-orphaned warm transcripts once their authoritative state settles
A gateway connection that dies mid-turn leaves cached session snapshots
whose busy/awaitingResponse flags can never settle: the respawned
backend re-mints runtime ids, so no terminal publish ever reaches the
orphaned snapshot again. #isWarmSettled treated those frozen flags as
live work, so every orphan pinned its full warm transcript until app
restart — roughly 5MB per reconnect cycle, which turned the restart
loop in #95189 into renderer OOM.

SessionStateCache now accepts an optional isAuthoritativelyActive
probe. When wired, in-flight flags only block eviction while the
authoritative $sessionStates record still claims work for the same
runtime id; without the probe the legacy always-block behavior is
preserved byte-for-byte. Eviction remains gated on needsInput, pending
drafts, and active references, so a genuinely running turn (which
re-asserts busy on every publish) is never a casualty.

The useSessionStateCache hook wires the probe to the store it already
imports. Reconnect reconciliation (reconcileBusyStatesOnReconnect)
settles the authoritative record, and the next prune drains the
orphaned cache entry through the normal LRU path, ownership included.
2026-08-26 04:49:22 -07:00
Teknium a7ea156470 fix(desktop): harden the dead-session poll guard per #94950 review
Two review-thread deltas on the salvaged #94950 latch:

- Match the gateway's structured 4001 code, not a message substring, when
  the rejection carries one (JsonRpcGatewayError). A coded error that
  merely mentions 'session not found' in wrapped text (e.g. a 5007 tool
  failure) must not latch the guard and freeze the status stack on a
  healthy session. The substring fallback survives only for codeless
  legacy errors.

- Reset the latch on runtime re-mint, not only on status-stack rebind:
  wire resetBackgroundPollingGuard() at both reconnect seams that already
  drop stale runtime bindings (use-gateway-boot's post-reconnect
  resetTileRuntimeBindings and gateway.ts's reopening path), so ids the
  dead runtime 4001'd resume polling once a respawned backend re-mints
  them.

Tests: code-specific match both directions; full-reset resumes every
latched session.
2026-08-26 04:49:22 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
Teknium 574bd7175c fix(desktop): unify boot-class getConnection() budgets on one shared 45s constant
Follow-up to the #95039 salvage: the cherry-picked bound used the 20s
RECONNECT_ATTEMPT_TIMEOUT_MS on boot()/softSwitch() getConnection(), but a
reviewer note (and the Phase A registry-restore work) established that
boot-class awaits must ride out a full backend cold spawn — main's spawn
budget is 45s (DEFAULT_BACKEND_READY_TIMEOUT_MS). A 20s renderer bound would
latch boot errors on healthy-but-slow cold boots.

Introduce BACKEND_BOOT_WAIT_TIMEOUT_MS (45s) in lib/with-timeout.ts as the
single shared boot-class budget, point boot()/softSwitch() getConnection()
and connections.ts BOOT_DESCRIPTOR_WAIT_TIMEOUT_MS at it, and keep the 20s
reconnect budget only for reconnect-class awaits against an already-spawned
backend. No magic-number drift: 45_000 now appears once in renderer code.
2026-08-26 04:49:22 -07:00
nftpoetrist 31f3de1f06 fix(desktop): bound getConnection() on the boot and soft-switch paths (#93454)
resolveGatewayWsUrl() in attemptReconnect() because a wedged IPC
round-trip into the main process (e.g. a stuck revalidation after a
liveness-probe trip) can hang these awaits forever. A later fix
(e8d5660bae) extended the bound to resolveGatewayWsUrl() in boot() and
softSwitch() too, but left the getConnection() call immediately above
it in both functions unbounded.

If that call wedges: during initial boot the 'Starting Hermes...'
screen never resolves (bootCompleted never flips, nothing hits catch),
and during a soft gateway/profile switch  latches
true forever since the try block's finally never runs.

Wrap both with the same withTimeout()/RECONNECT_ATTEMPT_TIMEOUT_MS
pattern already used for the sibling calls.
2026-08-26 04:49:22 -07:00
Teknium 20d33e385a fix(web_server): a dashboard started without a build recovers the moment one appears (#82614)
mount_spa's WEB_DIST.exists() check ran ONCE at mount time: a long-lived
'hermes dashboard --skip-build' that survived a git pull (or launched
before the first build) installed a permanent no_frontend catch-all and
answered 404 'Frontend not built' on every route forever — even after
npm run build completed. Remote Desktop clients saw ERR_EMPTY_RESPONSE.

The missing-dist branch is now reserved for the headless-serve contract
only. The SPA routes mount unconditionally and already cope with a
missing dist per-request (_serve_index returns the same 404 JSON when
index.html is unreadable; the /assets mount gains check_dir=False so
StaticFiles 404s instead of raising at mount). The dashboard recovers
the moment a build lands on disk — no restart needed.

Direction from #82666 by @codexbt (his PR's rebase dropped the product
hunk, leaving only the test; the test is cherry-picked as-is and this
commit restores the behavior it pins, adapted to the current mount_spa
shape: headless guard preserved, per-request recovery instead of a
per-request exists() check).
2026-08-26 04:48:58 -07:00
codexbt 3dea11d703 fix(web_server): recheck WEB_DIST existence dynamically in mount_spa
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".

- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.

Closes #82614
2026-08-26 04:48:58 -07:00
Shakti Prasad Mohapatra 98e87ac886 fix(cli): preserve stale positive behind-count on fetch failure (#92578)
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.

Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
2026-08-26 04:17:39 -07:00
Shakti Prasad Mohapatra 55d50d5c92 fix(cli): don't serve stale update-check results after fetch failure (#82166)
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.

Two fixes:

1. _check_via_local_git now detects fetch failure (returncode != 0 or
   exception) and returns None instead of falling through to stale
   refs. The caller treats None as 'check could not run' rather than
   'up to date'.

2. check_for_updates no longer caches None results. Previously, a
   None from a failed check was cached for 6 hours, suppressing
   retries until the cache expired. Now only conclusive results
   (0 or >=1) are cached, so the next check attempt runs immediately
   on the next call.

Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
2026-08-26 04:17:39 -07:00
kshitijk4poor d62a05e94c fix(checkpoints): surface skipped_oversize to users and stop misreporting failed deletes as restored
Follow-up to the salvaged #95207 fix, completing the misreport bug class:

- restore() now also drops delete_targets whose unlink failed (OSError
  swallowed) from restored_files — the sibling of the kept-oversize
  misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
  (slash_commands + gateway.rollback.kept_oversize locale key in all 17
  catalogs) now tells the user which files were kept because the size
  cap excluded them from every checkpoint; previously the file was
  correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
2026-08-26 16:44:43 +05:30
RickyYii 595b5ce68a refactor(checkpoints): call the size-cap predicate instead of restating it
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.

The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.

Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.

No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.

Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
2026-08-26 16:44:43 +05:30
RickyYii d28bf79927 fix(checkpoints): stop safe restore deleting files the size cap excluded
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.

Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.

Reproduced on main with shipped defaults:

    corpus.jsonl (2 MB), agent appends to it, then /rollback 1
    safe_restore_plan restore=['corpus.jsonl', 'notes.py']
    restore ok=True restored_files=['corpus.jsonl', 'notes.py']
    notes.py     exists=True   content="v1 = 'original source'"
    corpus.jsonl exists=False  <- deleted, was in no checkpoint

Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.

The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.

The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.

The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.

Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.

Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
2026-08-26 16:44:43 +05:30
Teknium 7a10d91b29 chore: map contributor email for notkisk 2026-08-26 04:14:16 -07:00