Commit Graph

14240 Commits

Author SHA1 Message Date
Teknium 03537d69dc feat(gateway): updaters pause gateways over the control socket instead of tree-killing them (#92091 step 2)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.

- gateway/run.py: pause-for-update verb handler registered on the
  existing control server; marshals onto the loop thread, ACKs with
  {pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
  no-answer (older gateway / no socket), so every caller keeps the
  legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
  per mapped profile gateway before the drain wait; positive ACKs extend
  the wait to the gateway's own declared drain budget (+ teardown grace)
  so a mid-turn gateway isn't force-killed at the end of a too-short
  local default. Marker write + force-kill ladder retained verbatim as
  the fallback.

Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
2026-08-26 09:59:17 -07:00
unsupportedpastels c4f376c19a fix(config): block generic Copilot ACP controls 2026-08-26 09:54:31 -07:00
unsupportedpastels 5425ba14f2 fix(config): harden MCP env policy on Windows 2026-08-26 09:54:31 -07:00
unsupportedpastels 08cf4fea5d fix(mcp): restrict catalog environment writes 2026-08-26 09:54:31 -07:00
pefontana b2c493c173 Assert the skip branch ran in the no-quiesce watcher test
The three existing assertions are absence checks, so the test also
passed when the loop got no iteration inside the sleep window (with
interval=5.0 it passes without the gate ever executing). Checking
_scale_to_zero_no_suspend_logged proves the branch was taken.
2026-08-26 13:02:04 -03:00
pefontana 85816595d2 Merge remote-tracking branch 'origin/main' into fix/scale-to-zero-no-pointless-quiesce 2026-08-26 13:02:03 -03:00
Teknium 30749ed9dc test(tools): guard _ever_connected set in reconnect regression mock so it bites on pre-fix code (#94671 hardening) 2026-08-26 08:40:07 -07:00
chelsealong 8f517f5ca6 chore: address AI-review nits on _ever_connected fix
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
2026-08-26 08:40:07 -07:00
chelsealong e8dc0af5b1 fix(tests): set _ever_connected in reconnect-scenario test mocks
These pre-existing tests fake a successful first connect by calling
only _ready.set(), which is what the real code did before this PR.
Now that run() gates the initial-vs-reconnect branch on the new sticky
_ever_connected flag instead, their later simulated reconnect failures
were misclassified as never-connected and hit the 3-attempt ladder,
failing test_reconnect_counter_resets_after_successful_session,
test_parked_server_self_probes_and_revives, and
test_retry_attempts_log_debug_transitions_warn in CI. Set the flag
alongside _ready.set() to mirror the real success sites, same as the
new test added in tools/mcp_tool.py's own PR.
2026-08-26 08:40:07 -07:00
chelsealong c7673f322b fix(tools): stop treating a post-registration reconnect drop as an initial-connect failure
MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.

Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.

Fixes #94654
2026-08-26 08:40:07 -07:00
liuhao1024 57309c0cbb fix(update): never respawn backends from a foreign HERMES_HOME (#94030)
The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.

Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
2026-08-26 08:39:04 -07:00
fangliquan 858916acc4 fix(update): preserve SSH ownership only during updates 2026-08-26 08:39:04 -07:00
fangliquan 1676c614b3 fix(update): preserve SSH-owned backends during cleanup 2026-08-26 08:39:04 -07:00
Teknium 306a096b4a test: re-pin --status output to the serve-inclusive contract (#81564)
The old assertions pinned the phrasing that HID serve backends — the
exact asymmetry #81564 reports. Re-pinned to the new message and
strengthened: a serve-mode row must now appear, tagged [serve].
2026-08-26 07:57:04 -07:00
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
beplee 19d8b87234 test(desktop): harden rebind assertions per #94417 review
Enough1122 review points on #94417:
1. Precedence hazard fixed: the busy-guard assertion now locates the
   rebind helper body precisely and asserts the guard INSIDE it, instead
   of a 2000-char window with an (m and X) or Y precedence trap.
2. stored_session_id guarantee: documented + pinned — the gateway always
   stamps it ('stored_session_id': session_key or "" in server.py), and
   the rebind's typeof check refuses non-string/empty values, so an
   unnamed rebuilt runtime is never adopted as lineage proof.
3. New third assertion pins that refusal contract.

Structural smoke tests remain structural by design; the behavior
contract for the rebind is exercised end-to-end by the model-switch
manual repro path — a vitest harness driving handleSessionInfoEvent is
the follow-up candidate noted in the reply.
2026-08-26 07:28:10 -07:00
beplee ec8ca8f2cb fix(desktop): re-bind open pane to rebuilt runtime after model switch
A mid-conversation model/provider switch rebuilds the agent runtime. The
rebuilt runtime emits session.info (and all later events) under a NEW
explicit session_id while the pane still holds the dead one as its
active id — isActiveEvent is false for the same conversation from that
moment on, so view-scoped updates stop and the chat freezes until a
full resume (#93942 scenario B; backend even logs 'client should resume
the stored session', but the client never does).

Fix: when a session.info event lineage-matches the selected conversation
(sessionMatchesStoredId over stored_session_id) but carries a different
runtime id, adopt the new runtime id as the active session id — keeping
the durable selection untouched — so every subsequent isActiveEvent gate
keeps matching without a resume. Guarded: the old runtime must show no
live turn (not busy/awaiting/streaming) or the adoption is refused, so
an overlapping manual switch can never split one conversation across
two panes.

The existing compression-rotation path does not cover this case: it
fires when the SAME runtime's stored id rotates, while a rebuild
produces a NEW runtime with a NEW stored id.

Together with #94255 (tile reconcile on sessions.changed), closes
#93942.

Regression tests verified failing pre-fix on 41447a6d70.
2026-08-26 07:28:10 -07:00
beplee db8ff4eb75 fix(desktop): reconcile workspace-tile transcripts on sessions.changed
Bot canonical chats open as workspace tiles (workspaceMode: 'bots') and
are deliberately hidden from $sessions/$messagingSessions, so the
sessions.changed transcript refresh skipped them twice over: it covers
only the main pane's selection, and its resolveSession() bails on hidden
sessions. A background delivery (bot-to-bot DM via bot_relay.deliver, a
cron run's output, another machine) therefore never reached an open bot
chat — the roster updated but the pane stayed stale until remount
(#93942 scenario A).

Fix: the sessions.changed tick now also reconciles every visible
workspace tile through a dedicated signature-gated path. Each tile
carries its own stored↔runtime id pair so no resolution step is needed;
per-tile signatures make no-change ticks free; busy tiles are skipped
(their own stream owns the view); closed/superseded tiles discard their
in-flight read.

Slice 1 of 2 for #93942 (scenario A only). Scenario B (stream re-key
after mid-conversation model switch) follows separately.

Fixes part of #93942
2026-08-26 07:28:10 -07:00
Teknium f4df86fe1a test: adapt summary-continuity + rotation-flush fixtures to the lean default
Continuity tests pin tail_mode=legacy (they assert the raw LLM text
terminates the stored summary; lean's verbatim-user appendix follows it by
design and the contract under test is mode-independent). The #57491
rotation fixture grows 200→2000 chars/message: at ~2.5K total tokens the
old fixture fit entirely inside lean's 10K tail floor, so the no-growth
guard correctly refused the rotation the test exercises.
2026-08-26 07:16:04 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
Teknium 84b91a1dc5 test(ssh-ownership): remove process-global patches that crashed sibling threads
De-flakes tests/hermes_cli/test_ssh_ownership_endpoint.py, which failed CI
twice on PR #95563 with teardown-time daemon-thread excepthook crashes — a
different test in the file each attempt, always green in isolation. Root
cause: three PROCESS-GLOBAL monkeypatches leaked into every other thread
sharing the per-file worker:
- monkeypatch.setattr(web_server.os, 'stat', ...) — web_server.os IS the os
  module; any daemon thread from an earlier test that stat()ed during the
  patch window got the fake 2-field stat and died in its excepthook, which
  fired at interpreter teardown.
- monkeypatch.setattr('builtins.open', ...) — same class, worse blast radius.
- monkeypatch.setattr(web_server.sysconfig, 'get_paths', ...) — sysconfig is
  process-global too.

Fixes, none of which weaken coverage:
- replaced-runtime test: a REAL tmp_path purelib whose recorded inode
  deliberately mismatches (st_ino + 1) — real os.stat, same code path.
- readonly-purelib test: chmod 0o555 on the real directory instead of an
  open() interceptor — exercises the genuine OSError branch (root-skipped,
  where mode bits aren't enforced).
- sysconfig patches swapped for a SimpleNamespace on the web_server module
  attribute — module-scoped, invisible to other threads.

Verified: 14 consecutive full-file runs green; sabotaging
_ssh_runtime_intact to always-True still fails 2 tests (coverage intact).
2026-08-26 07:03:04 -07:00
Teknium 2f9e187001 revert(macos): remove the TCC interpreter anchor — anchored copies could not load libpython
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.

Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).

Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.

Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
2026-08-26 06:50:53 -07:00
briandevans 33dce0eb7e refactor(update): fold the auto-restore sequence into a shared helper
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.

Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.

The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.

Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
2026-08-26 06:28:43 -07:00
briandevans 86d719067b fix(update): clear stale SQLite sidecars before auto-restoring state.db
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.

The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.

copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.

A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.

Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.

Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
2026-08-26 06:28:43 -07:00
webtecnica 88d55f31e4 fix(backup): restore state.db through SQLite backup API so live connections see restored data (#65942) 2026-08-26 06:28:43 -07:00
Teknium be85903234 feat(macos): one-switch Full Disk Access guidance in doctor and setup
The last piece of the macOS permissions campaign (#52010 follow-up): macOS
prompts per-folder (Desktop, then Downloads, then Documents, ...) as the
agent touches each one — a drip-feed of dialogs on first use. ONE Full Disk
Access grant covers all of them permanently, and with the stable signing
identities merged this week it survives every update. Nothing in Hermes
taught users that.

- hermes doctor: check_macos_full_disk_access() — prompt-free probe (the
  FDA-gated TCC db dir returns EPERM without a dialog; TCC only prompts on
  protected-CATEGORY paths), reports granted state or prints the one-switch
  setup with the Privacy_AllFiles deep link.
- hermes setup: same probe at the end of onboarding — the moment users are
  primed to do system setup — silent when already granted, indeterminate,
  or non-macOS.
- docs: desktop.md TCC section now leads with the one-switch guidance.
- 7 tests (granted / denied / indeterminate / non-macOS, both surfaces).
2026-08-26 06:26:07 -07:00
Teknium 979b7d14fd fix(search): path-scoped grep pruning + execution-backend gating for macOS TCC exclusions
Two fixups the #75785 review required before landing:
- grep fallback no longer uses --exclude-dir for protected dirs: grep
  matches exclude-dir globs against BASENAMES anywhere in the tree, so
  --exclude-dir=Downloads silently skipped every nested directory named
  Downloads (a repo's own Downloads/ included). Protected-dir searches now
  route through find's path-scoped -prune (same traversal-prevention the
  find backend uses) feeding grep via -exec. Regression test proves a
  nested work/repo/Downloads/notes.txt is still found while ~/Downloads is
  not (live filesystem, real find+grep).
- exclusions gated on env.is_local (new BaseEnvironment flag, True on
  LocalEnvironment): sys.platform/Path.home() describe the controller, not
  the execution host — a macOS controller driving a Linux SSH/container
  backend must not prune the remote's unprotected Downloads. Environments
  without the flag default to local semantics (warning-carrying skip,
  never data loss).

Both sabotage-verified: restoring basename --exclude-dir fails 2 tests.
2026-08-26 06:26:02 -07:00
takealook97 5fd6811dfe fix: avoid macOS privacy prompts during broad searches 2026-08-26 06:26:02 -07:00
Teknium cddb908aab fix(web_server): detect replaced venvs with a marker file — inode snapshots miss ext4 inode reuse
Follow-up on the cherry-picked #82644: the (st_dev, st_ino) snapshot of
site-packages does not survive contact with ext4 — a recreated directory
routinely REUSES the freed inode, so the exact reported repro
(rm -rf venv && uv venv) passed the intact check undetected. Proven live
during salvage: the E2E's replaced venv came back with the identical
inode and runtimeIntact stayed true.

Primary identity is now a marker file written into site-packages when
the SSH owner nonce activates: it deterministically dies with the old
tree on ANY replacement (same or different Python version) and survives
in-place pip/uv installs (no false stales). The stat snapshot remains as
the fallback for read-only site-packages, where it still catches
cross-device moves and version-bump path changes. Client classifier
semantics unchanged: only an explicit runtimeIntact:false rejects, so
older remotes stay compatible.

Three new tests: recreated-venv-with-reused-inode (the live-proven
case), in-place-install stays intact, read-only fallback arms the stat
tier.
2026-08-26 06:24:30 -07:00
toprakeker 8624c1e8f7 fix(desktop): reject SSH backends with replaced runtimes 2026-08-26 06:24:30 -07:00
kshitijk4poor d0351e3230 fix(checkpoints): display failed deletes to users and stabilize result keys
Folds review findings: surface failed_deletes in CLI and gateway
/rollback output (new gateway.rollback.failed_deletes locale key, 17
locales), emit skipped_oversize on the nothing-to-restore early return
too, document all three report keys in the restore() docstring, and pin
the failed_deletes contract from both sides in tests.
2026-08-26 18:02:17 +05:30
kshitijk4poor 37200847d1 fix(checkpoints): surface failed_deletes and make skipped_oversize unconditional
Follow-up to #95491. The restore result dict had two inconsistent
reporting surfaces: skipped_oversize was only present when non-empty
(unlike skipped_user_edits), and failed_deletes was filtered from
restored_files but never surfaced to the user at all (debug-level log
only). Both are the same silent-omission class #95491 fixed for
oversize files; this completes the cleanup.
2026-08-26 18:02:17 +05:30
kshitijk4poor dcfdc8deec fix(deadline): document the inline-mark contract; pin the ordering invariants (Phase 3a salvage round)
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.

New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
  invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
  async completion test)
2026-08-26 17:47:22 +05:30
Ayush Nangia 60b93eb521 test(deadline): Phase 3a poisoned-state coverage
- async timeout marks once with a label-carrying rounded-timeout reason
- completion never marks
- sync flavor marks on timeout
- non-adopting backends keep the real timeout result
- a raising mark_suspect cannot eat the timeout or the label
2026-08-26 17:47:22 +05:30
Dimar Anez 9eb13d07b6 fix(terminal): tolerate macOS TCC PermissionError in _safe_getcwd
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.

_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.

Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.

Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
2026-08-26 04:51:41 -07:00
Teknium bd134d0f30 test: loosen frozen bare-verdict dict in cua_0_9 sibling test to decision contract
The verify_fresh_state verdict now carries an optional human hint; assert
the decision + additive-field absence instead of the exact dict shape
(same contract loosening as test_computer_use_delivery_ladder.py).
2026-08-26 04:50:13 -07:00
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium 600d5166f0 test(gateway): prove delivered rows are never reclaimed by the reconnect sweep 2026-08-26 04:49:39 -07:00
milnerrad 8e1db41041 fix(gateway): redeliver transient failures after reconnect 2026-08-26 04:49:39 -07:00
Justin Johnson c19849cd02 fix(desktop): stop the status-stack poll storming a dead session with 4001s
The composer status stack polls `process.list` every 5s while a background
process row is on screen. `process.list` is session-scoped, so against a
runtime id the gateway no longer holds it returns 4001 "session not found".

`refreshBackgroundProcesses` swallowed *every* failure with a bare `catch {}`
commented "transient socket loss". A gone session is not transient: the poll
re-sent the same dead runtime id every 5 seconds for the lifetime of the
window. On one machine this produced 31,518 gateway rejections in a day
(vs 663 the day before), 18,614 of them against a single runtime id, and it
is what users see reported as "sessions stopped with a session not found
error" after an update.

The trigger is a reconnect, not the poll itself: anything that mints a fresh
runtime (gateway restart, the #94219 reconnect/replay work, an idle-reaped
pooled backend) strands the id the status stack is still holding, and nothing
in this path ever re-checked it.

Distinguish the two failure classes:

- 4001 / "session not found" is TERMINAL for that runtime id — latch the id
  and stop polling it.
- A timeout or transport error is transient — keep retrying, since the
  session may well still be alive. Misclassifying that direction would
  silently freeze the status stack on a healthy session.

The latch is cleared when the status stack (re)binds a session id, so a
session that comes back under a fresh runtime resumes polling normally
rather than staying dark for the life of the app.

Also name the method in the gateway's 4001 warning. That line was added in
c305839442 "for diagnosability", but without the RPC name it cannot say WHICH
client call is looping — the reason this storm could not be attributed from
the logs alone. A ContextVar set in `handle_request` carries it; it is
diagnostic only and never used for authorization.

Tests:
- composer-status: 4001 stops the poll, a timeout does not, one gone session
  never suppresses a healthy sibling, and a rebind resumes polling.
- tui_gateway: the rejection warning names the method.
2026-08-26 04:49:22 -07:00
codexbt 3dea11d703 fix(web_server): recheck WEB_DIST existence dynamically in mount_spa
When hermes dashboard --skip-build runs across agent updates, mount_spa checked WEB_DIST.exists() once at server startup and mounted an immutable 404 handler if the build was missing. As a result, subsequent builds while the server was running continued to serve 404 "Frontend not built".

- Removes early static return in mount_spa().
- Moves WEB_DIST.exists() check dynamically into _serve_index() and serve_spa().
- Mounts /assets StaticFiles with check_dir=False.
- Adds unit test test_mount_spa_dynamic_web_dist_recheck in tests/hermes_cli/test_web_server.py.

Closes #82614
2026-08-26 04:48:58 -07:00
Shakti Prasad Mohapatra 98e87ac886 fix(cli): preserve stale positive behind-count on fetch failure (#92578)
huklaa's review: a failed fetch makes origin/main stale, so a stale ref
cannot prove *currentness* (rev-list 0 is inconclusive), but a stale
positive count is still sound evidence an update exists. On fetch
failure, compute the stale behind-count and return it when > 0;
otherwise return None (inconclusive) and still skip the cache write.

Regression tests:
- fetch failure + stale rev-list 0 -> None (not 'up to date')
- fetch failure + stale rev-list 5 -> 5 (update evidence preserved)
- fetch failure + rev-list error -> None
2026-08-26 04:17:39 -07:00
Shakti Prasad Mohapatra 55d50d5c92 fix(cli): don't serve stale update-check results after fetch failure (#82166)
When _check_via_local_git's git fetch fails (timeout, offline, DNS),
the code silently fell through to compare HEAD against the stale
origin/main tracking ref, which can report 0 (up to date) even when
upstream has moved forward. Combined with the 6-hour cache in
check_for_updates, a single fetch failure could suppress update
notifications for days — the exact symptom in #82166 where the daily
cron reported 'up to date' for 4 days after v0.20.0 was released.

Two fixes:

1. _check_via_local_git now detects fetch failure (returncode != 0 or
   exception) and returns None instead of falling through to stale
   refs. The caller treats None as 'check could not run' rather than
   'up to date'.

2. check_for_updates no longer caches None results. Previously, a
   None from a failed check was cached for 6 hours, suppressing
   retries until the cache expired. Now only conclusive results
   (0 or >=1) are cached, so the next check attempt runs immediately
   on the next call.

Added regression tests:
- test_check_via_local_git_fetch_failure_returns_none
- test_check_for_updates_does_not_cache_none
2026-08-26 04:17:39 -07:00
kshitijk4poor d62a05e94c fix(checkpoints): surface skipped_oversize to users and stop misreporting failed deletes as restored
Follow-up to the salvaged #95207 fix, completing the misreport bug class:

- restore() now also drops delete_targets whose unlink failed (OSError
  swallowed) from restored_files — the sibling of the kept-oversize
  misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
  (slash_commands + gateway.rollback.kept_oversize locale key in all 17
  catalogs) now tells the user which files were kept because the size
  cap excluded them from every checkpoint; previously the file was
  correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
2026-08-26 16:44:43 +05:30
RickyYii 595b5ce68a refactor(checkpoints): call the size-cap predicate instead of restating it
Review feedback on #95207: `_exceeds_size_cap` and `_drop_oversize_from_index`
each computed the byte cap and compared against it. Both used `> cap`, so they
agreed, but only by coincidence of two independent expressions — nothing held
them together.

The coupling is the whole point of the fix. The checkpoint decides what to
store and safe restore decides what may be deleted; a threshold that drifted
between them would produce a file both absent from the checkpoint and not
recognised as capped at restore, which is exactly the deletion this branch
exists to prevent. `_drop_oversize_from_index` now calls the predicate.

Added a boundary case to TestSafeRestore that pins the round trip from both
ends: a file at exactly the cap is stored, so it must revert; one byte more is
excluded, so it must be kept. Mutation-checked — moving either side to `>=`
fails it, including the re-inlined-with-`>=` shape the reviewer described.

No behaviour change: the byte cap, the strict comparison and the
unstattable-path result are all as before.

Regression: the 10 test files covering checkpoint_manager / rollback, against
current main (1fe0f2f3a, 134 commits newer than the base measured on the first
commit) — 130 passed on main, 135 here (+5 new), zero failures either side.
2026-08-26 16:44:43 +05:30
RickyYii d28bf79927 fix(checkpoints): stop safe restore deleting files the size cap excluded
`/rollback <N>` runs `restore(..., safe=True)` — safe mode is the default,
`--all` opts out. Safe mode splits the changed files into two groups: those
present in the checkpoint are checked out, and those absent from it are treated
as files Hermes created during the turn and deleted, since deleting them is
what restores the pre-turn state.

Absence from the checkpoint is not proof of authorship. `max_file_size_mb`
(default 10) keeps large files out of every checkpoint via
`_drop_oversize_from_index`, so a file the agent appended to — a dataset, a
corpus, an export, a log — is absent for a completely different reason. Safe
mode deleted it. No checkpoint held a copy, so nothing could bring it back, and
`restored_files` listed the path, so the user was told it had been restored.

Reproduced on main with shipped defaults:

    corpus.jsonl (2 MB), agent appends to it, then /rollback 1
    safe_restore_plan restore=['corpus.jsonl', 'notes.py']
    restore ok=True restored_files=['corpus.jsonl', 'notes.py']
    notes.py     exists=True   content="v1 = 'original source'"
    corpus.jsonl exists=False  <- deleted, was in no checkpoint

Scope: this needs an agent write to the capped file. A large file Hermes never
touched is not in the ledger, lands in `skipped`, and was already safe.

The delete branch now asks whether the path is one the cap would have excluded,
using the same test `_drop_oversize_from_index` applies when building the
checkpoint, so "kept out of the checkpoint" and "refused deletion at restore"
share one definition. Such a path is reported under a new `skipped_oversize`
key and dropped from `restored_files`.

The classification keys on "absent from the checkpoint", not on "large now".
A file small enough to be checkpointed and later bloated past the cap does have
a stored version, and reverting to it is exactly what was asked for — it still
restores, and a test pins that.

The ledger records a content hash, not whether a write created or modified the
file, so an oversize path cannot be proven agent-created. Leaving one behind
costs a stale file the user can delete; removing it costs the file.

Tests: 4 cases in tests/tools/test_checkpoint_manager.py::TestSafeRestore. Two
fail on main — the deletion and the misreport. Two are guards: the
grew-past-the-cap revert, and the small agent-created file that must still be
removed.

Regression: the 10 test files covering checkpoint_manager / rollback —
130 passed on main, 134 with this change (+4 new), zero failures either side.
2026-08-26 16:44:43 +05:30
Teknium 9f8cdf89d6 fix(macos): keep the TCC anchor alive across CVE-repair rotations + sign anchor copies
Integration fixups so #82529's generation signing and #95131's interpreter
anchor cover each other's gaps (without these, each fix leaves the other's
rotation path broken):
- macos_tcc_anchor store detection now recognizes .hermes-runtime/python/
  generation-* stores: repair_vulnerable_runtime() rebuilds the venv against
  a generation interpreter, replacing the anchored bin/python with a fresh
  symlink — previously the anchor then read 'not uv-managed' and NEVER
  re-anchored, so every SQLite CVE repair silently orphaned terminal TCC
  grants (the exact #82427 scenario, path-keyed).
- _install_anchor signs the anchor copy with the same identifier-pinned DR
  (via managed_uv._macos_sign_managed_python) before it goes live: copy2
  carries the source build's cdhash-based signature, so an unsigned refresh
  would still change the stored csreq on every patch bump/repair despite
  the stable path. Best-effort, never blocks the anchor.
- Tests: generation-store recognition + repair-generation anchoring +
  sign-on-install call (sabotage-verified: dropping the generation root
  marker fails both new tests).
2026-08-26 04:14:16 -07:00
notkisk 8d6c0a3098 fix: preserve macOS TCC identity for managed Python 2026-08-26 04:14:16 -07:00
Teknium 19fde8a450 fix(dashboard): compare and spawn the venv interpreter by UNRESOLVED path
Follow-up on the cherry-picked #90030: candidate.resolve() breaks the fix
on the standard Linux venv layout, where venv/bin/python is a symlink to
the base interpreter. Resolving makes the venv python compare equal to
the dependency-less base (so the swap never happens), and returning the
resolved target would spawn the bare base interpreter, bypassing
pyvenv.cfg — the fix would silently not fix #90026 on the exact platform
it was reported from. Compare and return normalized UNRESOLVED paths:
the venv path IS the interpreter's identity. Adds the symlink-layout
regression test; live-E2E'd with a real dependency-less base runtime.
2026-08-26 04:03:21 -07:00
liuhao1024 ad8f995bcf fix(dashboard): spawn detached actions from the install's venv interpreter
Under an SSH remote backend the web server is launched by running the uv
BASE interpreter with the venv's site-packages injected into sys.path at
startup, so sys.executable is a dependency-less python. Detached
dashboard actions spawned from it (Update now, restart, anything routed
through _spawn_hermes_action) inherited neither the injected path nor a
PYTHONPATH and died on the first third-party import — 'Update now'
always failed instantly with ModuleNotFoundError: No module named 'yaml'
while 'hermes update' from the venv worked (#90026).

_dashboard_spawn_executable now prefers the install's own venv
interpreter (venv/bin/python, venv/Scripts/python.exe) when it differs
from sys.executable, resolving the same dependency set the venv launcher
provides. Same-interpreter launches return sys.executable unchanged,
preserving the Windows console-ownership behavior verbatim, and layouts
without an install venv keep the old fallback.
2026-08-26 04:03:21 -07:00