Commit Graph

14240 Commits

Author SHA1 Message Date
emozilla 5d5179d7a9 fix(windows): remove the hermes launchers on uninstall
Every uninstall mode deletes the code checkout, but the launchers in
the managed binary dir (%LOCALAPPDATA%\hermes\bin) live outside it and
survived -- so `hermes` in a new terminal resolved to a launcher whose
venv target was gone and errored, which reads worse than
command-not-found.

remove_windows_bin_launchers deletes both launcher forms (.exe/.cmd)
from the managed binary dir in every uninstall mode, anchored on the
default Hermes root so profile sessions cannot redirect the sweep into
profiles\<name>\bin. When the uninstall itself runs through the
launcher, that exe is mandatory-locked against deletion but not rename
(the same fact _quarantine_running_hermes_exe relies on), so it falls
back to renaming the launcher aside.

The managed uv (uv*.exe) in the same dir survives, and the hermes\bin
PATH entry is swept only on a full wipe from the default root
(include_managed_bin) -- a keep-data uninstall keeps the still-working
uv resolvable for reinstalls.

A lockstep test parses install.ps1's staging loop so the swept names
cannot drift from the staged names silently.
2026-08-22 13:39:04 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
poisdahl a36d6704b6 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-22 17:31:04 +02:00
poisdahl fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
Kshitij Kapoor 8ee0103ea2 fix(gateway): finite-bounded watchdog knob validation + wire keys through load_gateway_config
Addresses both review findings from @egilewski on #89134:

- Non-finite values: _coerce_int now degrades int(inf) (OverflowError
  previously ABORTED gateway config loading); the clamp requires
  math.isfinite plus sane upper bounds (interval <=3600s, timeout
  <=600s, strikes <=1000), falling back to the shutdown_watchdog
  constants.
- Loader wiring: load_gateway_config builds gw_data FLAT and never
  forwarded the yaml gateway: section, so loop_watchdog* keys —
  including the PRE-EXISTING loop_watchdog bool documented in
  config_defaults — were silently ignored on the real startup path.
  Bridged with the established top-level-wins/nested-fallback pattern.

E2E: config.yaml with loop_watchdog:false + strikes:12 + interval:.inf
now yields False/12/30.0 through the real loader.
2026-08-22 20:40:23 +05:30
Kshitij Kapoor 3616145723 fix(gateway): keep loop-watchdog default at 3 strikes; dedupe constants; register knobs in config defaults
Downscope of the salvaged #89134 per review: the 3->8 default raise was
symptom tolerance for the false-positive class the off-loop heartbeat +
two-witness probe fixes at the root — fleet-wide it would only delay
genuine-wedge recovery ~2.7x. The three tuning knobs keep independent
operator value and stay:

- default max_strikes back to 3 everywhere (constant, dataclass,
  from_dict fallback, floor clamp, tests)
- gateway/config.py + gateway/run.py now reference the
  shutdown_watchdog DEFAULT_* constants instead of duplicating literals
  in three places (drift hazard)
- knobs registered in hermes_cli/config_defaults.py alongside the
  sibling gateway.loop_watchdog bool
2026-08-22 20:40:23 +05:30
devops aa08cb8ccb fix(gateway): make loop-liveness watchdog tolerant of transient reconnect stalls
The event-loop liveness watchdog (gateway.shutdown_watchdog) hard-exited with
code 75 after 3 consecutive missed probes (probe_interval=30s, timeout=10s,
max_strikes=3), i.e. ~90-120s of loop block. Telegram/Discord reconnect during
a network blip does synchronous socket I/O on the loop and can block it for
60-90s; these stalls self-recover (recurring fleet incidents on 2026-08-17
stalled cron dispatch ~21h via restart churn, kanban t_0f76430f).

Raise the default max_strikes 3->8 so a transient reconnect stall is tolerated
while a genuine multi-minute wedge still escalates, and expose the three
tolerance knobs via config.yaml (gateway.loop_watchdog_probe_interval_s /
_probe_timeout_s / _max_strikes) so operators can tune per deployment.

Refs: kanban t_70483f23
2026-08-22 20:40:23 +05:30
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
Kshitij Kapoor 684e95a001 fix(gateway): claim ledger rows and clear resume_pending inline before the abandonable boot-send task
Post-merge follow-up to #92173. The claim + resume-clear lived inside
the boot-send task AFTER the restart notification — itself a
flood-controllable send. If that notification outlived the restore-gate
timeout, the gate opened with zero rows claimed and the resume
scheduler replayed turns whose answers were already in the ledger,
while the background task later redelivered them too (duplicate
delivery + re-paid turn).

Split _redeliver_pending_obligations into _claim_pending_obligations
(pure DB: sweep + resume clear, awaited inline before the send task
exists) and _redeliver_claimed_obligations (network half, stays inside
the bounded task). The original name remains as a composition wrapper.
Mutation-checked: both updated gate tests fail on pre-split run.py.
2026-08-22 16:26:51 +05:30
kshitijk4poor 4a6b362178 review follow-up: trim overreaching comment sentence, pin durable-copy assertion
- Drop the 'aborted before its tail' no-op sentence: early aborts are
  intercepted by the aborted/no-progress branches and never reach the
  would-grow check, so the framing overstated its relevance (2c finding).
- Test now also asserts the durable model_config copy still holds the
  armed runway after the refusal — locking in the memory==disk half of
  the contract, not just the in-memory value.
2026-08-22 16:07:12 +05:30
kshitijk4poor 4c76ec81a9 fix(compression): restore the prune runway when a would-grow refusal keeps the transcript
compress()'s successful tail zeroes _proactive_prune_rearm_tokens in
memory — correct for a committed compaction, whose boundary already broke
the prompt-cache prefix. But compress_context's anti-growth guard can then
REFUSE the result and keep the original transcript, whose cached prefix is
intact. The refusal returned with the in-memory runway still at 0 while the
durable model_config copy kept the old value, so:

- the next eligible iteration's proactive prune fired without the regrowth
  interval #79640 introduced — an immediate, unthrottled cache-breaking
  rewrite (#91830's bug class), and
- memory and disk disagreed until a restart silently re-armed the throttle
  from the stale durable row.

The refusal branch now restores the runway from the attempt snapshot — the
same targeted restore the rotation-failure rollback already performs.

Sibling non-commit branches audited: aborted (returns before the tail
zero), no-progress (tail zero only runs after a real boundary rewrite,
which no-progress by definition lacks), empty-transcript (built-in tail
never returns []), fence-denied (full snapshot restore already covers the
runway), in-place DB failure (in-memory transcript keeps the compacted
form, so the zeroed runway is consistent with it).

Fixes the reachable half of the structural asymmetry flagged in #91830.
2026-08-22 16:07:12 +05:30
fangliquan a4f16e3fef fix(gateway): retain failed replacement evidence 2026-08-22 15:27:26 +05:30
fangliquan 596bfc557f fix(gateway): distinguish failed systemd replacements 2026-08-22 15:27:26 +05:30
fangliquan 83b09ebd0a fix(gateway): make handoff recovery idempotent 2026-08-22 15:27:26 +05:30
fangliquan 91fb175188 fix(gateway): preserve systemd handoff recovery 2026-08-22 15:27:26 +05:30
fangliquanflq 5b024c7ccc fix(gateway): make systemd the sole restart owner 2026-08-22 15:27:26 +05:30
fangliquan f12cd04015 fix(gateway): isolate kanban dispatcher to_thread context
Spawn-time Context isolation cannot rewrite an already-running watcher task. Run dispatcher SQLite offloads in an empty Context so write_txn no longer false-trips after delegate_task, while real child callers still hit the mutation guard.
2026-08-22 15:25:50 +05:30
fangliquan bf3a0bb99d fix(gateway): isolate supervised watcher contexts 2026-08-22 15:25:50 +05:30
Jack Lau e173720774 fix(gateway): give supervision exhaustion an owner for queued platforms
Review of #90448 by @andrexibiza: adding _ensure_reconnect_watcher_running()
to the already-queued branch of a fatal callback is still an event-coupled
check. It needs a later fatal error from some other platform to arrive, and
#81036 makes that less likely rather than more -- it publishes the queue
before disconnect and drops the failed adapter from the live map, so after
the watcher's supervised restart budget is spent there may be no adapter
left to emit the event recovery is waiting on.

That is the state #72366 (salvage of #71867 by @ygd58) restored supervision
to close: queued work exists, the watcher is dead, and nobody owns the
invariant. Supervision being finite is correct; having no owner past the
budget is not.

_spawn_supervised now takes on_give_up, invoked when it abandons a task --
the supervisor is the only thing that knows it has. The reconnect watcher
uses it to hold:

  while _running and _failed_platforms is non-empty, either a reconnect
  watcher is live or a bounded respawn is scheduled.

Empty queue: leave it down and log; the enqueue path spawns a fresh watcher
the moment something depends on one. Non-empty: a bounded slow tier at
_RECONNECT_WATCHER_SLOW_RETRY_SECS (300s) for _MAX_SLOW_WATCHER_RESPAWNS (6)
attempts, standing down early if the queue drains or a watcher returns on
its own. Exhausted: one loud error naming the platforms left unattended.

The ceiling is (1 + _MAX_SUPERVISED_RESTARTS) x (1 + _MAX_SLOW_WATCHER_RESPAWNS)
spawns -- 42 across at least half an hour -- because each slow attempt hands
the watcher a fresh supervised budget. A test asserts that ceiling so it
cannot quietly become a restart loop.

Deliberately NOT included: requesting a process restart when the slow tier
is also exhausted. Taking down every healthy platform to heal a sick one is
a blast-radius policy decision for a maintainer.

Two things this turned up:

- _spawn_supervised did not thread on_give_up through its own backoff
  respawn, so the callback was lost after the first restart and the give-up
  branch had no owner at exactly the moment it needed one -- the same defect
  the on_spawn docstring warns about, one parameter over.
- Three call sites repeated the (factory, name, on_spawn) triple, whose
  on_spawn half is load-bearing. They now go through
  _spawn_reconnect_watcher().

_supervised_backoff() names the previously-inline exponential schedule so
the exhaustion tests can collapse it; production behaviour is unchanged.

Refs #90386
2026-08-22 15:25:45 +05:30
Jack Lau 92018e76a8 fix(gateway): heal a dead reconnect watcher when the platform is already queued
_ensure_reconnect_watcher_running() exists for one situation: the reconnect
watcher has exhausted _MAX_SUPERVISED_RESTARTS, so _spawn_supervised has logged
"giving up restarts" and will never bring it back on its own (#70344, and the
supervised-restart half of #71758). It had exactly one call site, inside the
newly-queued branch of _queue_retryable_fatal_platform.

That branch is unreachable for a platform already in _failed_platforms, which
is the only kind of platform the watcher can have been retrying long enough to
burn five rapid restarts on. So the backstop could not fire in the one state it
was written for.

The failure is silent by construction. The early return logs nothing, so there
is no "queued for background reconnection" line. The stranded check in
_handle_adapter_fatal_error_detached deliberately treats a queued platform as
safe, so the gateway does not exit for the service manager either. With another
platform still connected, self.adapters is non-empty and the "gateway staying
alive, watcher will retry in background" branch is skipped too. A retryable
fatal error can therefore produce a single ERROR line and then nothing: the
platform sits in the queue that nobody is draining until someone restarts the
process by hand (#90386 reports 4h17m of that, with cron unaffected throughout).

Call the ensure on the already-queued path as well. It is already idempotent
and already cheap: it returns immediately unless the tracked task is done, and
it routes through the same on_spawn handle tracking, so a live watcher is never
duplicated.

The queue entry itself is deliberately left untouched. Re-enqueueing would
reset attempts and next_retry, restarting the backoff ladder on every fatal
error and hammering a provider that is already refusing the connection.
2026-08-22 15:25:45 +05:30
liuhao1024 349d9aee43 fix(tui): log the refused shared-handle transfer and pin _get_db caching 2026-08-22 15:25:35 +05:30
liuhao1024 bd2afde48f fix(tui): never transfer the shared launch SessionDB to one agent
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in #81071).

Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
2026-08-22 15:25:35 +05:30
HexLab98 ce944a5a55 fix(gateway): do not let boot-path sends hold the inbound gate
Restart notification and obligation redelivery ran before the
startup-restore gate opened, so one hung Telegram send queued inbound
on every platform. Bound those sends with the same timeout the resume
gate already uses, and clear resume_pending before send so a timed-out
redelivery cannot also replay the turn.
2026-08-22 15:25:30 +05:30
HexLab98 a444b673ad fix(telegram): fail closed on long send-path flood waits
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
2026-08-22 15:25:30 +05:30
carryzuo00 8a963e8512 test(terminal): cover gateway ContextVar session-key path
The existing session-key regressions set HERMES_SESSION_KEY via os.environ,
which only exercises the os.getenv() fallback branch. Real gateway turns bind
the identity through gateway.session_context.set_session_vars() (a ContextVar)
and never write the process-global env var. Add two companion regressions that
bind via set_session_vars() with HERMES_SESSION_KEY absent from os.environ:

- test_session_key_from_contextvar_without_environ: container slot scopes to
  session:<key> purely through the ContextVar (subagent inheritance covered).
- test_contextvar_session_key_wins_over_environ: with a different value left in
  os.environ, the ContextVar-bound session wins, so two concurrent gateway
  sessions in one process cannot cross-contaminate via the process global.

Cleanup via clear_session_vars(tokens) in finally.
2026-08-22 15:07:05 +05:30
carryzuo00 a270c4adea fix(terminal): scope environment cache by session key to prevent cross-profile SSH leakage
_resolve_container_task_id always returned "default", so _active_environments
shared a single SSHEnvironment across all WebUI sessions. When a user switched
from profile A (ssh_host=10.0.0.1) to profile B (ssh_host=10.0.0.2), the new
session found _active_environments["default"] already set to A's SSHEnvironment
and reused it — silently running every command on the wrong remote host.

Fix: when HERMES_SESSION_KEY is present (set per-session by the WebUI streaming
layer and per-message by the gateway via contextvars), return "session:<key>"
as the cache key instead of "default". Each session now owns its own slot in
_active_environments and always creates an environment from its own profile's
TERMINAL_SSH_HOST / TERMINAL_ENV config.

Behaviour unchanged in CLI mode (no HERMES_SESSION_KEY → still "default").
RL/benchmark task overrides (register_task_env_overrides) are unaffected.
Subagent task_ids inside a WebUI session collapse to "session:<key>" so they
continue to share the parent session's container.

Five new regression tests added to test_shared_container_task_id.py.
2026-08-22 15:07:05 +05:30
Epic ab3e2f563b fix(discord): render provider model lists >25 options across multiple select menus
The Discord /model picker built a single discord.ui.Select filled with
models[:25], silently dropping any models beyond the first 25. Discord
caps a single select at 25 options but allows up to 5 component rows, so
partition the list across up to 3 select menus (25 each; Back/Cancel use
the other 2) instead of truncating.

This fixes providers like Nous (curated list + Portal recommendations
exceed 25) whose tail — including free-tier :free Portal picks — was
previously clipped on Discord while showing fine in the Portal UI / CLI.

- _build_model_select: slice into <=25-option chunks, one select per chunk
  (custom_id model_model_select_<i>), all via _on_model_selected. Multi-row
  menus get a (n/total) placeholder suffix.
- _on_provider_selected: 'N more available' count reflects models actually
  rendered across the partitioned menus.
- Add regression test covering the 37-model Nous case (no truncation/dupes,
  per-menu 25 cap holds).
2026-08-22 14:55:04 +05:30
kshitijk4poor 1fe8683e58 fix(state): split forensic-backup identity from repair-epoch fingerprint; publish backup bundle atomically
Addresses two data-integrity gaps @andrexibiza flagged reviewing #88425.

1. Forensic dedupe no longer reuses the repair-epoch fingerprint.
   _db_fingerprint masks SQLite's commit counters and samples only head/tail
   so an ordinary write does not re-key the repair budget — the right
   predicate for 'same damage epoch', the WRONG one for 'same recovery
   image'. A live writer committing rows into an interior page (size
   preserved, head/tail untouched) collided under it, so _backup_db_file
   handed back a STALE backup that predates real user data. New
   _backup_content_identity() digests the whole file + every sidecar; the
   dedupe uses it. The O(n) read is cheaper than the O(n) copy it avoids on a
   hit.

2. Backup bundle is now published atomically. The promotion loop replaced
   files one at a time (main first) and cleanup unlinked only staging srcs,
   so a sidecar os.replace failure after the main promotion left the
   final-prefix main backup on disk — a countable-but-incomplete bundle that
   passed the #69603 hard stop and deduped as legitimate next pass. Now
   sidecars publish first and the main DB last (its name is the commit
   marker _existing_malformed_backups counts), and cleanup rolls back every
   already-published destination.

Two regressions added (both mutation-checked — each fails on pre-fix code):
- test_backup_not_deduped_after_interior_page_write
- test_publication_failure_leaves_no_countable_partial_bundle

tests/test_state_db_repair_loop_mtime.py: 28 passed.
2026-08-22 14:33:20 +05:30
kshitij 5777e68b3d fix(state): include the rollback journal in the forensic backup
The pre-repair copy took only -wal/-shm. In rollback-journal (DELETE) mode --
Hermes's fallback on NFS/SMB/FUSE/ZFS and on WAL-reset-vulnerable SQLite builds
-- a hot <db>-journal exists on disk whenever a transaction was open, and that
file is what rolls the damaged bytes back to a consistent state. A forensic copy
without it cannot be recovered by hand, which is the entire purpose of taking
the copy before destructive surgery.

Verified the journal is really there:

  files while a txn is open: ['state.db', 'state.db-journal']
  files after commit:        ['state.db']

Add _DB_SIDECAR_SUFFIXES = ("-wal", "-shm", "-journal") and use it at the four
sites that must agree: the disk-guard sizing, the staging copy, the
backup-count exclusion in _existing_malformed_backups (so a copied journal is
not itself counted as a forensic backup), and _prune_malformed_backups (which
otherwise leaks one journal per pruned backup, quietly defeating the retention
cap this PR is partly about).

Matches the spelling hermes_cli/session_recovery.py:61 already uses for the
same concept.
2026-08-22 14:33:20 +05:30
kshitij 8779b782b3 fix(state): exclude SQLite's commit counters from the repair fingerprint
Third self-review pass found the content fingerprint was still defeated on
rollback-journal deployments, by the same mechanism as the original mtime bug.

The head sample starts at byte 0, so it covers the database header's file
change counter (bytes 24-27) and version-valid-for (92-95). In DELETE mode a
commit writes the main file directly and bumps both. A malformed-SCHEMA DB
still accepts writes -- that is the whole premise of this PR -- so any ordinary
session write between passes re-keyed the ledger:

  DELETE, 18MB db, one peer UPDATE between passes (before this commit)
    pass 1..6: attempts=1 every pass, exhausted=False -> unbounded loop

  after
    pass 1..3: attempts=1,2,3   pass 4: BLOCKED

WAL is unaffected (commits land in -wal; the main header only moves on
checkpoint), so this was invisible on a WAL host and reproducible on every
NFS/SMB/FUSE/ZFS or WAL-reset-vulnerable host -- exactly the deployments the
earlier lock-safety commit was written for.

Mask the two volatile ranges out of the sample. Page 1's sqlite_master b-tree
sits after byte 100 and stays in, so genuine recovery still resets the budget:
verified schema rewrite, index rebuild, VACUUM and truncation all change the
key, while a bare utime and an ordinary commit do not.

Test-cost cleanup in the same file, since the new tests needed a
larger-than-sample fixture and the file was already slow:
  - the two guard tests that allocated 450MB of os.urandom now use sparse
    truncate (both only ever read st_size), and the new fixtures use 600 rows
    rather than 40k;
  - file runtime 127s -> 35s.
2026-08-22 14:33:20 +05:30
kshitij 602c45e45e fix(state): never let a peer connection reset the repair budget
Self-review of the previous commit found it reintroduced the bug this PR
exists to fix, by a different route.

`_db_fingerprint` fell back to `size:mtime_ns` when a live connection made the
content read unsafe. The ledger compares keys for EQUALITY, and the two keys
have different SHAPES, so a gateway peer connecting between passes flipped the
shape and the counter reset to 1 every time:

  pass 1 [offline] attempts=1  fp=8192:58c7924f0fba...
  pass 2 [LIVE   ] attempts=1  fp=8192:1786972039271402096
  pass 3 [offline] attempts=1  fp=8192:58c7924f0fba...
  ... never reaches _MAX_PERSISTENT_REPAIR_ATTEMPTS

Return None instead, and teach the two ledger helpers to cope:

- `_persistent_repair_attempts_exhausted` falls back to the recorded key's
  SIZE prefix (the one component both shapes share and that needs no raw
  read) rather than reading as "not exhausted" — otherwise a peer connection
  hides an exhausted budget on every pass, same loop.
- `_record_repair_outcome` keeps the key already on record and still
  increments, rather than dropping the pass.

  pass 1 [offline] attempts=1  pass 2 [LIVE] attempts=2
  pass 3 [offline] attempts=3  pass 4 [LIVE] BLOCKED

Intra-pass flips were already safe (the probe and the record are both reached
with the same liveness within one `repair_state_db_schema` call); it is the
cross-pass change that desynced.

Also drops two `type: ignore` directives `ty` flagged as unused, and replaces
the `LiveConnectionError = ()` / `nullcontext()` shim with a real no-op
contextmanager + exception class so the scaffold-install path is honest.
2026-08-22 14:33:20 +05:30
kshitij 8de64b1634 fix(state): stop backup staging from posing as a forensic copy
The staging name was derived from the backup name
(`<db>.malformed-backup-<stamp>.incomplete`), which still matches the prefix
`_existing_malformed_backups` selects on -- it excludes only `-wal`/`-shm`.
Three consequences, all reproduced:

  - it is COUNTED as a forensic backup;
  - it sorts NEWEST (`.incomplete` > the bare stamp), so prune's
    keep-3-newest slice retained partials and deleted intact copies -- the
    exact inversion the staging change was meant to prevent;
  - worst, the dedupe ran BEFORE the sweep, and a staging file orphaned by a
    kill mid-copy is a byte-identical copy of the damaged DB, so its
    fingerprint MATCHES and it was handed back as the official `backup_path`.
    Repair then passed the #69603 hard-stop gate and ran destructive surgery
    believing a forensic copy existed, and the next pass's sweep deleted that
    very file.

Move staging outside the prefix (`<db>.backup-staging-<stamp>`) and sweep
before the dedupe. The sweep also matches the pre-merge `.incomplete`
spelling so a host that ran the earlier build does not keep prefix-matching
debris that sorts newest and survives prune forever.

Before / after on the same fixture (orphaned staging + a later pass):

  before  backup_path = ...malformed-backup-<stamp>.incomplete   (staging!)
          pass-1 forensic copy deleted by the next sweep
  after   backup_path = ...malformed-backup-<stamp>              (real copy)
          debris swept, pass-1 forensic copy preserved
2026-08-22 14:33:20 +05:30
kshitij b3f14c8534 fix(state): keep the repair fingerprint from cancelling POSIX advisory locks
The content fingerprint takes a raw descriptor, and close() on ANY descriptor
cancels every POSIX advisory lock the process holds on that file. The
exhaustion probe runs before _backup_db_file's has_live_connection guard, so
the read happened even when a peer SessionDB held a write lock.

Verified end-to-end (journal_mode=DELETE, gateway mid-turn write, peer in a
subprocess):

  before   peer BLOCKED -> repair -> peer BLOCKED, holder COMMIT ok
  unfixed  peer BLOCKED -> repair -> peer STOLE the lock,
                                    holder COMMIT: disk I/O error

WAL is immune (it coordinates through -shm), but DELETE is what Hermes falls
back to on NFS/SMB/FUSE/ZFS and on SQLite builds vulnerable to the WAL-reset
bug, so this is a real deployment shape.

Run the read under offline_file_access and fall back to size:mtime_ns when a
connection is live. That keeps the ledger counting instead of returning None
(which reads as "not exhausted" and would restore the unbounded loop), and the
content key stays load-bearing on the offline repair path -- the only path
where surgery actually runs.

Also fail the free-space guard CLOSED: a nearly-full volume is exactly where
statvfs is likeliest to fail, and proceeding is the multi-GB copy that finishes
off the disk.
2026-08-22 14:33:20 +05:30
jirathip-k c914a9ac4b fix(state): make backup atomic and the disk guard proportional
Follow-up to adversarial review of the first commit. Three findings, two
confirmed by test and fixed here, one disproven and left alone.

CONFIRMED — the free-space guard was a threshold, not cleanup. Prune runs
only on the success path, so any copy that failed partway (ENOSPC, sidecar
copy failure, kill mid-copy) left a file matching the `malformed-backup-`
prefix that nothing ever removed. Measured on the unpatched tree: backups
capped at 3 while copies succeed, but 13+ and climbing once copy2 raises —
self-reinforcing, since each partial consumes the space that guarantees the
next failure. Worse, partials sort newest-by-name, so a later successful
prune KEPT the garbage and deleted the intact forensic copies.
Fix: copy to a `.incomplete` staging name that does not match the backup
prefix, os.replace into place only after every copy succeeds, unlink staging
on failure, and sweep stale staging debris on entry.

CONFIRMED — the 2GiB floor was a small-volume regression. A 50MB DB on a
10GB volume with 1.5GB free (30x headroom) was refused, and since a refused
backup is a HARD STOP (#69603) that silently converts "repair loops" into
"repair never runs". Fix: require the copy itself (now including its
-wal/-shm sidecars, which the old check ignored) plus proportional headroom
— max(256MiB, 2% of volume).

DISPROVEN — the review claimed a refused backup skips _record_repair_outcome
so the loop never terminates. It does not: repair_state_db_schema records the
outcome on the result returned by _repair_state_db_schema_locked, which is
where the hard stop returns. Verified on a simulated low-disk host: terminal
at pass 4 with zero backups written. No change made.

Tests: 5 new (small-volume allow, proportional headroom, sidecar accounting,
failed-copy leaves no countable debris + staging swept). 23 pass with the
#86747 suite; test_hermes_state.py 252 passed. Pre-existing unrelated
failures unchanged.
2026-08-22 14:33:20 +05:30
jirathip-k 27d661e171 fix(state): stop unbounded state.db repair loop from filling the disk
A malformed-schema state.db sent Hermes into a repair loop that wrote a
fresh full-size forensic backup every ~10s: 31 copies / 2.3GB in 20
minutes, free space heading to zero on a host running an agent fleet.

The #86747 guards for exactly this were already present and did not hold.
Both keyed on `size:mtime_ns`:

  * `_db_fingerprint` -> the ledger's attempt counter reset to 1 on every
    pass, so `_MAX_PERSISTENT_REPAIR_ATTEMPTS` was never reached and the
    loop never terminated;
  * `_backup_db_file`'s dedupe compared mtime, so it never matched and
    each pass wrote another full-size copy.

The assumption behind that key -- "nothing can successfully write to a
damaged file" -- holds for the b-tree damage of #86747 but not for the
malformed-SCHEMA class: the DB still opens and accepts writes (only
sqlite_master is unreadable), so live writers, WAL checkpoints and the
in-place repair strategies themselves all move mtime between passes.

Fixes:

  * fingerprint on size + a bounded head/tail content sample instead of
    mtime. Stable across passes that merely touch the file, still changes
    on genuine repair/truncation/restore (so recovery resets the budget),
    and stays O(1) on a multi-GB DB.
  * dedupe the forensic backup on that same fingerprint.
  * add the missing free-space guard: refuse the pre-repair copy when it
    would leave under 2GiB free, with an actionable error. The backup is a
    full raw copy of the damaged DB, so a repair loop is a disk amplifier
    that can take down every process on the host -- and the refusal path
    already hard-stops the repair (#69603) rather than mutating the only
    remaining copy.

Tests fail on the unfixed tree and pass here; the pre-existing failures in
test_state_db_malformed_repair.py and TestFTS5Search are unrelated and
reproduce on the base commit.
2026-08-22 14:33:20 +05:30
Teknium a9860d413d fix(bot-mode): the canonical Bot Chat is found by NAME — session-id pins removed
A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.

Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.

Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.

No migration shims: this IS the system.
2026-08-22 01:23:39 -07:00
Gille 9782275b2a fix(windows): restore dedicated CLI launchers on update 2026-08-22 00:05:51 -07:00
ethernet 969094e4d2 fix(tests): remove four shared-state and lifetime faults at high concurrency
The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.

1. Every pytest subprocess shared one temp root.

pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.

scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.

Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.

2. The config read guard walked directories that other tests were writing.

tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.

The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.

The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.

3. A PTY test waited for a file to exist, and not for its content.

tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.

The test now waits for the expected bytes, with a bounded deadline.

Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.

4. A dialog close timer outlived the test that started it.

ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:

    ReferenceError: window is not defined
     at resolveUpdatePriority (react-dom-client.development.js:1308)
     at dispatchSetState
     at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574

The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.

Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.

Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.

Verification:
- The affected Python files and the tests of the runner itself pass under
  scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
  "window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
  before this change. The two cleanup effects carry an eslint-disable line for
  the ref-mirror rule. They write a timer handle, and not a mirror of a
  reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
  no python3 outside the nix store, and the test uses the literal `python3`.
  The child exits 127 there. The fix rests on the 25-run measurement above and
  on CI.
2026-08-22 02:25:12 -04:00
Teknium 0a9a449a32 fix(desktop): Send Diagnostics review fixes — consent accuracy, log-grade redaction, dismissal guard, linkless-success (review feedback)
Addresses @helix4u's review on #92020:
- Consent notice now matches the real --nous contract: full logs up to
  512KB each, likely conversation content/tool outputs/file paths, viewable
  by Nous staff AND allowlisted Discord moderators (all 5 locales).
- Client-supplied text (error_context + extra_files) rides _redact_log_text
  — the same upload-safe redactor as backend logs (secrets + email masking),
  not the weaker bare secret pass; regression test covers both.
- ok:true without view_url or id becomes a structured failure; a returned
  id without a link renders an upload-ID fallback the user can quote.
- Generation guard in the store: dismissal is immediate in every phase
  (incl. mid-upload); a stale completion can no longer resurrect or
  overwrite the dialog. Cancel button never disabled.
2026-08-21 23:01:30 -07:00
Teknium 8f30e9c77a feat(desktop): Send Diagnostics — one-click redacted debug-bundle upload from the error card
New diagnostics.share_nous RPC reuses the CLI --nous pipeline
(collect_share_bundle → build_nous_bundle → share_to_nous) with redaction
forced on; accepts redacted error context + client-side extra files
(local desktop.log on remote connections) with sanitized labels and size
caps. Desktop: Send Diagnostics action on the failed-turn error card →
consent modal (privacy notice, explicit Upload) → private view link +
GitHub Issues / Nous Portal Support / Discord handoff. CLI --nous success
output gets the same three-destination pointer. i18n en/ja/zh/zh-hant/ar;
docs updated.
2026-08-21 23:01:30 -07:00
Teknium 9ddb6547a0 test(update): re-pin ZIP-fallback desktop test to the preserve-through-swap contract
The old contract WAS the bug (#70337): exe deleted by the swap, then
rebuilt from scratch. With the release-dir graft the exe survives the
swap; the test now asserts survival + original bytes.
2026-08-21 22:08:16 -07:00
JonthanaHanh 01c14ad7f3 fix(update): ZIP swap preserves the built desktop app (apps/desktop/release)
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).

Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-08-21 22:08:16 -07:00
kshitijk4poor eac3f645ef fix(update): don't ZIP-fallback on dependency failures or dirty trees
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.

- ZIP fallback now keys on git ACTUALLY having failed
  (_should_zip_fallback_on_update_error): a dependency-install failure
  after a successful pull can't be fixed by re-downloading source and
  would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
  (-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
  re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
  'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
2026-08-21 22:08:16 -07:00
Teknium 7d6db4efb8 fix(update): holder classifier derives value-flags from the real parser; de-flake goal-resume fixture
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.

De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
2026-08-21 19:11:55 -07:00
Teknium 8131b0a29f test(windows): #87594 probe asserts on the gateway ANCESTOR, not the direct parent
Diagnostic run showed the venv shim makes every spawn a launcher/worker
chain: the child's direct parent is its own launcher (python.exe
child_scan.py), and the gateway-argv process is the grandparent. The
probe now finds the gateway ancestor by argv — the same way the pause
machinery would — and asserts THAT pid is visible to the scan.
2026-08-21 19:11:55 -07:00
Teknium 4d3a61b63b test(windows): diagnostics in the #87594 probe — parent cmdline/exe + matcher verdict 2026-08-21 19:11:55 -07:00
Teknium 4c922a9348 test(windows): realistic gateway-parent argv in the #87594 live probe (child code via file, one-line -c) 2026-08-21 19:11:55 -07:00
Teknium c02cac00ce fix(update): venv-holder labels parse the real subcommand; gateway ancestors stay visible to the scan
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.

#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.

15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
2026-08-21 19:11:55 -07:00
Teknium f9aed7d7f6 test(windows): on-demand live venv-holder E2E lane + probe suite (#91277)
On-demand workflow (fires only on wine2e/** pushes, never on PRs/main)
that runs a live venv-holder E2E on windows-latest: real spawned
processes with Hermes argv shapes, real detection/classification/
message code against the live process table. Tests pin CORRECT behavior
for the cluster issues (#90778 mislabeling, #78089 long-path exemption,
#87594 ancestor-exclusion, #81774 serve premise), so unfixed bugs fail
on the runner — empirical premise-check before the consolidation fix.
2026-08-21 19:11:55 -07:00
Gille e9a7c7aa4d fix(telegram): omit topic routing from rich edits 2026-08-21 19:07:14 -07:00