Commit Graph

63 Commits

Author SHA1 Message Date
Teknium cfa327e5dc refactor(hclib): remaining hermes_cli library modules — dead code, unified helpers, flattened branches 2026-09-02 15:03:42 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium 045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
Konstantin Khlopkov aea2daee29 fix(backup): import into the active HERMES_HOME instead of the root home
run_import() resolved its restore target through get_default_hermes_root(),
which maps a profile home (<root>/profiles/<name>) back to <root>. A
profile-scoped import then overwrote the live root's config.yaml while the
profile directory stayed empty — while the 'Target:' line printed the
profile path, so the overwrite was silent.

Restore into the home the command operates under (get_hermes_home()), and
skip the automatic gateway service install when the restore landed in a
non-default home and a live default install exists: a second gateway on the
default service name would shadow or hijack the machine's primary install.

Fixes #99839
2026-09-01 09:28:46 -07:00
Teknium ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium cdd3b84c13 fix(state): reject special files in zeroed probe; real schema-bytes decode fixture
Follow-ups on the salvage: regular-file guard before the zeroed byte-probe (a FIFO at the state.db path would block startup forever — #98017 review P2), plus an on-main-reproducing UnicodeDecodeError fixture for #98924 (raw bytes in sqlite_master, not messages.content, are what reach pysqlite error-message decode).
2026-08-31 09:56:43 -07:00
loulanyue 69245e65bd fix(state): serialize startup across zero-byte check, quarantine, connect, and schema commit (#97568)
- Guard against concurrent-opener race where newly created 0-byte state.db was falsely quarantined before first schema write
- Wrap startup in quarantine_cross_process_lock when database is uninitialized or zeroed
- Guard is_zeroed_sqlite_file and is_zeroed_state_db against active live connections in current process
- Add concurrent-opener and live-connection regression tests
2026-08-31 09:56:43 -07:00
loulanyue a071fc80da fix(state): quarantine 0-byte truncated state.db and record store provenance (#97568) 2026-08-31 20:44:26 +05:30
Teknium f1d05ce7d8 fix(browser): real-profile snapshot is a first-class secret store + preserve channel identity
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:

Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
  (_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
  same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
  (honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod

Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
  Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
  stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots

Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
2026-08-26 19:25:33 -07:00
Teknium 7125e839d3 fix(state): fail closed when a live process still holds state.db during destructive restore (#90950)
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:

- restore_quick_snapshot's unlink+move fallback replaced the inode and
  deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
  the live writer's next checkpoint writes wrong-offset pages (the
  page-1 compression_locks clobber from the report).

Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
2026-08-26 06:28:43 -07:00
dsad a7988b830b fix(backup): clear stale SQLite sidecars on snapshot restore
`restore_quick_snapshot` swaps only the main `.db` file: copy to a temp name,
`unlink()` the destination, `move()` the temp into place. The destination's
`state.db-wal` / `state.db-shm` are left untouched.

The snapshot is a checkpointed `sqlite3.backup()` image (`_safe_copy_db`) that
owns no WAL, so a leftover `-wal` describes the database that was just
unlinked. An ungracefully killed gateway (SIGKILL / OOM / power loss /
container stop) strands exactly those sidecars -- and that is precisely when an
operator reaches for a snapshot restore. SQLite then replays that foreign WAL
over the restored file on the next open.

The module already documents this hazard for the backup archive:
`_EXCLUDED_SUFFIXES` says shipping sidecars next to a `backup()` copy "would
pair a fresh snapshot with stale sidecar state and produce a torn restore on
the next open." The same reasoning was never applied to the restore
destination.

Reproduced against the real `create_quick_snapshot` / `restore_quick_snapshot`:

  origin/main  restore_returned=True  sidecars_left=['-wal','-shm']
               integrity=*** in database main *** Tree 2 page 295:
               btreeInitPage() returns error code 11   rows=DatabaseError
  patched      restore_returned=True  sidecars_left=none
               integrity=ok   rows=2000

The restore reports success and returns True while leaving the database
malformed. `SessionDB` then fails to open on every subsequent start, and
`repair_state_db_schema` only rebuilds `sqlite_master`/FTS -- the damage is in
the `sessions`/`messages` b-trees, so it re-raises. The pre-restore `state.db`
is already unlinked, and because the stale `-wal` survives the restore,
re-running it corrupts the file again identically: the documented last-resort
recovery is wedged.

Every other DB-move site in the repo already handles sidecars --
`hermes_state.py:742` and `:1698`, `hermes_cli/kanban_db.py:1759`,
`hermes_cli/session_recovery.py:60`. `restore_quick_snapshot` was the outlier.

The two `hermes update` auto-restore sites (`hermes_cli/main.py`) have the same
shape and run unattended: they `copy2` the snapshot over `state.db` with the
sidecars still present, then `verify_sqlite_integrity` fails and prints
"Auto-restore FAILED -- restored copy also failed integrity". Same fix.
2026-08-26 06:28:43 -07:00
webtecnica 88d55f31e4 fix(backup): restore state.db through SQLite backup API so live connections see restored data (#65942) 2026-08-26 06:28:43 -07:00
xthezealot d422f7103e fix(backup): don't hang forever on locked SQLite sources
hermes backup freezes mid-archive when a .db file under HERMES_HOME is
locked by another process — e.g. a live Chromium profile database held
with an exclusive lock by a running browser. sqlite3.Connection.backup()
retries SQLITE_BUSY indefinitely and never honors the connection's busy
timeout, while a plain statement on the same source fails cleanly after
~5s with "database is locked".

Fixes:
- Probe the source with a cheap read before snapshotting, so a locked
  database fails fast instead of hanging the whole backup.
- Add a watchdog that interrupts the source connection after 15 minutes
  as a last resort for pathological cases.
- Exclude browser-profiles/ from full backups: the CDP browser profile is
  live, regenerable (cache + re-login), and unsafe to snapshot while
  running. On a real install this cut the backup from 28,396 files /
  1.1 GB to ~4,000 files / 548 MB, completing in ~33s instead of hanging.

The pre-update automatic backup shares this code path and was equally
at risk.

Adds a regression test that holds an EXCLUSIVE transaction in a separate
process and asserts _safe_copy_db returns False in bounded time.
2026-08-21 14:39:02 -07:00
Simon f9849c43a2 fix(backup): don't nest state-snapshots/ into full backups
`hermes backup` already skips `backups/` so a full zip never re-ships
earlier pre-update zips. `state-snapshots/` (written by `hermes backup
--quick`, `/snapshot create`, and the pre-update safety net) has the same
shape — every retained snapshot holds its own copy of state.db — but was
not in `_EXCLUDED_DIRS`, so a full backup shipped the DB once per
retained snapshot on top of the live one.

Two places hit this in practice:

- `hermes update` in `full` mode takes the quick snapshot *before* the
  full zip, so the pre-update zip always nests the snapshot it just made
  (state.db twice in every pre-update-*.zip).
- Any recurring `hermes backup --quick` (default keep=20) makes a daily
  `hermes backup` grow by roughly one compressed state.db per retained
  snapshot; a 750 MB state.db with two snapshots on disk pushed a daily
  zip from 1.8 GB to 2.3 GB.

Add `_QUICK_SNAPSHOTS_DIR` to `_EXCLUDED_DIRS` (moving the constant up
next to the exclusion rules so there is one source of truth). Both walk
sites and `_should_exclude` share the set, so `hermes backup`, the
pre-update zip and the auto-backup path all pick it up. Restoring
snapshots after a machine move was never the point of the full backup —
`profiles.py` already excludes `state-snapshots/` from `--clone-all` for
the same reason.

Tests: unit case next to the `backups/` one, plus two end-to-end cases
that use the real `create_quick_snapshot` producer and assert the zip
carries exactly one state.db (full backup and pre-update-order).
2026-08-21 14:39:02 -07:00
Teknium 1575116629 fix(update): pre-update snapshots now cover every profile, not just the invoking one (#66140)
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.

- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
  snapshot set, per-file 1GiB cap, and keep policy as the invoking
  profile (no partial tier, no new restore-coherence class), each into
  the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
  runs the #34600 cron-loss safety net per profile against its OWN
  snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
  profile's (best-effort, receipt-recorded); post-update cron restore
  extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
  behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
  jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
2026-08-21 13:01:35 -07:00
Teknium 0a8a4cdb3d fix(backup): friendly error on unwritable output path instead of raw traceback 2026-08-20 01:47:06 -07:00
fangliquanflq 076b8a5aa8 fix(backup): bound locked database snapshot waits 2026-08-16 02:00:50 -07:00
briandevans 66356fe24b fix(backup): drop setuid/setgid from the mode restored onto imported files
``_extract_member_atomically`` carries the replaced file's permissions
across the publish so that routing through mkstemp does not change what
the caller would otherwise have produced. But ``_preserve_file_mode``
returns ``stat.S_IMODE``, which is all twelve bits, and this restore is
deliberate on both sides of the replace: the mode is fchmod'd onto the
temp before ``atomic_replace`` and re-applied afterwards because chown
clears the elevated bits. So a target sitting at 0o4755 comes out of
``hermes import`` still at 0o4755 — with contents supplied by the zip.

That is a regression introduced by the atomic rewrite rather than a
pre-existing one. The overwrite it replaced was an in-place
``open(target, "wb")``, and an in-place write by a process without
CAP_FSETID has the elevated bits stripped by the kernel, so the old path
left 0o4755 as 0o755.

The blast radius is not limited to Hermes' own state: the ``_external/``
branch of ``run_import`` publishes members anywhere under ``$HOME``, and
this is the path that documents ``sudo`` use so ownership survives a
restore. An archive that happens to contain a member matching some
existing privileged file would take over the identity that file runs as.

Mask the two bits off the preserved mode. The masking happens once,
before the temp file is chmod'd, so there is no transient elevation
either. The sticky bit is kept — it is inert on a regular file. The
ordinary permission bits are unaffected, so the Docker/NAS installs the
preservation exists for still get their broader modes back.

This is the one write path in the repo where the bytes are untrusted;
the ``utils`` writers that preserve the full mode re-serialize content
the process itself produced, and are correct as they stand.
2026-08-14 21:47:23 -07:00
briandevans 60f86662e3 docs(backup): note that atomic_replace's cross-device fallback still truncates
The atomicity claim in _extract_member_atomically's docstring holds on the
os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses
shutil.copyfile and so opens the destination 'wb'. That is pre-existing
behaviour shared by every atomic writer in the repo, and it is reachable here
for a symlinked target whose real file lives on another filesystem. Scope the
docstring to what the helper actually guarantees instead of overstating it;
the fallback itself is a utils.atomic_replace change.
2026-08-14 21:47:23 -07:00
briandevans 1c3c1f4d71 fix(backup): preserve owner on atomic import writes and close the 0600 transit window
Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.

Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).

Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.

This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728a5 and 43fc86562; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.

Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
  need root; asserts chown fires once, with the captured owner, on the
  pre-existing file only (a newly created member has no prior owner).
  Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
  on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
  without the fix, 0o644 with it. Mutation-checked the same way.
2026-08-14 21:47:23 -07:00
briandevans e88c9f0ef2 fix(backup): restore import members atomically so a failed import can't erase config
`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.

Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.

`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.

This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
2026-08-14 21:47:23 -07:00
Teknium 715d26cdf4 feat: auto-install gateway service during setup and import
Users who install Hermes and then restore a backup (hermes import) ended
up with bot tokens and cron jobs fully registered but nothing running
them: the setup wizard's service-install prompt lived at the end of the
Messaging Platforms section, so skipping messaging (the normal case on a
box whose tokens arrive with the import afterward) skipped the service
entirely, and run_import never touched the service layer at all.

A platform-less gateway is already a supported mode (gateway/run.py runs
the cron scheduler and picks platforms up as tokens appear), so there is
no reason to gate the service on messaging config — or to ask at all.

- hermes_cli/gateway.py: new ensure_gateway_service() — prompt-free,
  never-raising install+start of the user-scope service (systemd /
  launchd / Scheduled Task), no-op in containers and on hosts without a
  service manager, refuses to pile onto conflicting user+system units.
- hermes_cli/setup.py: setup_gateway() service block now runs
  unconditionally (zero platforms included) and auto-installs instead of
  prompting; restart-on-config-change keeps its prompt. Quick-setup and
  migrated-config paths that skip the messaging section now call
  ensure_gateway_service() so they can no longer skip the service.
- hermes_cli/backup.py: run_import() ends by installing/starting the
  service when none is running, with a manual fallback hint on failure.
- tests: new tests/hermes_cli/test_ensure_gateway_service.py (9 cases)
  + 3 run_import wiring tests; existing backup tests get an autouse
  fixture so they never touch the host's real service manager.
2026-08-12 16:59:37 -07:00
kshitij 7289898494 refactor: consolidate five duplicate byte formatters into hermes_cli.sizefmt
Five modules each carried a private near-identical human-readable byte
formatter (backup._format_size, checkpoints._fmt_bytes,
doctor._human_bytes, context_references._human_bytes,
curator_backup.format_size). Three of them silently topped out at GB and
rendered a 1 TiB value as '1024.0 GB'. All five now alias one shared
format_bytes in hermes_cli/sizefmt.py (sibling of timefmt.py, same
zero-dependency rationale), keeping each module's established local name
so no caller churns.

Deliberately NOT migrated (behavior differs on purpose):
- session_recovery._format_bytes: binary suffixes (KiB/MiB/GiB)
- qqbot chunked_upload.format_size: '100.0 B' one-decimal style, pinned
  by its protocol tests

Net -33 production LOC before the new module; parity verified over a
16-value corpus against all five verbatim originals (only divergence:
the TB tier fix). Contract tests mutation-checked red-green.
2026-08-08 15:10:35 +05:30
Hao Wang aad8f7412c fix(backup): serialize and atomically publish snapshots 2026-08-03 23:48:55 +05:30
teknium1 fbd5e5772b fix(state): stop cancelling our own POSIX locks on live SQLite databases
CI / Check uv.lock (push) Has been cancelled
CI / Detect affected areas (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Docs Site (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Build&Test Docker image (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
CI / CI review comment (live) (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
Build Skills Index / build-index (push) Has been cancelled
Build Skills Index / trigger-deploy (push) Has been cancelled
`hermes sessions optimize` could corrupt state.db. Root cause is Hermes,
not the SQLite WAL-reset bug (#69784).

close() on ANY file descriptor for a SQLite database cancels every POSIX
advisory lock the process holds on that file, including a running VACUUM's
EXCLUSIVE lock (sqlite.org/howtocorrupt.html section 2.2). Hermes byte-probed
live databases in several hot paths: the zeroed-state.db detector runs on every
SessionDB construction (and the gateway builds those constantly), and kanban's
post-commit invariant check ran after every COMMIT. While VACUUM rewrote the
file, those probes dropped its lock and let other processes write into it.

A/B against the real code, only variable being the raw read:

  SQLite 3.50.4, VACUUM + concurrent writers, DELETE mode
    raw open/close during VACUUM   8 vacuums, 319 vacuum errors, 2/2 corrupt
    no raw read (control)        229 vacuums,   0 vacuum errors, 0/2 corrupt

  SQLite 3.53.1 (WAL-reset FIXED) reproduces identically: 2/2 corrupt.
  After this change: 0/4 corrupt, 0 vacuum errors.

Because the upgraded runtime corrupts too, replacing the embedded SQLite does
not fix this class; and because DELETE is where it reproduces, #70055's
"force DELETE on vulnerable builds" mitigation steered users into the failing
mode. That gate is reverted here: vulnerable builds get WAL again and still
warn so operators can upgrade.

- add hermes_cli/sqlite_safe_read.py: read page_count via PRAGMA over the
  existing connection instead of open()+seek(28); byte-level probes are
  restricted to before any connection exists and refused once one is live,
  with an explicit force= escape for offline artifacts (snapshots, archives)
- track live connections in SessionDB and kanban's connect so that guard is
  enforced rather than merely documented
- kanban's torn-extend check now only applies under a rollback journal; in WAL
  a committed page may still legitimately sit in the -wal file
- revert the force-DELETE WAL gate and update the tests that pinned it

Regression tests assert the behavioural contract (an external process stays
locked out across Hermes' inspection calls) and were verified to fail when the
old raw-open behaviour is restored.
2026-07-25 21:44:43 -07:00
Jasmine Naderi 6048696ed0 fix(state): cross-process quarantine lock + oversized DB pruning suppression (#68805)
Reviewer egilewski identified two recovery-loss paths:

Path A — quarantine race (hermes_state.py):
SessionDB checks state.db before quarantine_zeroed_state_db() without a
shared cross-process lock, and Path.rename() may replace an existing
destination. Two writable startups can therefore move the first
instance's newly created database over the .zeroed-*.bak, erase the
original damaged-file evidence, and replace the live database with
another empty one.

Fix: add a cross-process lock (msvcrt on Windows, fcntl on POSIX)
around quarantine_zeroed_state_db() with a 5s bounded timeout. Under
the lock: re-check is_zeroed_state_db (another process may have already
quarantined it and created a fresh DB), use a PID-suffixed unique
destination, and non-clobbering rename with counter fallback.

Path B — size-cap pruning gap (hermes_cli/backup.py):
_too_large() runs before failed-database tracking. With keep=1 and a
size cap, an oversized state.db is omitted while failed_dbs stays empty,
so automatic pruning deletes the older complete snapshot that may
contain the only recoverable database.

Fix: track oversized DB files in a new oversized_skipped list (both in
the directory walk and top-level file loop). The manifest now records
oversized_skipped. Pruning is suppressed when failed_dbs or
oversized_skipped is non-empty, preserving the older complete snapshot
as recovery source.

Tests:
- test_concurrent_quarantine_no_clobber: two threads racing on the
  same zeroed state.db — verifies quarantine backup survives with
  original bytes and live DB is valid.
- test_oversized_db_suppresses_pruning: keep=1 + oversized state.db
  verifies the older complete snapshot is not pruned.

All 33 tests pass (3 zeroed_state_db + 25 TestQuickSnapshot + 5
quarantine_forensic_logging).
2026-07-24 23:06:05 -07:00
Jasmine Naderi fc99b549ee fix(backup): skip prune when DB capture failed — preserve recovery source
Address #68805 review: when failed_dbs is non-empty, skip
_prune_quick_snapshots so the incomplete snapshot does not delete
the older snapshot containing the last good database. The updater's
keep=1 would otherwise evict the recovery source.
2026-07-24 23:06:05 -07:00
Jasmine Naderi ed5e41ddd6 fix(state): loud failed state.db snapshot + zeroed-file quarantine
Hardening for the Windows zeroed-state.db class (#68474):
- Surface critical stdout when pre-update/quick snapshot cannot copy a
  present *.db (was log-only; update still looked successful).
- Detect all-NUL SQLite header on SessionDB open, quarantine the bytes,
  and open a fresh DB with recovery guidance to state-snapshots.

Does not claim storage-stack root cause.
2026-07-24 23:06:05 -07:00
teknium1 01232e8e21 fix(update): stop hermes update stalling for minutes on a large state.db
The post-update state.db integrity guard called verify_sqlite_integrity()
with max_bytes=0, which disables the size ceiling and forces a full
PRAGMA integrity_check. That pragma walks every b-tree page in the file,
so its cost scales with database size — measured on a real 30 GB state.db:
143.5s with a cold page cache (worse under an update's memory pressure),
with zero output on screen. The update looks hung right after
"✓ Code updated!" and a CPU sits pegged.

Multi-GB session databases are normal for heavy users, so a
size-unbounded check is never an acceptable default on the update path.

- verify_sqlite_integrity(): max_bytes now defaults to
  DEFAULT_INTEGRITY_CHECK_MAX_BYTES (2 GiB) instead of 0. max_bytes=0
  remains the explicit opt-in for a full scan.
- The oversized path no longer degrades to a header-only check: it adds a
  constant-time structural probe (read-only open + schema_version +
  sqlite_master read) so the malformed-schema class is still caught, not
  just the #68474 zeroed-file signature.
- Drop the explicit max_bytes=0 at the post-update guard and in
  copy_db_and_verify() so both inherit the bounded default.

Measured on the reporter's real 30 GB state.db: 143.5s cold → 0.001s,
still valid=True. Corruption detection verified at multi-GB scale for
both classes (zeroed header, malformed schema) — both still fail closed.

Tests: default-is-bounded invariant, oversized probe catches malformed
schema, max_bytes=0 still forces the full check.
2026-07-24 21:39:32 -07:00
teknium1 a5f9ea2741 fix(update): correct integrity-guard bugs from #70553 salvage + tests
Follow-up fixes on top of the cherry-picked guard:
- verify_sqlite_integrity(): an oversized (max_bytes-exceeding) database
  now still fails on a zeroed/invalid header — previously valid=True was
  set before the header check, so a >1GiB zeroed state.db (the exact
  #68474 signature at 95MB scale) passed as valid.
- _run_pre_update_backup(): the guard referenced get_hermes_home before a
  later function-local 'from hermes_constants import get_hermes_home'
  shadowed it → UnboundLocalError swallowed by the snapshot try/except,
  silently disabling the post-snapshot check AND the snapshot-id output.
  Alias the import explicitly.
- Import _quick_snapshot_root where used (was NameError in all 3 guards).
- tests/hermes_cli/test_state_db_guard.py: real-SQLite E2E coverage —
  valid/zeroed/truncated files, oversized header gate, copy+verify
  roundtrip, snapshot restore flow, and the live _run_pre_update_backup
  path against a temp HERMES_HOME with mid-flight zeroing.
2026-07-24 15:59:32 -07:00
webtecnica d68e043ba7 fix(desktop,update): prevent silent state.db zeroing during Windows update (#68474)
Problem:
On Windows, state.db could be silently replaced with 95MB of null bytes
during a desktop update (v0.19.0). The pre-update snapshot was valid, but
the live file was destroyed and the update reported exit code 0, masking
the data loss. Sessions between the snapshot and the update were
irrecoverable.

Root cause analysis:
The update flow (Desktop Electron → hermes-setup.exe → hermes update)
kills the backend process tree via taskkill /T /F, then pauses Windows
gateways, creates a pre-update snapshot, runs git pull + pip install, and
resumes gateways. On Windows, a force-killed process holding state.db
(SQLite WAL mode) can leave the file open to races with antivirus/NTFS
filter drivers, or the gateway resume can encounter a partially-recovered
WAL state that results in a zeroed file — all while exit code 0 reports
success.

Fix — three layers of defense:

1. Emergency desktop-side backup (pre-flight):
   - New  function in Electron main.ts reads the
     SQLite header, logs it, and takes a timestamped emergency copy of
     state.db BEFORE the backend is killed or the updater is spawned.
     Runs in both the Tauri-updater path (Windows) and the in-app update
     path (Posix). Prunes to the 2 most recent emergency backups.

2. Pre-update integrity verification (Python CLI):
   - After  creates the pre-update snapshot,
      checks the LIVE state.db file (header +
     PRAGMA integrity_check). If corrupted, checks whether the snapshot
     copy is valid and warns the user. The update still proceeds because
     the snapshot is the recovery path.

3. Post-update auto-restore (Python CLI):
   - After the update completes (both git-pull and ZIP paths), verify
     state.db integrity. If corrupted/zeroed, automatically restore from
     the pre-update snapshot and re-verify. This catches the exact case
     where state.db was destroyed mid-update but the snapshot was valid.

New functions in hermes_cli/backup.py:
  - verify_sqlite_integrity(path, check_header, run_pragma, max_bytes)
    → Three-stage check: file size, SQLite header magic, PRAGMA
      integrity_check. Configurable max_bytes to avoid reading huge DBs.
  - copy_db_and_verify(src, dst)
    → Like _safe_copy_db() but verifies the destination after backup.

Fixes #68474
2026-07-24 15:59:32 -07:00
Teknium 2278f2cb7e fix(discord): harden reconnect message recovery
Route recovered messages through the live Discord ingress policy, preserve dedup and completion invariants, bound and retain the recovery ledger, and expose the opt-in config with docs and backup coverage.
2026-07-18 14:01:33 -07:00
teknium1 edfa4cd9b7 fix(cron): widen UTF-8 BOM tolerance to backup/curator jobs.json readers
Follow-up to the salvaged #66609 (4 primary readers) and #41604 (context
files): two more jobs.json readers rejected a BOM'd file —

- hermes_cli/backup.py _count_cron_jobs: a BOM made the count None,
  silently disabling the post-update cron-loss auto-restore safety net
- agent/curator_backup.py _backup_cron_jobs_into: BOM broke the job
  count (spurious parse_warning) and propagated the BOM into snapshots

Both now read utf-8-sig; curator snapshots are written BOM-free so
rollback restores a file load_jobs can read. AUTHOR_MAP entry added
for deacon-botdoctor.

Tests: BOM'd-live-file auto-restore + BOM'd snapshot count/BOM-free copy.
2026-07-18 02:31:20 -07:00
teknium1 abc22cdf1a fix(cron): harden execution attempt ledger 2026-07-17 04:58:35 -07:00
teknium1 332fbadd7b fix(backup): fail closed on sqlite snapshot errors 2026-07-17 04:55:19 -07:00
Teknium d0dcb9a5fd fix(update): consolidate pre-update backups into one gated mechanism (#65754)
hermes update ran TWO separate pre-update backup mechanisms: the
config-gated full zip (updates.pre_update_backup, default off) and an
unconditional quick state snapshot added for #15733 that ignored the
user's setting entirely. On a large state.db (observed: 24 GB) the
'cheap' snapshot silently added ~60s to every update and ate 24 GB of
disk in state-snapshots/.

Now there is ONE mechanism, gated by updates.pre_update_backup with
three modes:

- quick (new default): state snapshot of critical small files (pairing
  JSONs, cron jobs, config, auth, per-profile DBs). Files over 1 GiB
  are skipped with a warning so a bloated state.db can never stall the
  update again.
- full: the quick snapshot plus the HERMES_HOME zip (old 'true'
  behavior; --backup forces it for one run).
- off: nothing runs — an explicit opt-out now disables the quick
  snapshot too (--no-backup does the same per-run).

Legacy booleans are honored: true -> full, false -> off.

_run_pre_update_backup() now returns the quick-snapshot id so the
post-update cron-jobs restore safety net (#34600) keeps working; the
snapshot moved from the post-fetch site to the pre-mutation site,
which also covers the zip-fallback update path it previously missed.
2026-07-16 08:47:25 -07:00
Teknium 55e3ee1ab8 fix: remove dead f-string prefixes via ruff F541 (216 sites) (#52336)
ruff check --fix --select F541 . on current main. Pure prefix removals;
adjacent-string concatenations keep the f only on interpolating fragments.
No string content or live placeholder altered.
2026-07-05 13:42:46 -07:00
0xDevNinja 9ef49cd78f fix(backup): include projects.db, kanban boards, and sibling stores in pre-update snapshot (#52889)
projects.db (per-profile project store) and kanban.db were missing from
_QUICK_STATE_FILES, so the pre-update quick snapshot never backed them up.
On a desktop upgrade, when the update flow removes/replaces the file and the
post-update schema-init re-creates an empty one, all user-created projects,
folder mappings, the active-project pointer, kanban board bindings, and tasks
vanish silently — no error.

Add the per-profile user-created stores to the snapshot set:
- projects.db               — project store
- response_store.db         — gateway conversation history / tool payloads (WAL)
- memory_store.db           — holographic memory facts/entities (WAL)
- verification_evidence.db  — agent verification audit trail
- kanban.db                 — default board (back-compat <root>/kanban.db)
- kanban/boards             — non-default boards (<root>/kanban/boards/<slug>/kanban.db
                              + metadata); workspaces/ and attachments/ subtrees
                              are skipped as large + regenerable.

Also: the directory-branch of create_quick_snapshot now routes *.db through the
WAL-safe _safe_copy_db (SQLite backup() API), matching the top-level file path —
previously a non-default board DB with an open WAL could be copied inconsistently.

Salvaged from #52930 by @0xDevNinja (authorship preserved via cherry-pick).
On top of the original (which covered only projects.db + the default kanban.db),
this adds: non-default-board coverage, the three sibling per-profile DBs that
meet the same upgrade-wipe criteria, WAL-safe directory copies, and a
workspaces/attachments skip to avoid snapshot bloat (×20 retained). 8 tests,
all mutation-verified; E2E verified snapshot→wipe→restore preserves all six
store types on the real code path.

Closes #52889. Supersedes #52930.
2026-06-26 19:23:33 +05:30
liuhao1024 56cf517ccd fix(cron): detect partial job loss in restore_cron_jobs_if_emptied (#52144)
The desktop scheduler can overwrite cron/jobs.json with its own small
set of internally-tracked crons after an update/restart, causing
partial loss of tool-created cron jobs. The previous guard only
checked for total loss (live_count == 0), missing the case where
live_count > 0 but less than the pre-update snapshot count.

Compare live_count against snap_count instead of checking for zero,
so both total loss (0 vs N) and partial loss (1 vs 19) trigger
restoration.

Salvaged from #52161 by @liuhao1024.

Closes #52144
2026-06-25 18:49:18 -07:00
memosr ae46699905 fix(security): validate snapshot_id and file paths in restore_quick_snapshot to prevent path traversal 2026-06-21 12:44:22 -07:00
Teknium 587b5b9ac2 fix(backup): capture memory-provider state stored outside HERMES_HOME (#50325)
hermes backup only walks HERMES_HOME, so memory providers that keep
config/credentials in home-anchored dotdirs (honcho -> ~/.honcho,
hindsight -> ~/.hindsight, openviking -> ~/.openviking) lost that data
across a backup/import cycle — the peer IDs, session pairings, and API
keys never made it into the archive.

Add an optional MemoryProvider.backup_paths() hook (default []). The
active provider declares its external paths; backup resolves them from
config only (no init, no network), archives the ones under the home dir
into a reserved _external/ subtree encoded relative to home, and import
restores them to their original location with a home-anchored traversal
guard and 0600 on credential-shaped files. Paths outside home are
skipped as non-portable.

honcho, hindsight, and openviking override the hook. E2E-validated full
backup->import cycle plus 7 new tests.
2026-06-21 12:03:46 -07:00
xxxigm e738c08336 fix(backup): exclude regeneratable dependency and cache dirs
`hermes backup` walked every file under HERMES_HOME, excluding only
hermes-agent / node_modules / __pycache__ / backups / checkpoints. Python
dependency trees (plugin and MCP-server venvs, site-packages) and pip/uv
tool caches that live under HERMES_HOME were swept in file-by-file,
ballooning a backup to hundreds of thousands of entries that crawl for
hours — the reported "backup stuck for days / 426543 files" symptom.

Add the canonical regeneratable-dir names (.venv, venv, site-packages,
.tox, .nox, .pytest_cache, .mypy_cache, .ruff_cache — mirroring
agent.skill_utils.EXCLUDED_SKILL_DIRS) plus .cache to the backup's
exclusion set, used by both run_backup and the pre-update/pre-migration
_write_full_zip_backup. .archive is intentionally left in so the curator's
restorable archived skills still get backed up.

Tests cover each new dir name (excluded at any depth), that .archive and
cache-resembling files are kept, and an integration check that a planted
venv/site-packages/cache is pruned from the actual backup zip while
skills/config survive.
2026-06-19 14:37:41 +05:30
Ben Barclay 9c3c5da356 fix(backup): hermes import never overwrites volatile gateway runtime state (NS-501) (#48243)
Importing a backup wrote every file from the zip over the target home
wholesale. On a hosted instance this clobbered gateway_state.json with the
source machine's last recorded run/desired state — driving the container-boot
reconciler (container_boot._read_desired_state, which only auto-starts a
gateway whose state is "running") off stale/foreign state and leaving the
gateway stuck "starting", disconnected from the Nous portal.

Add _IMPORT_SKIP_NAMES (gateway_state.json, gateway.pid, cron.pid,
gateway.lock, processes.json) and skip them by basename in run_import, so both
the root profile and named profiles preserve the target's own runtime state.
This mirrors what container_boot._STALE_RUNTIME_FILES already sweeps on every
container boot, and protects against older backups that predate the
backup-side exclusions. The import summary reports which files were preserved.

This is the second half of NS-501 (filed separately as NS-508): the upload
502 was fixed in #47663; this fixes the import-breaks-the-instance half.
2026-06-18 15:27:45 +10:00
Teknium 0d82060c74 fix: harden WhatsApp target alias salvage
Add a parser-only routing regression that proves raw WhatsApp group JIDs bypass channel-directory resolution and home-channel fallback, include channel_aliases.json in quick state snapshots, harden malformed alias handling, and map Keiron McCammon for release attribution.
2026-06-15 05:51:47 -07:00
kshitijk4poor ed2b9e43c8 fix(backup): stage SQLite snapshots beside output zip in pre-update path too
The pre-update / pre-migration backup path (_write_full_zip_backup) had the
same /tmp staging bug as run_backup: a small tmpfs at the default tempfile
location silently drops large *.db files from the archive. Route its SQLite
staging temp files to the output zip's directory as well, and add regression
tests (mutation-verified) for both staging paths.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-06-11 12:45:40 +05:30
liuhao1024 dd40600e0a fix(backup): stage SQLite snapshots alongside output zip and stop excluding nested hermes-agent skill dirs
Two bugs in the backup routine:

1. SQLite safe-copy used tempfile.NamedTemporaryFile() which defaults to
   the system temp directory (/tmp).  When /tmp is a small tmpfs and the
   database is large, the copy silently fails and the resulting zip is
   missing state.db, kanban.db, and response_store.db.

   Fix: pass dir=out_path.parent so the temp file is staged alongside the
   output zip on the same filesystem.

2. _EXCLUDED_DIRS contained "hermes-agent" which matched at ANY path
   depth, accidentally excluding the Hermes Agent skill directory at
   skills/autonomous-ai-agents/hermes-agent/.

   Fix: special-case "hermes-agent" to only match when it is the first
   path component (the root-level code checkout).  All other excluded dir
   names continue to match at any depth.

Regression tests added for both fixes.
2026-06-11 12:43:39 +05:30
Bartok9 3845d86b93 fix(cron): restore jobs.json emptied by config migration on update
Config-version migrations have been observed to leave cron/jobs.json
valid-but-empty after `hermes update`, silently dropping every scheduled
job (#34600). The existing malformed-shape guards in cron/jobs.py don't
catch this because {"jobs": []} is valid JSON.

Add restore_cron_jobs_if_emptied() as a post-migration safety net: if the
live cron/jobs.json now has zero jobs while the pre-update snapshot held
one or more, restore the snapshot copy in place and warn loudly. The
check is conservative — it only restores on unambiguous evidence of loss
(snapshot had jobs, live file readable-and-empty), so a user who genuinely
cleared their jobs is never second-guessed and an unreadable live file is
left untouched so real corruption still surfaces.

Wired into _cmd_update_impl after migrate_config(), reusing the existing
pre-update quick snapshot (which already captures cron/jobs.json).

Closes #34600
2026-05-29 13:22:54 -07:00
Aditya Rajesh Gadgil 031983bbf8 fix: limit pre-update state snapshots 2026-05-28 02:45:25 -07:00
nguyen binh 0d55315c36 fix(backup): skip symlinked files in zip archives (#25289) 2026-05-25 05:07:52 -07:00
kshitij 2ec8d2b42f chore: ruff auto-fix PLR6201 — tuple → set in membership tests (#23937)
Replace  with  for all literal-tuple
membership tests. Set lookup is O(1) vs O(n) for tuple — consistent
micro-optimization across the codebase.

608 instances fixed via `ruff --fix --unsafe-fixes`, 0 remaining.
133 files, +626/-626 (net zero).
2026-05-11 11:13:25 -07:00