Commit Graph

41 Commits

Author SHA1 Message Date
jango 5a2f512390 fix(cli): exit non-zero from sessions repair --check-only on an unhealthy store
`hermes sessions repair --check-only` printed the corruption reason and
exited 0, so scripts and the console wrapper gating on the status read a
broken state.db as healthy. Return 1 from the CLI handler (main.py already
sys.exits a truthy return) and from the console handler, where
`_capture_output` turns the status into a ConsoleCommandError carrying the
printed reason.

Salvaged from PR #103321 (the check-only reporting part only; the probe
rewrite and connection-tracking changes were not taken). Refs #63386.
2026-09-11 06:37:27 -07:00
teknium1 496eb13bd7 fix(state): one corrupt timestamp row no longer kills sessions list, export or insights
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).

Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.

Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.

Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
2026-09-11 06:24:54 -07:00
Teknium eeb7671e69 simplify(compat): hermes_cli small facades — drop 7 re-exports/aliases (+relay_runtime alias module), repoint 12 callers/tests 2026-09-03 13:05:57 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0fc204042d fix(integration): sessions CLI — close db via try/finally, complete test doubles
cmd_sessions used 'with db:' which breaks test doubles lacking the context
manager protocol (13 reds in test_sessions_pin/delete/export). Restore the
explicit try/finally db.close() (same semantics for real SessionDB). Add
get_session/count_prune_matches to the FakeDB doubles in test_sessions_delete
instead of re-adding getattr guards to production code.
2026-09-03 02:06:29 -07:00
Teknium c88d60551e refactor(hermes_cli): group module-level constants in session modules 2026-09-03 00:06:12 -07:00
Teknium 92cd2a76ef refactor(hermes_cli): _cmd_recover report writing and delete pinned-note inline 2026-09-02 23:57:08 -07:00
Teknium f1b080b91d refactor(hermes_cli): markdown single-export summary line 2026-09-02 23:52:56 -07:00
Teknium 10cd4dcf20 refactor(hermes_cli): _cmd_repair import hoisting and merged prints 2026-09-02 23:49:49 -07:00
Teknium 251a82b46e refactor(hermes_cli): sessions_cmd pin/retitle/prune preview tightening 2026-09-02 23:37:04 -07:00
Teknium 3cc998a4b8 refactor(hermes_cli): table-driven _export_flat replaces three parallel single-file exporters 2026-09-02 23:32:59 -07:00
Teknium 2b848064a6 refactor(hermes_cli): sessions_cmd export collector and pinned-note tightening 2026-09-02 23:15:10 -07:00
Teknium 7d675b9132 refactor(hermes_cli): sessions_cmd storage/dispatch tail compaction 2026-09-02 23:10:33 -07:00
Teknium 68d746df2a refactor(hermes_cli): sessions_cmd list/recover minor tightening 2026-09-02 23:02:46 -07:00
Teknium 679b303ba1 refactor(hermes_cli): drop intra-function blank separators in session modules (AST-identical) 2026-09-02 22:55:35 -07:00
Teknium df39de8738 refactor(hermes_cli): sessions_cmd export/prune flow tightening 2026-09-02 22:48:19 -07:00
Teknium cfcafb5b23 refactor(hermes_cli): merge sessions_cmd filter-arg tables into one _FILTER_ARGS predicate 2026-09-02 22:42:47 -07:00
Teknium 9dfdbe7190 refactor(hermes_cli): sessions_cmd second pass — merged printer calls, _export_dir helper, inverted guards 2026-09-02 22:19:58 -07:00
Teknium cec7daa07d refactor(hermes_cli): pack exploded signatures/calls in session recovery modules (AST-identical) 2026-09-02 22:09:28 -07:00
Teknium 542819c468 refactor(hermes_cli): compact sessions_cmd docstrings, comments and printer bodies 2026-09-02 21:39:50 -07:00
Teknium 86d111d77c refactor(hermes_cli): collapse defensive layers in sessions_cmd, join bracket groups in session recovery modules 2026-09-02 21:12:14 -07:00
Teknium 5a9970ad4b refactor(cli/sessions_cmd): lift recover progress printer + verdict out of _cmd_recover 2026-09-02 16:04:45 -07:00
Teknium 1d2b39305f refactor(cli/sessions_cmd): unify output writer, compact browser class, tighten comments 2026-09-02 15:57:43 -07:00
Teknium 1172735295 refactor(cli/sessions_cmd): split cmd_sessions into per-subcommand handlers + dispatch dicts; extract browse picker to sessions_cmd_browse 2026-09-02 15:45:35 -07:00
Teknium ed8aa5e3e4 refactor(cli): move session browse/status/relative-time helpers from main.py into sessions_cmd.py
_relative_time, _session_status_tag, _annotate_session_statuses,
_session_browse_picker and _size_delta_label (410 LOC) replace the
call-time delegating wrappers in sessions_cmd; main.py re-exports them via
_LAZY_COMMAND_EXPORTS so hermes_cli.main.<name> imports/patches still work.
2026-09-02 13:29:44 -07:00
Teknium aca40d1d63 fix(sessions): error paths return non-zero exit codes (delete/rename/prune/import) 2026-08-20 02:06:05 -07:00
Teknium 76653a8eba fix(sessions): prune/archive spare pinned sessions by default (data loss) 2026-08-20 01:46:20 -07:00
Teknium 35598d8e8e Inspired by Perplexity Computer: sessions pin/unpin/pinned CLI (#52955)
Perplexity Computer's July update let its agent manage sessions
conversationally from any surface — pin, archive, rename, fork — treating
session organization as operational infrastructure rather than a GUI
nicety. Hermes already has the durable pinned flag in state.db (Desktop
sidebar writes it; auto-archive honors it), but no CLI access existed:
GUI-only management was a single point of failure and blocked scripting
(issue #52955).

- hermes sessions pin <id...> / unpin <id...>: set/clear the durable keep
  flag via SessionDB.set_session_pinned (whole compression lineage,
  prefix resolution, multi-id, exit 1 on any miss)
- hermes sessions pinned [--json]: list all pinned conversations via the
  include_pinned back-fill (old pins can't fall off a paging window);
  --json enables backup/restore scripting
- docs: user-guide/sessions.md section
- tests: 6 tests covering prefix resolution, multi-id partial failure,
  pinned-only filtering, JSON shape, empty hint
2026-08-16 22:09:17 -07:00
fangliquanflq 0a42bc7113 fix(sessions): align prune filter derivation 2026-08-16 01:55:25 -07:00
fangliquanflq 0b8a09759c fix(sessions): address prune skip review notes 2026-08-16 01:55:25 -07:00
fangliquanflq 29dfbf2d6a fix(sessions): surface open sessions skipped by prune 2026-08-16 01:55:25 -07:00
Teknium 933ef69470 feat: session picker lifecycle status + delete 2026-08-15 18:15:22 -07:00
Teknium 04c61f2949 feat: import and resume Claude Code / Codex CLI sessions 2026-08-15 18:15:13 -07:00
joaomarcos 174ce8770d fix(state): arm the live-DB guard by process ancestry, not env alone
Production `state.db` files accumulate zero-message "open" gateway session
rows carrying test-fixture identities (`chat-1` / `user-1` / `wx-chat`), with
matching `gateway_routing` scopes pointing at `pytest-of-*` temp directories.

The escape is structural. Hermetic isolation rides entirely on the process
environment: `HERMES_HOME` says *where* to write, `PYTEST_CURRENT_TEST` /
`PYTEST_VERSION` say *whether the guard is armed*. Both travel in the same
carrier, so a child spawned with a rebuilt environment loses them together —
it resolves the developer's real `state.db` *and* silences the only check
that would have stopped it, in one step. The guard is a no-op in precisely
the situation it was written for.

Back the env probe with process ancestry, which survives an env rebuild:

* `_process_looks_like_pytest()` matches a pytest launcher by argv token
  basename, so `/tmp/pytest-of-dev/...` paths in real argv cannot
  false-positive, and an unreadable process is never assumed to be a test.
* `_has_pytest_ancestor()` walks parents via psutil, memoised, and fails
  open when psutil is unavailable — a real `hermes` run pays for at most
  one walk and keeps the previous behaviour if the walk errors.
* `_in_test_context()` checks env first (two dict lookups, covers the
  in-process case) and only then ancestry.

`_STATE_DB_GUARD_BYPASS` is a module global and cannot cross a process
boundary, so ancestry-armed children would have had no way to opt out at
all; `HERMES_STATE_DB_GUARD_BYPASS=1` is the env-carried twin.

Also sweeps the rows already written. Bulk prune/archive cannot reach them:
their shared selector is pinned to `ended_at IS NOT NULL` so a live session
is never picked, which permanently excludes every never-closed row. Adds a
narrower selector — keyed, still open, and with no messages, tokens, tool
calls, API calls, activity or title — behind
`hermes sessions prune --never-active` (default floor 30 days, honours
--dry-run/--yes). Routing entries naming a deleted row go with it, so the
gateway is never left resuming a session id that no longer exists; `pinned`
and `archived` rows are excluded as explicit user intent.

Closes #82770
2026-08-15 02:20:13 -07:00
RelaxJonh 0d91ab8889 fix: close leaked SessionDB connections on /insights and sessions-repair exception paths (#83226)
Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command - db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair - SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Salvage note: the original PR (#83237) also added a __del__ safety-net
finalizer to SessionDB; review showed the atexit hook registered by
queue_token_counts() strongly retains the instance, so the finalizer never
fires for the leak class it claimed to cover. Dropped here in favor of the
deterministic constructor-finally ownership repair salvaged from #83620.
2026-08-14 21:41:26 -07:00
Teknium 6dad74596e fix(sessions): recover budget exhaustion + lost_and_found last-resort lane
Fixes #80205: when one ordered rowid-edge probe failed,
_salvage_rowid_bounds() substituted the whole SQLite rowid domain and
_copy_table_salvage() burned the 10,000-query budget bisecting a
synthetic tail that could not contain rows, silently omitting readable
boundary rows (field case: message 76882 of 76882). Two-part fix:

* _probe_populated_edge(): gallop outward from the surviving edge with
  doubling offsets; a clean 'no rows beyond X' probe caps the domain in
  O(log range) queries instead of exhausting the budget on it.
* exact-key singleton salvage: a one-row range scan must advance the
  cursor past the hit into the damaged sibling page to prove exhaustion,
  which discards the already-produced row; 'WHERE rowid = ?' stops at
  the hit, recovering the boundary row exactly like sqlite3 .recover.
* the strict-path refusal now points users at --allow-partial.

New last-resort lane for --allow-partial when the sessions/messages
table schemas themselves are unreadable (previously a hard refusal even
though page-level salvage recovers the rows fine). If a sqlite3 CLI is
on PATH, shell out to '.recover --ignore-freelist' into a scratch
lost_and_found DB, then heuristically map rows back into a fresh
SessionDB-schema database (hermes_cli/session_lost_and_found.py):
classification keyed on nfield counts + sentinel columns (session ids
matching ^\d{8}_\d{6}_, roles in user/assistant/tool/system, known
source strings), covering the current 54-col sessions layout, the
52-col historical layout, a 14-col legacy identity-only salvage,
rowid-alias messages rows and 18-col session_model_usage rows. Missing
parent sessions are stubbed (children are never deleted for FK
cleanup), FTS is rebuilt at the end, and output is labeled BEST-EFFORT
everywhere. Without the CLI the error names the sqlite3 requirement
with actionable guidance. Mirrors a successful manual recovery of a
real corrupt state.db (2026-08-12), and this lane was validated against
that preserved file: 32 sessions / 7 messages / 4 usage rows mapped,
integrity_check ok, opens via SessionDB.

Also fixes #72291: the source-fingerprint 'bundle changed while it was
being copied' error now enumerates that the parent interactive CLI
session itself counts as a Hermes process and suggests a fresh shell or
an immutable snapshot.

Tests use real physical page corruption (flipped b-tree/schema header
bytes), skip the CLI-dependent path cleanly when sqlite3 is absent, and
keep the mapper unit tests binary-independent via a synthetic
lost_and_found DB. Sabotage-verified: reverting the fixes makes the
regression tests fail with the exact field failure shape.
2026-08-12 19:43:47 -07:00
joaomarcos c790ed2a5d fix(state): recover gateway sessions stranded without a routing identity
When state.db's write path fails (corrupt FTS, or a crash landing between
routing publication and row creation), the live gateway conversation can end
up in a session row that never received its identity columns: session_key,
chat_id, chat_type and origin_json are all NULL. In-memory routing hides the
damage for as long as the gateway stays up. After a restart the chat is
resolved from the DB, and find_latest_gateway_session_for_peer cannot see
that row — both of its queries match on the very columns it lacks — so the
chat resumes the last keyed sibling instead, days older. The messages were
never lost, only unreachable.

Hardening the write side cannot reach a row that is already damaged, so add
the offline repair path the tracking issue asks for:

- SessionDB.find_orphaned_gateway_sessions() reports message-bearing rows
  with no session_key, and names the predecessor each one continues only
  when the evidence is unambiguous — a recorded parent_session_id
  ("lineage"), or exactly one keyed row of the same source and compatible
  user_id that fell quiet within 15 minutes of the orphan's start
  ("contiguity"). Contested pairs are reported with a reason and left alone:
  a wrong adoption would splice one person's conversation into another
  person's chat. Branch, delegate and tool rows are excluded — they are
  unkeyed by design, not by damage.
- SessionDB.adopt_orphaned_gateway_session() stamps the orphan from the
  predecessor (never overwriting a column that already has a value), records
  the lineage, and retires the predecessor under end_reason
  'superseded_by_repair' — a reason recovery does not treat as resumable, so
  the repaired row wins the chat from then on. The pair is re-verified inside
  the write transaction, making a concurrent heal a no-op rather than a
  conflicting write.
- `hermes sessions repair-routing` drives both. It reports without touching
  the database; --apply confirms first and warns that a running gateway
  still holds the old mapping in memory.

Refs #82616.
2026-08-09 14:06:06 -07:00
Brooklyn Nicholson f726090d48 feat(sessions): name a session the moment it starts
Titling fired on the first response, so a session sat unnamed for the whole
opening turn - p50 151s, p90 1212s across real sessions, because a turn is
tool calls, not one round-trip. A turn that failed or was interrupted never
got a title at all. Four surfaces each carried their own copy of the call.

Move it into the shared turn prologue and split it in two: a deterministic
title derived from the user's opening message, written inline before the
model runs, then one small-model call that upgrades it. The response is
constrained to a JSON object so there is no preamble to strip, and control
wrappers are stripped rather than refused, so a slash command titles as
what the user asked for instead of the command itself.
2026-08-08 17:07:21 -05:00
joaomarcos e18c040c3d fix(cli): back up state.db before clean-markers writes by default
purge_stale_tool_call_markers ran a permanent, irreversible UPDATE with
no backup — inconsistent with repair_state_db_schema's backup-by-default
convention for destructive state.db operations elsewhere in this file.

Take a full snapshot via VACUUM INTO (safe against a live connection,
unlike the raw-copy _backup_db_file used for malformed-schema repair)
before the write, timestamped beside state.db. Skipped when dry_run or
when there's nothing to change. Add --no-backup to `hermes sessions
clean-markers`, mirroring `sessions repair`.

Verified end-to-end: the CLI run against a real temp state.db produces
the backup file before printing the cleared-row count.
2026-08-04 11:26:15 +05:30
joaomarcos e1a2739692 feat(cli): add sessions clean-markers to permanently purge stale tool-call markers (#78148)
The load-on-read repair (_strip_stale_tool_call_markers) fixes affected
sessions in memory on every resume, but never touches the DB — long-lived
sessions re-scan and re-repair the same rows on every load, and the
contaminated bytes stay in state.db (and any backup/cache snapshot of it)
indefinitely.

Add SessionDB.purge_stale_tool_call_markers(dry_run=False): a one-time,
idempotent UPDATE that permanently blanks the content column on affected
rows. Only content is touched — tool_calls and every other column are
left untouched, so provider tool_call/tool_result pairing survives.
dry_run reads through the no-lock read path and never writes.

Wire it up as `hermes sessions clean-markers [--dry-run]`, mirroring the
existing optimize/repair subcommands. Verified end-to-end against a real
temp state.db: dry-run reports the row without writing, the real run
clears it and preserves tool_calls, and a second run is a no-op.
2026-08-04 11:26:15 +05:30
teknium1 0e7c4018f7 refactor: hoist cmd_sessions out of main() into sessions_cmd.py 2026-07-29 10:59:54 -07:00