Commit Graph

49 Commits

Author SHA1 Message Date
Teknium 68a6945ae8 fix: preserve per-match context decode failure isolation 2026-09-07 06:04:23 -07:00
Teknium 7e73ef4676 test: preserve batched context ties and duplicate hits 2026-09-07 06:04:23 -07:00
Teknium 20af238599 perf: batch search context using indexed neighbor seeks
Keep projections query-free and timestamp/id neighbor ordering. Bound parameter batches at 500 and avoid scanning complete sessions with LAG/LEAD. Based on the N+1 analysis in #104296; no additional YAML cache or durability/freshness changes.

Co-authored-by: DevvGwardo <25094504+DevvGwardo@users.noreply.github.com>
2026-09-07 06:04:23 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 2b55ded1ac perf(state): keep delegate-child transcripts out of the trigram FTS index (schema v30)
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.

Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.

The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
2026-09-03 02:35:37 -07:00
Andrew Wikel 46bcad0c24 fix(state): finalize empty FTS rebuild markers
Clear zero-row rebuild markers so empty databases can finish teardown and
stamp the new layout. Extend the migration regression through close/reopen
recovery for both empty and populated databases.
2026-09-03 11:35:13 +05:30
Andrew Wikel 5a7bee0fa8 fix(state): exclude tool calls from trigram FTS
Keep structured tool_calls searchable through the standard FTS index while
removing their repetitive JSON from the trigram projection. Reuse the
existing optimize-storage rebuild path for deployed v1 layouts.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-03 11:35:13 +05:30
Teknium cb6cc64700 refactor(state): search/schema/registry/portability/telegram/usage/titles — inline single-use helpers, contextlib.suppress ladders, pack wrappers around unchanged SQL literals 2026-09-02 21:57:21 -07:00
Teknium c1620901ae refactor(hermes_state): AST-neutral packing of state mixins (120 cols) 2026-09-02 19:10:31 -07:00
Teknium 561d1a6f8b refactor(hermes_state_search): regex quoted-phrase protection (fuzz-verified), drop unreferenced cjk aliases 2026-09-02 18:54:49 -07:00
Teknium 6da196e39a refactor(hermes_state_search): env_float threshold, collapse rebuild-status/teardown guards, compact docstrings 2026-09-02 18:45:38 -07:00
Teknium d2614f435e refactor(hermes_state): AST-neutral line packing across state mixins 2026-09-02 18:30:39 -07:00
Teknium d7fe7a47fd refactor(hermes_state): drop constant-false v10 trigram branch, shared search SELECT builder 2026-09-02 18:28:57 -07:00
Teknium 58accf5782 refactor(hermes_state_search): shared cjk/token helpers, reuse _row_to_message_dict, compact docstrings 2026-09-02 18:13:31 -07:00
Teknium eb8d628c97 refactor(hermes_state): restore WHY comments dropped by round-2 sub-branches
Comment/docstring-only (AST-identical): surrogate-scrub rationale, persisted
marker stripping invariant, generation counter upgrade semantics, CJK marker
empty-vs-populated rule, WAL 0-page ordering precondition, repair backup
live-connection case, telegram topic delete precondition, mixed-mode
corruption definition, and similar.
2026-09-02 16:48:30 -07:00
fangliquanflq ea65fcd980 perf(state): exclude cron sessions from trigram FTS 2026-09-03 05:08:22 +05:30
Teknium 3a9940fd6f refactor(state): hoist context-window SQL to a module constant 2026-09-02 16:37:20 -07:00
fangliquanflq 57162d0cc1 fix(state): bound FTS indexing for large tool results 2026-09-03 05:03:04 +05:30
Teknium 96508304ec refactor(state): trash-teardown drop helper; simplify deferral-record parsing 2026-09-02 16:31:45 -07:00
Teknium d86ebe66c0 refactor(state): CJK range table; shared _coerce_or for import numeric coercion 2026-09-02 16:29:12 -07:00
Teknium 6983794cb7 refactor(state): lift optimize_fts_storage vacuum/settle phases into helpers 2026-09-02 16:24:28 -07:00
Teknium a138533157 refactor(state): fold docstring closers (whitespace only) 2026-09-02 16:21:39 -07:00
Teknium b422113061 refactor(state): compact long rationale comments (all rules kept) 2026-09-02 16:20:35 -07:00
Teknium 5340109e61 refactor(state): compact verbose docstrings (rules/invariants kept) 2026-09-02 16:14:33 -07:00
Teknium 95714f6d93 refactor(state): AST-neutral line packing; derive v22 session_model_usage DDL from the heal DDL 2026-09-02 16:13:05 -07:00
nftpoetrist c7429f60ca fix(state): close the SessionDB lock gate's blind spot on its own mixin files
test_no_locked_readers_gate.py (#97676) parses hermes_state.py's SessionDB
class body with ast and flags any method that holds the writer lock
around a pure-read query — Pattern C, where every concurrent turn's
persistence convoys behind an unrelated read. #97676 converted 39 such
methods and closed detection blind spots for alias/variable-SQL readers.

But SessionDB is declared as
`class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)`,
and the gate only ever opened hermes_state.py — it never parsed the three
mixin files those base classes are defined in, so a locked reader
declared there was structurally invisible to the scanner regardless of
how good the alias/variable-SQL detection got.

Applying the gate's exact scanning logic to the three mixin files
directly turns up 9 genuine pure-read methods still holding the writer
lock, none in #97676's converted list:

- hermes_state_search.py: _fts_teardown_trash_step, fts_optimize_available,
  optimize_fts_storage, list_recent_user_messages
- hermes_state_portability.py: distinct_session_cwds, list_cron_job_runs,
  _get_session_rich_rows_batch, list_skill_scaffolded_sessions,
  get_first_assistant_text

_get_session_rich_rows_batch is a hot path: it backs list_sessions_rich's
compression-tip resolution and the web server's session-search hydration
across every gateway install — its own docstring already claims "same
read-your-writes guarantee as list_sessions_rich", but list_sessions_rich
was already using _read_ctx() (its guarantee comes from flush_token_counts()
before the read, not from holding the writer lock) while this method's
implementation never caught up to match.

Converted all 9 to `with self._read_ctx() as conn:`, the exact pattern
#97676 used, verified each is a genuine pure read with no hidden writes
by tracing every helper call it makes.

Extended the gate itself (_ALL_STATE_SOURCES) to scan all three mixin
files under their own class names, plus hermes_state.py, so this blind
spot can't silently reopen. Added a regression test
(test_scan_all_state_sources_visits_every_mixin_file) that plants a
synthetic violation in a mixin-shaped file and asserts the scanner still
catches it — a change that reverts the file list back to one file passes
the existing sabotage test but fails this one.

Mutation-verified: with the gate's new scope but the old (unconverted)
mixin sources, test_no_locked_pure_readers fails and names all 9 real
violations with correct file/line. Restored the fix; it passes clean.
2026-09-03 04:29:06 +05:30
Teknium 31729ddcff refactor(state): unify search routes in hermes_state_search (_match_rows/_search_cjk, dispatch table for sort) 2026-09-02 15:46:58 -07:00
Teknium d15c61b5dc refactor(state): split SessionDB into domain mixins and free-function modules; unify SQL boilerplate
hermes_state.py 17,220 -> 6,442 LOC. Behavior-neutral: every moved body is
AST-identical to the original, verified per extraction.

SessionDB core
- _write_sql / _write_rowcount / _read_one / _read_all replace ~120 copies of
  the `def _do(conn): conn.execute(...)` + `_execute_write(_do)` and
  `with self._read_ctx() as conn: row = conn.execute(...).fetchone()` shapes.
- _set_lineage_column replaces four copies of the recursive compression-lineage
  UPDATE (archived / pinned / hidden / last_read_at).
- _read_session_number unifies the three compression counter readers.
- Dead (zero refs repo-wide): restore_rewound, delete_gateway_routing_entries,
  _is_duplicate_replayed_user_message, SessionPortabilityMixin.get_first_assistant_text.

New mixins bound onto SessionDB via the MRO (logger name stays "hermes_state"):
  hermes_state_messages    SessionMessagesMixin       48 methods
  hermes_state_compression SessionCompressionMixin    30
  hermes_state_gateway     SessionGatewayMixin        26
  hermes_state_maintenance SessionMaintenanceMixin    13
  hermes_state_usage       SessionUsageMixin          12
  hermes_state_titles      SessionTitlesMixin         13
  hermes_state_telegram    SessionTelegramTopicsMixin 11
Origin-internal symbols resolve through a lazy `from hermes_state import ...`
inside the few methods that need them (no import cycle).

New free-function modules, every name re-imported into hermes_state so
`hermes_state.<name>` (and test monkeypatches on it) keep working; intra-module
calls to patched helpers go through the lazy origin import:
  hermes_state_repair   repair/backup/preflight (43 defs)
  hermes_state_wal      journal-mode / PRAGMA policy (33 defs)
  hermes_state_dbfile   header probes, zeroed-db quarantine, stats, holders (21 defs)

Existing mixins: search — shared FTS MATCH/LIKE builders, unified rebuild
status/step/finish engines, state_meta helpers; schema — one legacy/v23 FTS init
branch, shared _live_pk_columns, Row/tuple dual access dropped; portability —
shared _PREVIEW_RAW_SUBQUERY_SQL and _rich_row; common — single
stat_db_file_identity (was 3 copies), AUTO_VACUUM_MIN_FREELIST_RATIO.

Docstrings/comments hand-compacted (AST-identical) keeping every invariant,
ordering rule, failure mode and WHY. Schema SQL, migration order and PRAGMAs
untouched. test_repair_path_has_no_bare_connects repointed to hermes_state_repair.
2026-09-02 13:32:13 -07:00
teknium1 fd05029430 fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:

* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
  EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
  ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
  environment failures that polling cannot fix — `_acquire_db_flock` and
  both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
  now defer immediately with the real errno instead of burning the full
  120s / holder timeout and then logging a fake "held by another process".

* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
  open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
  lock) stayed `_fts_stale` — LIKE-only search — until the process
  reopened state.db. Short-lived CLIs reopen every run; the gateway opens
  once and stays up for days, so the deferral was effectively permanent
  (#100108). The retry runs from the EXISTING gateway housekeeping tick
  (`_start_gateway_housekeeping`, 60s) against the shared SessionDB
  instances via `hermes_state_registry.live_shared_session_dbs()`:
  non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
  bounded backoff 60s -> 1h, no new thread, still fails closed on live
  holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.

* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
  line can be matched to the interpreter that actually linked it
  (#100108 point 3).

Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 04:15:02 -07:00
the3asic 18ac3c4fb6 fix(state): defer corrupt FTS rebuilds past live operations 2026-08-31 12:08:30 -07:00
Teknium 22dcbdece6 fix(state): contain post-commit FTS maintenance errors + lock-audit the writer conn (salvage #90734)
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
  (including the bare SystemError CPython's sqlite3 layer raises under
  cross-thread errmsg scrambling) are contained and logged instead of
  escaping and making the caller replay an ambiguous, possibly-durable
  write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
  hammer + AST sweep freezing the no-unlocked-writer-conn invariant.

On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-31 09:56:04 -07:00
Teknium 9d0727d49b fix(state): single fail-closed cross-process authority for all full FTS rebuilds
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:

- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
  30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
  _rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()

Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.

Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
2026-08-23 19:00:36 -07:00
jackijianxa 0f33c207e6 fix(state): serialize cross-process FTS rebuild with file lock
When two Hermes processes (e.g. gateway + serve) detect FTS corruption
simultaneously, both run rebuild_fts() on the same database file in
parallel. rebuild_fts() only holds an in-instance threading lock, so the
concurrent rebuilds collide on write and structurally corrupt the
database ('file is not a database' / 'database disk image is malformed').

This happened twice in production (2026-08-15 and 2026-08-23), each time
requiring a full page-level salvage of state.db: sessions b-tree
clobbered, 5507 messages recovered row-by-row.

Fix: acquire an exclusive fcntl.flock on <db_path>.fts_rebuild.lock
before rebuilding, with a bounded 30s wait. The SQLite writer lock
remains the final backstop. POSIX-only; no-op elsewhere.
2026-08-23 19:00:36 -07:00
Adolanium ee9ec6164c perf(state): stop selecting full message content in session search
Every search route in _search_messages_impl (FTS, CJK bigram, trigram,
LIKE fallback, rebuild-gap supplement) selected m.content, then the
result tail popped it unread. On DBs with multi-MB tool rows, each
search read and materialized up to `limit` full rows only to discard
them. Snippets come from snippet()/substr() in SQL and the context
window is re-fetched by id, so no code path ever read the column.

Drop the column from all six SELECT lists. Returned dicts are
unchanged: content was never part of the public result (the pop ran
before return), and tests/test_hermes_state.py already documents that
contract.

(cherry picked from commit d0c3af167e7dd4eb18e1bea29ba107a94911ea24)
2026-08-15 00:34:29 -07:00
lkz-de ba80f3b86d fix(state): PASSIVE not TRUNCATE for all state.db checkpoints (#45383)
SessionDB.close() ran `PRAGMA wal_checkpoint(TRUNCATE)`. Every cron
run_agent opens and closes its own transient SessionDB, so on a busy
fleet this fired a full WAL reset many times an hour, racing the
gateway's long-lived writer on a large WAL database and tearing hot
B-tree pages -- structurally the same corruption this module's own
periodic checkpoint was already switched to PASSIVE to avoid (#45383).
Only close() and two manual-maintenance paths still used TRUNCATE.

Route every checkpoint on the shared state.db through PASSIVE:
  - close()                    (hermes_state.py)
  - pre-VACUUM in vacuum()     (hermes_state.py)
  - post-optimize-storage      (hermes_state_search.py)

PASSIVE never resets/truncates the WAL and never takes the exclusive
checkpoint lock, so it cannot lose a transient closer's race with the
live writer. The WAL is instead bounded by `journal_size_limit` and the
writer's natural post-checkpoint reset. TRUNCATE belongs only on a
sole-opener/quiescent connection (e.g. offline maintenance); this change
does not try to detect that -- PASSIVE is the safe default.

Diagnosed as the root cause of three state.db B-tree corruptions in
2026-08: damage localized to the hottest-written pages (gateway_routing
and the sessions indexes), with whole zero-filled pages still live and
off the freelist -- the checkpoint/reset-race signature, not disk or
application SQL.

Tests: tests/test_wal_checkpoint_strategy.py now asserts PASSIVE at
close(), before vacuum(), and after optimize_fts_storage() VACUUM;
tests/test_hermes_state.py asserts close() likewise. Focused run:
226 passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 19:43:22 -07:00
izumi0uu 1527a81b5e fix(state): keep canonical writes available when FTS is corrupt 2026-08-09 14:10:21 -07:00
Teknium 90311ee75f fix(search): strip % from non-CJK FTS5 queries
Closes the residual the contributor's own triage comment flagged: % was
excluded from the special-char class to protect the CJK LIKE fallback,
but a non-CJK query never reaches that fallback (is_cjk gates it), so
'50%' still hit MATCH raw and silently returned zero results. Strip %
whenever the sanitized query contains no CJK; the CJK path keeps its
pre-existing contract. Regression tests for both directions.
2026-08-08 19:17:09 -07:00
Drexuxux c595dcb955 fix(search): strip the FTS5 special characters the sanitizer was missing
_sanitize_fts5_query's strip step only removed +{}():"^ . Every other
character FTS5's grammar rejects outside a quoted phrase reached MATCH
raw and raised, and — as the step's own comment says about the colon it
was fixed for — the execute site swallows that into zero results. Session
search silently found nothing for ordinary queries:

  it's            fts5: syntax error near "'"
  gateway/run.py  fts5: syntax error near "/"
  user@host       fts5: syntax error near "@"
  a,b             fts5: syntax error near ","
  why?            fts5: syntax error near "?"
  e=mc2           fts5: syntax error near "="

Complete the class and assemble it with re.escape, because written as a
regex literal the backslash was eaten as an escape and never made it in
(C:\path\file still raised after the first pass).

Measured against a real FTS5 table over 651 realistic queries:
373 unparsable before, 77 after. The remainder is leading/trailing "." and
"-", which #43889 already covers.

% is deliberately left in: the CJK path falls back to a LIKE search that
needs it literal and escapes wildcards itself, so stripping it widened
those queries onto unrelated rows (test_cjk_like_escapes_wildcards).
2026-08-08 19:17:09 -07:00
kshitij fecba5afcc refactor(agent): fold simplify findings — DB picker parity, single scan, canonical strip delegation
Review-pass follow-ups (three parallel reviewers, findings verified):

- hermes_state_search.py list_recent_user_messages now drops legacy
  standalone compaction handoffs in the decode loop (SQL can't see them:
  durable role=user, no display_kind). Closes the /undo N pairing skew
  where the in-memory count (new predicate) and the DB soft-delete pick
  (old predicate) targeted different turns on legacy sessions. Fetches
  with headroom so the requested limit is still honored. 3 new tests,
  mutation-checked (no-op'ing the skip fails 2/3).
- _should_skip_model_call_for_reference_handoff: single drive-check scan
  (was two — once inside the restore helper, once after); the restore
  helper no longer re-scans and its return value now decides the verdict.
- _final_response_from_messages replaced by the _HANDOFF_SKIP_FINAL_RESPONSE
  constant it always returned (parameter was unused).
- _handoff_carries_live_user_content delegates to the canonical
  _strip_context_summary_handoff_message — also fixes the edge where a
  merged-shaped row with an EMPTY preserved prior tail was wrongly
  treated as carrying live content.
- Site-level guard test for rollback.restore with a legacy handoff row
  (predicate-in-context, complements the unit tests).
2026-08-07 19:44:35 +05:30
Chen Jin 23dce021a5 perf(fts): drain trash tables with a high-water marker instead of re-scanning
_fts_teardown_trash_step deleted rows via 'WHERE key IN (SELECT key
LIMIT N)' — each chunk's subquery re-scanned from the start of the
table, so chunk k skipped past (k-1)xN already-deleted rows: O(n²)
total row visits. On a v22 shadow table with ~230K rows that is on the
order of 10^8 row visits, turning optimize-storage teardown into a
multi-hour grind on slow disks, with a write lock held per chunk.

Single-column INTEGER-PK trash tables now drain via a fts_teardown_<tbl>_progress
high-water marker mirroring fts_rebuild_step: each chunk claims rows
past the marker (SELECT ... WHERE key > ? ORDER BY key LIMIT N), deletes
the claimed range, and publishes the new marker in the same transaction.
Per-chunk work is bounded → O(n) total.

TEXT-PK tables (the FTS config shadow table, pk like 'version') and
compound-key tables fall back to the legacy chunked delete — those are
small by construction.

Fixes #79324
2026-08-07 18:40:47 +05:30
kshitij 52a5fc0048 refactor(state): consolidate SQL LIKE escaping onto one shared helper
Follow-up to #79722, which introduced _escape_like in hermes_state.py for
the prune/archive filter fix. The same three-replace escape chain existed
as five more inline copies in hermes_state.py and two in
hermes_state_search.py (which must not import hermes_state — cycle).

Move the helper to hermes_state_common.escape_like (the module that exists
for exactly this) and route every copy through it:

- hermes_state.py: session-ID prefix resolution, find_session_by_title,
  get_next_title_in_lineage, the _like_pattern closure in list projection,
  and the kanban cwd retag
- hermes_state_search.py: the two LIKE-fallback token escapes

hermes_state re-imports it as _escape_like for back-compat. No behavior
change: every site produces byte-identical SQL patterns.
2026-08-06 04:28:44 +05:30
Kyzcreig 2f32092b38 hermes sessions optimize-storage aborts with
```
Error: optimization failed: no such table: messages_fts_trigram
No data was lost. Re-run to resume.
```

on any install where the trigram FTS index is legitimately absent. The failure is
deterministic — re-running can never make progress, because the crash happens at the same
point every time — so the database is permanently stuck on the legacy high-footprint FTS
layout with no supported way forward.

Observed on a 5.4 GB production `state.db`. After the fix the same database optimized
successfully and shrank to 3.3 GB.

The trigram index is absent whenever the runtime cannot maintain it. On a SQLite build
without the `trigram` tokenizer, `_ensure_fts_schema()` returns `False`, so `__init__`
leaves `self._trigram_available = False` and no `messages_fts_trigram` table on disk. This
is a **supported degraded runtime**, not damage — CJK/substring search falls back to
`LIKE` and everything else works normally. `_is_fts5_unavailable_error()` and
`_warn_trigram_unavailable()` exist specifically to make this path graceful.

Two code paths write the boundary sweep for the deferred FTS rebuild, and only one of them
respects that flag:

| Function | Trigram `INSERT` guarded? |
|---|---|
| `fts_rebuild_step()` | ✅ `if include_trigram:` where `include_trigram = self._trigram_available` |
| `_fts_rebuild_finish()` | ❌ unconditional |

`_fts_rebuild_finish()` runs the boundary sweep at the *end* of the backfill. Its
unguarded `INSERT INTO messages_fts_trigram …` raises `OperationalError`, which propagates
out of `optimize_fts_storage()` and aborts the entire optimization — *after* the backfill
has already completed. Hence the characteristic output showing 100% progress immediately
before the error:

```
Rebuilding index: 100% (909,671/909,671)
Error: optimization failed: no such table: messages_fts_trigram
```

There is a second, quieter consequence. The teardown phase that reclaims the demoted
`fts_v22_trash_*` shadow tables runs *after* the backfill phase in
`optimize_fts_storage()`. Because the crash happens before teardown is ever reached, those
tables are never emptied or dropped — so the space the migration was supposed to reclaim
stays allocated indefinitely, and the leftover trash tables look (misleadingly) like
evidence of a half-finished migration.

Build a populated v23 database, set the deferred-rebuild markers, then reopen it on a
runtime where `_ensure_fts_schema('messages_fts_trigram', …)` returns `False` (exactly
what a SQLite build without the trigram tokenizer produces) and call
`optimize_fts_storage()`:

```
[precondition] trigram absent, _trigram_available=False, rebuild pending  ✓

RED  ✗ optimize_fts_storage raised OperationalError: no such table: messages_fts_trigram
```

With this patch applied, unchanged harness:

```
     optimize_fts_storage returned {'ok': True, 'vacuumed': None}
GREEN ✓ optimize ok; markers cleared; base FTS 'zebra' -> 200 hits
```

Full harness and transcripts in `TEST-EVIDENCE.md`.

Gate the sweep on `self._trigram_available`, exactly as `fts_rebuild_step()` already does:

```python
include_trigram = self._trigram_available

def _do(conn):
    ...
    if include_trigram:
        conn.execute("INSERT INTO messages_fts_trigram(...) ...")
```

The base `messages_fts` sweep and the marker cleanup are untouched, so the rebuild still
finalizes correctly and the index remains complete for every row it is responsible for.
The fix does not disable or weaken search to dodge the error — the regression tests assert
that base FTS still returns results afterwards.

`TestFtsRebuildFinishWithoutTrigram` in `tests/test_hermes_state.py`:

- `test_rebuild_finish_skips_trigram_when_unavailable` — drives `_fts_rebuild_finish()`
  directly on a trigram-less runtime; asserts it completes, clears both rebuild markers,
  and leaves base FTS searchable.
- `test_optimize_fts_storage_succeeds_without_trigram` — end-to-end through the public
  `optimize_fts_storage()` entry point; asserts `ok=True`, markers cleared, search intact.

Both use the existing `_NoTrigramConnection` helper already in the file. Both fail on
`main` with `no such table: messages_fts_trigram` and pass with this patch.

`tests/test_hermes_state.py` passes in full (463 tests → 465 with these two). `ruff` clean.

This PR is the crash only.

A companion PR narrows `_db_opens_cleanly()` so that
`hermes sessions repair --check-only` stops reporting a write-broken FTS schema as
healthy — the gap that makes this class of problem hard to diagnose in the first place.
The two are independent and can land in either order.
2026-08-03 19:09:15 +05:30
Jakub Wolniewicz ffb54305c4 perf(session-search): project fields before enrichment 2026-08-03 17:50:58 +05:30
kshitijk4poor 1e2e69db98 perf(state): truncate FTS index with 'delete-all' instead of plain DELETE
_reset_fts_index_to_empty used a no-WHERE DELETE, whose docstring
claimed FTS5 treats it as an efficient drop-all. That's true only for
ordinary rowid tables — on external-content FTS5 each deleted row's
tokens are regenerated from the content table, making it O(rows)
(measured ~12us/row: 0.22s @100K, 5.2s @400K, ~25s projected @2M) while
holding the write lock. It also corrupts the index when indexed rows
have diverged from messages — exactly the broken-bookkeeping shape this
repair path handles. The FTS5 'delete-all' special command is the
documented O(1) truncate for external-content tables (measured 1.6ms
@100K) and truncates unconditionally regardless of divergence.
2026-08-02 21:36:11 +05:30
kshitijk4poor 2febb5823c perf(state): probe empty FTS index with EXISTS instead of COUNT(*)
_fts_external_index_empty_with_messages runs on every writable open via
the _init_schema fts_storage_version stamp condition. COUNT(*) is a full
b-tree scan on both messages and messages_fts_docsize (~100ms per open
on a 2M-row DB, measured); the function only ever compares against
zero, so EXISTS(SELECT 1 ...) gives the identical boolean in O(1).
2026-08-02 21:36:11 +05:30
Adolanium b2d5995fc6 fix(state): do not stamp empty FTS after interrupted optimize-storage demote
Demote wrote the empty v23 schema via executescript inside BEGIN IMMEDIATE,
which commits early and can leave trash + empty indexes without rebuild
markers. Re-run then tore down trash and stamped fts_storage_version with
docsize=0, permanently losing historical session search.

Stage markers with the demote, create schema only after they are durable,
heal empty-index bookkeeping on resume, and refuse settle until the base
index is populated. Settle refusal returns ok=False instead of raising,
and resume fails fast if the base v23 table cannot be re-created.

Orphan-marker repair only resets a missing fts_rebuild_progress to 0 once
the index is known empty: the chunk worker replays its whole selected id
range without an anti-join, so a partially indexed DB that lost only its
progress key is first reset to a known-empty surface, then rebuilt.

Ported onto the SessionDB mixin split (hermes_state_search.py /
hermes_state_schema.py).
2026-08-02 21:36:11 +05:30
teknium1 21c7ae8563 refactor: split SessionDB into Search/Schema/Portability mixins (mechanical move, ~2.9K LOC out of hermes_state.py) 2026-07-29 10:14:19 -07:00