b10952e9c6
unicode61 indexes a CJK run as ONE token, so 2-char Korean terms (일본, 구글, 우리, ...) can never match it and the trigram tokenizer needs >=3 chars per term — any query containing a 1-2 char CJK token falls through to a LIKE full-table scan (measured 3-6.4s CPU per query on a 6.8GB production state.db; the #1 base cost behind a 12.4s session_search average on CJK workloads). This ships a ~250-line loadable FTS5 tokenizer (no deps) that wraps unicode61: maximal CJK runs inside its tokens are re-emitted as overlapping character bigrams (Lucene CJKAnalyzer semantics), everything else passes through unchanged. FTS5 phrase semantics turn consecutive sub-tokens into exact substring matching down to 2-char terms at index speed. Build: native/fts5_cjk/build.sh -> ~/.hermes/lib/libfts5_cjk.so (override: HERMES_FTS5_CJK_SO). Salvaged from PR #65544; the schema integration lands separately on the v23 external-content layout.
fts5_cjk — cjk_unicode61 FTS5 tokenizer
unicode61 + CJK character bigrams (Lucene CJKAnalyzer semantics). Fixes 2-char Korean/CJK terms falling through to LIKE full-table scans.
Build & install to ~/.hermes/lib/:
./build.sh
Then run scripts/fts_v2_migrate.py to create + backfill messages_fts_v2,
and set HERMES_FTS_V2_READ=1 in ~/.hermes/.env to cut reads over.
Override the .so location with HERMES_FTS5_CJK_SO.