export_session_lineage spread segments[-1] over the merged dict, so the top-level `timings`
described only the last compression segment while `messages` spanned the whole lineage — a
reader would see a 500 ms wall clock over a lineage that ran for hours. Compute the block over
the merged message list (segments keep their own).
_validate_import_session measured the raw session JSON, so the derived `timings.intervals`
(one entry per message pair) counted toward the 5 MiB per-session limit and could reject a
long lineage whose actual content fits. Strip `timings` before measuring; it is rebuilt from
the messages on the next export anyway.
`_export_timings` parsed message timestamps with a bare `float()`, so a
pre-existing out-of-range cell (e.g. 8.4e252) reached
`datetime.fromtimestamp` and raised OverflowError, killing the whole
`export_session`/`export_all` call — a regression caught by
tests/hermes_state/test_corrupt_row_robustness.py. Route the cell through
`coerce_epoch`, the reader every other timestamp surface already uses, so a
bad row degrades to a `missing` count (with the usual warning naming the
session) and the export still lands.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.
Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".
gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.
Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).
Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.
Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.
Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
Browse the backend host's foreign CLI session logs, preview a bounded read-only
transcript, and continue a copy in Hermes under the selected profile. Reuses the
hermes_cli.foreign_sessions parsers and the portability validator/writer;
imports are transactional and deduplicated on the recorded origin.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
test_no_locked_readers_gate.py (#97676) parses hermes_state.py's SessionDB
class body with ast and flags any method that holds the writer lock
around a pure-read query — Pattern C, where every concurrent turn's
persistence convoys behind an unrelated read. #97676 converted 39 such
methods and closed detection blind spots for alias/variable-SQL readers.
But SessionDB is declared as
`class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)`,
and the gate only ever opened hermes_state.py — it never parsed the three
mixin files those base classes are defined in, so a locked reader
declared there was structurally invisible to the scanner regardless of
how good the alias/variable-SQL detection got.
Applying the gate's exact scanning logic to the three mixin files
directly turns up 9 genuine pure-read methods still holding the writer
lock, none in #97676's converted list:
- hermes_state_search.py: _fts_teardown_trash_step, fts_optimize_available,
optimize_fts_storage, list_recent_user_messages
- hermes_state_portability.py: distinct_session_cwds, list_cron_job_runs,
_get_session_rich_rows_batch, list_skill_scaffolded_sessions,
get_first_assistant_text
_get_session_rich_rows_batch is a hot path: it backs list_sessions_rich's
compression-tip resolution and the web server's session-search hydration
across every gateway install — its own docstring already claims "same
read-your-writes guarantee as list_sessions_rich", but list_sessions_rich
was already using _read_ctx() (its guarantee comes from flush_token_counts()
before the read, not from holding the writer lock) while this method's
implementation never caught up to match.
Converted all 9 to `with self._read_ctx() as conn:`, the exact pattern
#97676 used, verified each is a genuine pure read with no hidden writes
by tracing every helper call it makes.
Extended the gate itself (_ALL_STATE_SOURCES) to scan all three mixin
files under their own class names, plus hermes_state.py, so this blind
spot can't silently reopen. Added a regression test
(test_scan_all_state_sources_visits_every_mixin_file) that plants a
synthetic violation in a mixin-shaped file and asserts the scanner still
catches it — a change that reverts the file list back to one file passes
the existing sabotage test but fails this one.
Mutation-verified: with the gate's new scope but the old (unconverted)
mixin sources, test_no_locked_pure_readers fails and names all 9 real
violations with correct file/line. Restored the fix; it passes clean.
Closes the TOCTOU window flagged in review on #93369 (merged via
#93430): the divergence guard compared EXPORT-TIME message counts, but
another backend can append donor messages between the export snapshot
and the retire loop — that growth would be stamped behind the
non-recoverable adopted_by_profile archive, the exact H2 class the
guard exists to prevent, just via a narrower race.
The retire loop now re-reads live donor vs local counts immediately
before end_session and leaves the donor unretired (donor_retired=False,
warn-logged) on any donor-ahead signal; the next resume's export-time
guard then handles the divergence normally. Equal-count CONTENT
divergence (donor rewind+rewrite) remains invisible to count comparison
— documented as accepted: bytes stay in the donor store either way.
New red-first-verified regression simulates the exact race by appending
to the donor from inside an export_session_lineage wrapper.
adoption+ownership suites: 25 passed; ruff clean.
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
an UNRELATED default-store conversation (bot titles collide by design;
get_session_by_title has no archived filter/ordering). Donor probe is
now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
accumulated NEWER messages than the profile copy (skip-based
idempotency never merges). New divergence guard compares message
counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
failed under suppress. Now per-segment tracked + warn-logged;
True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.
5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.
- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
composes the existing export_session_lineage()/import_sessions()
primitives; donor rows are archived (never deleted) with
end_reason=adopted_by_profile, which is deliberately NOT in
RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
to adoption from the default store right before the 4007; ids unknown
to BOTH stores still 4007 exactly as before, and launch-profile
resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
unit adoption + non-resurrectable archive; 3 handler-level through
server.handle_request incl. the live repro shape); db-ownership
leak test taught that the shared launch handle probe is by design.
Follow-up to #93296/#93311; part of #93091.
Rework of salvaged PR #6372 (@ag9920) onto current main:
- /save promoted from CLI-only JSON snapshot to a cross-platform session
export: `/save [json|md|html] [filename] [redact]` on CLI and every
gateway platform (sent as a document via adapter.send_document).
- Rendering routes through the existing shared renderers
(hermes_cli/session_export.py + session_export_html.py) instead of the
PR's new hermes_state formatter — new helpers normalize_save_format /
render_session_for_save / default_save_filename are shared by both
surfaces.
- `redact` arg runs the export through the force-mode secret redaction
pass (session_export_md.redact_session_data) before writing.
- Gateway handler awaits AsyncSessionDB correctly, sanitizes user-supplied
filenames with basename, and lands in gateway/slash_commands.py (the
handlers moved out of gateway/run.py since the PR was authored).
- /export stays profile export (name collision resolved: session export
lives on /save).
- Slack 50-slash cap curation: /platform moves to the /hermes-only set to
free a native slot for /save (parity test updated rationale comment).
- Folds in PR #62268 (@briandevans): None title/model coalescing in the
single-session HTML export.
Closes#4249. Closes#51200.
Simplify-pass fold: SQLITE_MAX_VARIABLE_NUMBER is 999 on pre-3.32\nbuilds (which the repo still supports — the trigram-availability\nmachinery exists for exactly that class), and limit=10000\nlist_sessions_rich callers exist in web_server. Chunk inside the\nbatch helper — the single choke point — so no call site can overflow.
list_sessions_rich()'s compression-root projection called
_get_session_rich_row() once per root — a separate single-row query per
compression root on every session-list render. Resolve every tip id
first, then fetch all tip rows in one WHERE id IN (...) query via the
new _get_session_rich_rows_batch().
_get_session_rich_row() is now a thin wrapper over the batch method, so
the enriched SELECT (preview + last_active) lives in exactly one place —
future column changes (e.g. #42196's include_system_prompt) only touch
one query.
get_compression_tip()'s chain walk is untouched; it's a genuine
per-session graph walk with branch/delegate-exclusion and race handling,
and batching it safely is out of scope here.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Config docs now describe session_stall_timeout precisely: a RECOVERY
notifier for an in-process AIAgent with an adapter-queued follow-up —
not a general gateway/session stall detector — with a per-AIAgent scan
cadence (not globally coordinated per durable session).
- import_sessions documents the deliberate export-includes /
import-resets asymmetry for the activity fields (no resurrected
'working' labels on machines where no agent runs), with a regression
pinning both halves.
- Strip trailing whitespace in contributors/emails/fangliquan@qq.com
(git diff --check housekeeping).
PR #76354 review, scope/contract items + housekeeping.
Three mechanisms to detect and notify when gateway sessions stall silently:
1. Mid-turn activity heartbeats stamped to SessionDB so hermes sessions list
and hermes status show progress during long turns without new message rows.
2. Stall watchdog: when a busy session has pending inbound and the shared
activity clock is idle past agent.session_stall_timeout (default 300),
log a WARNING and notify the user once to try /new. Notify-only; does
not kill the turn.
3. Compaction timeout: fenceless compress_context callers get a progress-aware
host budget (compression.context_timeout_seconds default 120 idle,
compression.context_total_ceiling_seconds default 600 ceiling). On timeout,
cancel via commit fence, skip compaction without dropping messages, and
continue the turn.
Closes#72016 (slices 1-3; slice 4 cumulative SSE stream-retry deadline
remains a follow-up).
Cherry-picked from PR #72424 by @fangliquanflq.