`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Three-reviewer pass (reuse/quality/efficiency) on the guard file:
- Drop dead `agent._interrupt_requested = False` setup: the
`_record_streamed_assistant_text` chain only consults
`_stream_writer_superseded()` (stream-writer TLS token), never
`_interrupt_requested` — verified by reading both call sites.
- Deduplicate the two trace-callback blocks into one
`_count_writer_statements` helper using the house idiom
(`statements.append` — 9 existing uses in tests/test_hermes_state.py)
instead of a mutable-dict counter closure; failure messages now dump
the captured SQL for direct diagnosis.
Dropped after verification: reviewer suggestion to flip xfail to
strict=True — its premise ("the fix PRs already remove the markers")
is wrong: #92166/#95380 predate this file and cannot remove markers
they don't contain, so strict=True would redden main's CI the moment
either merges. strict=False + follow-up marker removal is the
deliberate no-red-window ratchet.
Efficiency reviewer: no material findings (GC delta 0.06 on the ratio,
36.9MB peak, sqlite trace API stable since 3.14, dir convention OK).
Re-verified post-fold: 3x main runs (2 passed, 2 xfailed), flip checks
on both fix branches still pass with --runxfail.
Pattern B (O(N²) rebuild-per-delta in hot paths) has no lintable
signature, unlike Pattern A's ASYNC ruff gate: `s += frag` is quadratic
in a hot loop and harmless elsewhere. The only durable prevention is
behavioral — pin the scaling SHAPE of each known hot path and fail CI
when it regresses.
New tests/perf_guards/test_pattern_b_scaling.py, three guards:
1. Streamed-text accumulation (fix in flight: #92166)
Self-normalizing ratio: time at 16k deltas over 4k deltas.
Linear ≈ 4x, quadratic ≈ 16x; main measures 9.6x → strict bound 7.0.
xfail on today's main, PASSES on the #92166 branch (verified).
2. list_sessions_rich statement count (fix in flight: #95380)
Deterministic — counts writer-connection SQL statements via sqlite
trace callback, zero timing. Bounded-constant guard xfails on main
(measured N+1: 26 stmts / 12 sessions), PASSES on #95380 (verified).
A second, weaker budget guard (≤4·N+8) passes TODAY and catches a
regression from N+1 to N·M immediately.
3. Tool-call fragment assembly (#92242 shape)
Ratio guard on the buffered-parts accumulator model; sized so the
small case takes ≥3ms (sub-ms bases jitter on CI runners).
Flake hardening: min-of-K timing, ratio thresholds with ≥2x separation
from both measured-good and measured-bad, operation counts preferred
over timing. 10/10 identical outcomes across repeated local runs.
The xfail markers are the ratchet contract: each names its fix PR and
must be removed when that PR merges, flipping the guard to enforcing.