`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
The classic CLI froze for ~0.7-2s between the banner and the first prompt. py-spy +
strace on real PTY startups showed the main/REPL threads inside
refuse_deleted_wal_generation -> _iter_proc_fd_targets: a second full SessionDB open.
_init_session_store built a bare SessionDB(); a moment later the goal/loop/heartbeat
managers acquired the same state.db through hermes_state_registry from the REPL thread,
which is a different handle, so the whole open ran again — including the /proc-wide
deleted-WAL sidecar scan (~4.4k readlinks). Each readlink drops and re-takes the GIL
while the startup threads (plugin discovery, MCP, skill sync, banner git) are busy, so an
11ms scan stretched to 1.3s per pass, and the second pass landed exactly where the
prompt should have appeared.
Route the CLI's handle (init + the two re-open sites) through the registry so every
in-process consumer shares one writer. One scan per startup; live A/B on the same box,
interleaved x6: banner->prompt gap 0.37s mean -> 0.15s mean (plain), 1.77s -> 0.39s
under strace. The registry release path replaces close(), so /quit, /snapshot restore
and /handoff keep their semantics.