Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.
- mcp-research-data/: 224K of July tool-search bench result rows. The
harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
gitignored dir; the rows were committed by hand once and the headline
numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
WebResearchEnv that no longer exists; the yaml paths point at a
configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
for a resolved bug, two RFCs whose work shipped, an unimplemented
profile-builder proposal, the kanban dialog mock HTML and the kanban v1
spec PDF. profile-routing.md duplicated the profile_routes section of
website/docs/user-guide/multi-profile-gateways.md.
Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
Replays the actual tool schemas captured from Epic's UE 5.8
ModelContextProtocol + AllToolsets plugins (830 tools / 52 toolsets) as
live registry tools with mocked editor responses, then benchmarks
eager vs bare-bridge vs bridge+listing at two scales (62-tool editor
subset, full 830) on Claude Opus 4.8 (1M ctx; eager at 830 does not fit
any 200K model — first call requests ~266K tokens).
Headline (full 830, mean per task, rescored): eager 8/8 at 810,578
input tokens ($4.05); bare bridge 16/16 at 160,844 ($0.80); listing
16/16 at 257,264 ($1.29). Frontier model erases the accuracy gap in
every mode; cost is the differentiator. At 62 tools eager wins on cost
— consistent with the auto-threshold design.
Also parameterizes livetest harness model + listing_max_tokens via
env/args (TS_UE_MODEL, TS_UE_SCALE, TS_UE_MODES, TS_UE_LISTING_MAX).