5 Commits

Author SHA1 Message Date
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
teknium1 0c0875b746 chore: delete orphaned bench data, datagen examples and stale one-off docs
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.

- mcp-research-data/: 224K of July tool-search bench result rows. The
  harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
  gitignored dir; the rows were committed by hand once and the headline
  numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
  WebResearchEnv that no longer exists; the yaml paths point at a
  configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
  for a resolved bug, two RFCs whose work shipped, an unimplemented
  profile-builder proposal, the kanban dialog mock HTML and the kanban v1
  spec PDF. profile-routing.md duplicated the profile_routes section of
  website/docs/user-guide/multi-profile-gateways.md.

Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
2026-09-13 06:06:46 -07:00
teknium1 2643ea17fb bench: discovery-bound suite — paraphrase/absence/survey tasks isolate the listing's structural advantage
Bridge vs listing only (Opus 4.8, 830 real UE schemas, 3 reps/cell).
Excluding one both-modes mock artifact: listing 24/24 vs bridge 20/24,
searches/task 0.2 vs 4.0. Bridge failures: core-tool substitution at
frontier tier (ran the host test suite via terminal instead of
discovering RunTests, 2/3 reps), up to 8 searches to prove a negative,
and search-vocabulary misses on paraphrase. Listing asserts absence in
zero searches and answers a 5-way capability survey in 1 API call.
2026-07-26 08:26:09 -07:00
teknium1 21cc643ac2 bench: adversarial 830-tool gauntlet — confusion clusters, type-aware error mocks, strict scoring
Scenarios target real confusion clusters in Epic's UE 5.8 catalog
(StaticMesh vs SkeletalMesh set_material, three tag systems, CurveTable
vs DataTable rows, Niagara Component vs System variables, four capture
variants, zero-keyword phrasing). Mocks return realistic editor errors
on wrong-type calls; scoring is strict (clean solve = correct tool with
zero distractor calls; first-call accuracy tracked separately).

Key result: first-call selection is unreliable in EVERY mode — eager
with all 199K of schemas in context managed 2/10 — but clean solves stay
75-95% because agents probe (get_components, get_material_slots) before
committing. The probe loop works through the 3-tool bridge at 1/4 the
cost of eager ($1.60-1.69 vs $6.49/task, Opus 4.8). On Haiku the
listing beats bare bridge 18/20 vs 15/20 (core-tool substitution again).
Zero distractor invocations across all 50 Opus runs.
2026-07-26 08:26:09 -07:00
teknium1 3a1bbfac61 bench: Unreal-scale live benchmark — Epic's real 830 UE 5.8 schemas replayed (Opus 4.8)
Replays the actual tool schemas captured from Epic's UE 5.8
ModelContextProtocol + AllToolsets plugins (830 tools / 52 toolsets) as
live registry tools with mocked editor responses, then benchmarks
eager vs bare-bridge vs bridge+listing at two scales (62-tool editor
subset, full 830) on Claude Opus 4.8 (1M ctx; eager at 830 does not fit
any 200K model — first call requests ~266K tokens).

Headline (full 830, mean per task, rescored): eager 8/8 at 810,578
input tokens ($4.05); bare bridge 16/16 at 160,844 ($0.80); listing
16/16 at 257,264 ($1.29). Frontier model erases the accuracy gap in
every mode; cost is the differentiator. At 62 tools eager wins on cost
— consistent with the auto-threshold design.

Also parameterizes livetest harness model + listing_max_tokens via
env/args (TS_UE_MODEL, TS_UE_SCALE, TS_UE_MODES, TS_UE_LISTING_MAX).
2026-07-26 08:26:09 -07:00