3 Commits

Author SHA1 Message Date
ouyangbo 012c9c6bed chore: anchor fresh-start history to upstream 2026-09-16 15:06:54 +08:00
teknium1 0c0875b746 chore: delete orphaned bench data, datagen examples and stale one-off docs
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.

- mcp-research-data/: 224K of July tool-search bench result rows. The
  harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
  gitignored dir; the rows were committed by hand once and the headline
  numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
  WebResearchEnv that no longer exists; the yaml paths point at a
  configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
  for a resolved bug, two RFCs whose work shipped, an unimplemented
  profile-builder proposal, the kanban dialog mock HTML and the kanban v1
  spec PDF. profile-routing.md duplicated the profile_routes section of
  website/docs/user-guide/multi-profile-gateways.md.

Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
2026-09-13 06:06:46 -07:00
teknium1 21cc643ac2 bench: adversarial 830-tool gauntlet — confusion clusters, type-aware error mocks, strict scoring
Scenarios target real confusion clusters in Epic's UE 5.8 catalog
(StaticMesh vs SkeletalMesh set_material, three tag systems, CurveTable
vs DataTable rows, Niagara Component vs System variables, four capture
variants, zero-keyword phrasing). Mocks return realistic editor errors
on wrong-type calls; scoring is strict (clean solve = correct tool with
zero distractor calls; first-call accuracy tracked separately).

Key result: first-call selection is unreliable in EVERY mode — eager
with all 199K of schemas in context managed 2/10 — but clean solves stay
75-95% because agents probe (get_components, get_material_slots) before
committing. The probe loop works through the 3-tool bridge at 1/4 the
cost of eager ($1.60-1.69 vs $6.49/task, Opus 4.8). On Haiku the
listing beats bare bridge 18/20 vs 15/20 (core-tool substitution again).
Zero distractor invocations across all 50 Opus runs.
2026-07-26 08:26:09 -07:00