bench: discovery-bound suite — paraphrase/absence/survey tasks isolate the listing's structural advantage

Bridge vs listing only (Opus 4.8, 830 real UE schemas, 3 reps/cell).
Excluding one both-modes mock artifact: listing 24/24 vs bridge 20/24,
searches/task 0.2 vs 4.0. Bridge failures: core-tool substitution at
frontier tier (ran the host test suite via terminal instead of
discovering RunTests, 2/3 reps), up to 8 searches to prove a negative,
and search-vocabulary misses on paraphrase. Listing asserts absence in
zero searches and answers a 5-way capability survey in 1 API call.
This commit is contained in:
teknium1
2026-07-19 01:41:46 -07:00
committed by Teknium
parent 21cc643ac2
commit 2643ea17fb
2 changed files with 1543 additions and 0 deletions
File diff suppressed because it is too large Load Diff