refactor: move live-model benchmark harnesses from scripts/ into evals/

scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".

- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
  -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
  the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
This commit is contained in:
teknium1
2026-09-13 05:37:53 -07:00
committed by Teknium
parent 0c0875b746
commit b980495847
14 changed files with 15 additions and 15 deletions
+2 -2
View File
@@ -190,8 +190,8 @@ docs/superpowers/*
/.install_method
# Tool Search live-test harness output — non-deterministic model transcripts,
# regenerated by scripts/tool_search_livetest.py. Never an artifact of the repo.
scripts/out/
# regenerated by evals/tool_search/tool_search_livetest*.py. Never an artifact of the repo.
evals/tool_search/out*/
# Per-release changelog drafts. These exist only transiently during a release
# cut (passed to `gh release create --notes-file`); the GitHub Release itself
+1 -1
View File
@@ -21,7 +21,7 @@ web tasks.
rating aggregation, JS/delayed render, login chain, cross-category
compare).
- **Resume-safe.** Completed cells in `results/*.jsonl` are skipped on rerun
(same pattern as `scripts/toolperf_abeval`).
(same pattern as `evals/toolperf_abeval`).
- **Backend matrix.** `orchestrate.py` drives a local headless-Chrome CDP;
`orchestrate_cloud.py --backend nous-cloud|browserbase` provisions a real
cloud browser per cell through the same provider plumbing the product uses.
@@ -4,7 +4,7 @@ Runs both paths against the same live Chrome and prints a comparison table.
Not a pytest — a script you run manually for the PR description.
Usage:
.venv/bin/python scripts/benchmark_browser_eval.py [--iterations N]
.venv/bin/python evals/browser_use/benchmark_browser_eval.py [--iterations N]
"""
from __future__ import annotations
+1 -1
View File
@@ -1,7 +1,7 @@
"""Local-CDP battery orchestrator: tasks x arms x models x reps.
Resume-safe: completed cells in results.jsonl are skipped, so a killed
battery continues where it left off (same pattern as scripts/toolperf_abeval).
battery continues where it left off (same pattern as evals/toolperf_abeval).
Usage:
# start a headless Chrome first:
@@ -2,14 +2,14 @@
Runs five scenarios against a real model (Claude Haiku 4.5 via OpenRouter) to
verify that the bridge tools work end-to-end. Records transcripts in
`scripts/out/`.
`evals/tool_search/out/`.
## Running
```bash
cd <repo root>
python3 scripts/tool_search_livetest.py # runs all 5 scenarios x 2 modes
python3 scripts/analyze_livetest.py # side-by-side report
python3 evals/tool_search/tool_search_livetest.py # runs all 5 scenarios x 2 modes
python3 evals/tool_search/analyze_livetest.py # side-by-side report
```
Requires `OPENROUTER_API_KEY` set or present in `~/.hermes/.env`.
@@ -34,7 +34,7 @@ A/B baseline. The harness records:
## Output structure
```
scripts/out/
evals/tool_search/out/
<scenario>__enabled.json # tool_search ON
<scenario>__disabled.json # tool_search OFF
_summary.json # one-line summary across all runs
@@ -37,7 +37,7 @@ ORIGINAL_HOME = os.environ.get("HERMES_HOME")
ORIGINAL_AUTH = Path.home() / ".hermes" / "auth.json"
_THIS_DIR = Path(__file__).resolve().parent
_WORKTREE_ROOT = _THIS_DIR.parent
_WORKTREE_ROOT = _THIS_DIR.parents[1]
sys.path.insert(0, str(_WORKTREE_ROOT))
# ---------------------------------------------------------------------------
@@ -16,7 +16,7 @@ from pathlib import Path
from typing import Any, Dict, List
_THIS_DIR = Path(__file__).resolve().parent
_WORKTREE_ROOT = _THIS_DIR.parent
_WORKTREE_ROOT = _THIS_DIR.parents[1]
sys.path.insert(0, str(_WORKTREE_ROOT))
sys.path.insert(0, str(_THIS_DIR))
@@ -23,7 +23,7 @@ from pathlib import Path
from typing import Any, Dict, List
_THIS_DIR = Path(__file__).resolve().parent
_WORKTREE_ROOT = _THIS_DIR.parent
_WORKTREE_ROOT = _THIS_DIR.parents[1]
sys.path.insert(0, str(_WORKTREE_ROOT))
sys.path.insert(0, str(_THIS_DIR))
@@ -24,7 +24,7 @@ from pathlib import Path
from typing import Any, Dict, List
_THIS_DIR = Path(__file__).resolve().parent
_WORKTREE_ROOT = _THIS_DIR.parent
_WORKTREE_ROOT = _THIS_DIR.parents[1]
sys.path.insert(0, str(_WORKTREE_ROOT))
sys.path.insert(0, str(_THIS_DIR))
@@ -27,7 +27,7 @@ from pathlib import Path
from typing import Any, Dict, List
_THIS_DIR = Path(__file__).resolve().parent
_WORKTREE_ROOT = _THIS_DIR.parent
_WORKTREE_ROOT = _THIS_DIR.parents[1]
sys.path.insert(0, str(_WORKTREE_ROOT))
sys.path.insert(0, str(_THIS_DIR))
@@ -55,7 +55,7 @@ in real production traffic.
## Run
```bash
cd scripts/toolperf_abeval
cd evals/toolperf_abeval
export ABEVAL_ROOT=/tmp/abeval-workspace # results + sandboxes land here
export ABEVAL_HOME=/tmp/abeval-home
./run_all.sh /tmp/abeval-baseline /path/to/fixes-tree 3 \