refactor: move live-model benchmark harnesses from scripts/ into evals/
scripts/ is for repo tooling (tests runner, release, installers, CI checks). The tool-search live tests, the toolperf A/B eval and the browser eval benchmark are offline benchmarks that spend model budget, which is exactly what evals/ holds; evals/browser_use already cited scripts/toolperf_abeval as "the same pattern". - scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for the extra directory level; gitignore now covers evals/tool_search/out*/ - scripts/toolperf_abeval/ -> evals/toolperf_abeval/ - scripts/benchmark_browser_eval.py -> evals/browser_use/
This commit is contained in:
+2
-2
@@ -190,8 +190,8 @@ docs/superpowers/*
|
||||
/.install_method
|
||||
|
||||
# Tool Search live-test harness output — non-deterministic model transcripts,
|
||||
# regenerated by scripts/tool_search_livetest.py. Never an artifact of the repo.
|
||||
scripts/out/
|
||||
# regenerated by evals/tool_search/tool_search_livetest*.py. Never an artifact of the repo.
|
||||
evals/tool_search/out*/
|
||||
|
||||
# Per-release changelog drafts. These exist only transiently during a release
|
||||
# cut (passed to `gh release create --notes-file`); the GitHub Release itself
|
||||
|
||||
@@ -21,7 +21,7 @@ web tasks.
|
||||
rating aggregation, JS/delayed render, login chain, cross-category
|
||||
compare).
|
||||
- **Resume-safe.** Completed cells in `results/*.jsonl` are skipped on rerun
|
||||
(same pattern as `scripts/toolperf_abeval`).
|
||||
(same pattern as `evals/toolperf_abeval`).
|
||||
- **Backend matrix.** `orchestrate.py` drives a local headless-Chrome CDP;
|
||||
`orchestrate_cloud.py --backend nous-cloud|browserbase` provisions a real
|
||||
cloud browser per cell through the same provider plumbing the product uses.
|
||||
|
||||
@@ -4,7 +4,7 @@ Runs both paths against the same live Chrome and prints a comparison table.
|
||||
Not a pytest — a script you run manually for the PR description.
|
||||
|
||||
Usage:
|
||||
.venv/bin/python scripts/benchmark_browser_eval.py [--iterations N]
|
||||
.venv/bin/python evals/browser_use/benchmark_browser_eval.py [--iterations N]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
"""Local-CDP battery orchestrator: tasks x arms x models x reps.
|
||||
|
||||
Resume-safe: completed cells in results.jsonl are skipped, so a killed
|
||||
battery continues where it left off (same pattern as scripts/toolperf_abeval).
|
||||
battery continues where it left off (same pattern as evals/toolperf_abeval).
|
||||
|
||||
Usage:
|
||||
# start a headless Chrome first:
|
||||
|
||||
@@ -2,14 +2,14 @@
|
||||
|
||||
Runs five scenarios against a real model (Claude Haiku 4.5 via OpenRouter) to
|
||||
verify that the bridge tools work end-to-end. Records transcripts in
|
||||
`scripts/out/`.
|
||||
`evals/tool_search/out/`.
|
||||
|
||||
## Running
|
||||
|
||||
```bash
|
||||
cd <repo root>
|
||||
python3 scripts/tool_search_livetest.py # runs all 5 scenarios x 2 modes
|
||||
python3 scripts/analyze_livetest.py # side-by-side report
|
||||
python3 evals/tool_search/tool_search_livetest.py # runs all 5 scenarios x 2 modes
|
||||
python3 evals/tool_search/analyze_livetest.py # side-by-side report
|
||||
```
|
||||
|
||||
Requires `OPENROUTER_API_KEY` set or present in `~/.hermes/.env`.
|
||||
@@ -34,7 +34,7 @@ A/B baseline. The harness records:
|
||||
## Output structure
|
||||
|
||||
```
|
||||
scripts/out/
|
||||
evals/tool_search/out/
|
||||
<scenario>__enabled.json # tool_search ON
|
||||
<scenario>__disabled.json # tool_search OFF
|
||||
_summary.json # one-line summary across all runs
|
||||
@@ -37,7 +37,7 @@ ORIGINAL_HOME = os.environ.get("HERMES_HOME")
|
||||
ORIGINAL_AUTH = Path.home() / ".hermes" / "auth.json"
|
||||
|
||||
_THIS_DIR = Path(__file__).resolve().parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -16,7 +16,7 @@ from pathlib import Path
|
||||
from typing import Any, Dict, List
|
||||
|
||||
_THIS_DIR = Path(__file__).resolve().parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||
sys.path.insert(0, str(_THIS_DIR))
|
||||
|
||||
@@ -23,7 +23,7 @@ from pathlib import Path
|
||||
from typing import Any, Dict, List
|
||||
|
||||
_THIS_DIR = Path(__file__).resolve().parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||
sys.path.insert(0, str(_THIS_DIR))
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ from pathlib import Path
|
||||
from typing import Any, Dict, List
|
||||
|
||||
_THIS_DIR = Path(__file__).resolve().parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||
sys.path.insert(0, str(_THIS_DIR))
|
||||
|
||||
+1
-1
@@ -27,7 +27,7 @@ from pathlib import Path
|
||||
from typing import Any, Dict, List
|
||||
|
||||
_THIS_DIR = Path(__file__).resolve().parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
||||
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||
sys.path.insert(0, str(_THIS_DIR))
|
||||
|
||||
@@ -55,7 +55,7 @@ in real production traffic.
|
||||
## Run
|
||||
|
||||
```bash
|
||||
cd scripts/toolperf_abeval
|
||||
cd evals/toolperf_abeval
|
||||
export ABEVAL_ROOT=/tmp/abeval-workspace # results + sandboxes land here
|
||||
export ABEVAL_HOME=/tmp/abeval-home
|
||||
./run_all.sh /tmp/abeval-baseline /path/to/fixes-tree 3 \
|
||||
Reference in New Issue
Block a user