refactor: move live-model benchmark harnesses from scripts/ into evals/
scripts/ is for repo tooling (tests runner, release, installers, CI checks). The tool-search live tests, the toolperf A/B eval and the browser eval benchmark are offline benchmarks that spend model budget, which is exactly what evals/ holds; evals/browser_use already cited scripts/toolperf_abeval as "the same pattern". - scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for the extra directory level; gitignore now covers evals/tool_search/out*/ - scripts/toolperf_abeval/ -> evals/toolperf_abeval/ - scripts/benchmark_browser_eval.py -> evals/browser_use/
This commit is contained in:
+2
-2
@@ -190,8 +190,8 @@ docs/superpowers/*
|
|||||||
/.install_method
|
/.install_method
|
||||||
|
|
||||||
# Tool Search live-test harness output — non-deterministic model transcripts,
|
# Tool Search live-test harness output — non-deterministic model transcripts,
|
||||||
# regenerated by scripts/tool_search_livetest.py. Never an artifact of the repo.
|
# regenerated by evals/tool_search/tool_search_livetest*.py. Never an artifact of the repo.
|
||||||
scripts/out/
|
evals/tool_search/out*/
|
||||||
|
|
||||||
# Per-release changelog drafts. These exist only transiently during a release
|
# Per-release changelog drafts. These exist only transiently during a release
|
||||||
# cut (passed to `gh release create --notes-file`); the GitHub Release itself
|
# cut (passed to `gh release create --notes-file`); the GitHub Release itself
|
||||||
|
|||||||
@@ -21,7 +21,7 @@ web tasks.
|
|||||||
rating aggregation, JS/delayed render, login chain, cross-category
|
rating aggregation, JS/delayed render, login chain, cross-category
|
||||||
compare).
|
compare).
|
||||||
- **Resume-safe.** Completed cells in `results/*.jsonl` are skipped on rerun
|
- **Resume-safe.** Completed cells in `results/*.jsonl` are skipped on rerun
|
||||||
(same pattern as `scripts/toolperf_abeval`).
|
(same pattern as `evals/toolperf_abeval`).
|
||||||
- **Backend matrix.** `orchestrate.py` drives a local headless-Chrome CDP;
|
- **Backend matrix.** `orchestrate.py` drives a local headless-Chrome CDP;
|
||||||
`orchestrate_cloud.py --backend nous-cloud|browserbase` provisions a real
|
`orchestrate_cloud.py --backend nous-cloud|browserbase` provisions a real
|
||||||
cloud browser per cell through the same provider plumbing the product uses.
|
cloud browser per cell through the same provider plumbing the product uses.
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ Runs both paths against the same live Chrome and prints a comparison table.
|
|||||||
Not a pytest — a script you run manually for the PR description.
|
Not a pytest — a script you run manually for the PR description.
|
||||||
|
|
||||||
Usage:
|
Usage:
|
||||||
.venv/bin/python scripts/benchmark_browser_eval.py [--iterations N]
|
.venv/bin/python evals/browser_use/benchmark_browser_eval.py [--iterations N]
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
"""Local-CDP battery orchestrator: tasks x arms x models x reps.
|
"""Local-CDP battery orchestrator: tasks x arms x models x reps.
|
||||||
|
|
||||||
Resume-safe: completed cells in results.jsonl are skipped, so a killed
|
Resume-safe: completed cells in results.jsonl are skipped, so a killed
|
||||||
battery continues where it left off (same pattern as scripts/toolperf_abeval).
|
battery continues where it left off (same pattern as evals/toolperf_abeval).
|
||||||
|
|
||||||
Usage:
|
Usage:
|
||||||
# start a headless Chrome first:
|
# start a headless Chrome first:
|
||||||
|
|||||||
@@ -2,14 +2,14 @@
|
|||||||
|
|
||||||
Runs five scenarios against a real model (Claude Haiku 4.5 via OpenRouter) to
|
Runs five scenarios against a real model (Claude Haiku 4.5 via OpenRouter) to
|
||||||
verify that the bridge tools work end-to-end. Records transcripts in
|
verify that the bridge tools work end-to-end. Records transcripts in
|
||||||
`scripts/out/`.
|
`evals/tool_search/out/`.
|
||||||
|
|
||||||
## Running
|
## Running
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd <repo root>
|
cd <repo root>
|
||||||
python3 scripts/tool_search_livetest.py # runs all 5 scenarios x 2 modes
|
python3 evals/tool_search/tool_search_livetest.py # runs all 5 scenarios x 2 modes
|
||||||
python3 scripts/analyze_livetest.py # side-by-side report
|
python3 evals/tool_search/analyze_livetest.py # side-by-side report
|
||||||
```
|
```
|
||||||
|
|
||||||
Requires `OPENROUTER_API_KEY` set or present in `~/.hermes/.env`.
|
Requires `OPENROUTER_API_KEY` set or present in `~/.hermes/.env`.
|
||||||
@@ -34,7 +34,7 @@ A/B baseline. The harness records:
|
|||||||
## Output structure
|
## Output structure
|
||||||
|
|
||||||
```
|
```
|
||||||
scripts/out/
|
evals/tool_search/out/
|
||||||
<scenario>__enabled.json # tool_search ON
|
<scenario>__enabled.json # tool_search ON
|
||||||
<scenario>__disabled.json # tool_search OFF
|
<scenario>__disabled.json # tool_search OFF
|
||||||
_summary.json # one-line summary across all runs
|
_summary.json # one-line summary across all runs
|
||||||
@@ -37,7 +37,7 @@ ORIGINAL_HOME = os.environ.get("HERMES_HOME")
|
|||||||
ORIGINAL_AUTH = Path.home() / ".hermes" / "auth.json"
|
ORIGINAL_AUTH = Path.home() / ".hermes" / "auth.json"
|
||||||
|
|
||||||
_THIS_DIR = Path(__file__).resolve().parent
|
_THIS_DIR = Path(__file__).resolve().parent
|
||||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
@@ -16,7 +16,7 @@ from pathlib import Path
|
|||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
_THIS_DIR = Path(__file__).resolve().parent
|
_THIS_DIR = Path(__file__).resolve().parent
|
||||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||||
sys.path.insert(0, str(_THIS_DIR))
|
sys.path.insert(0, str(_THIS_DIR))
|
||||||
|
|
||||||
@@ -23,7 +23,7 @@ from pathlib import Path
|
|||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
_THIS_DIR = Path(__file__).resolve().parent
|
_THIS_DIR = Path(__file__).resolve().parent
|
||||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||||
sys.path.insert(0, str(_THIS_DIR))
|
sys.path.insert(0, str(_THIS_DIR))
|
||||||
|
|
||||||
+1
-1
@@ -24,7 +24,7 @@ from pathlib import Path
|
|||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
_THIS_DIR = Path(__file__).resolve().parent
|
_THIS_DIR = Path(__file__).resolve().parent
|
||||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||||
sys.path.insert(0, str(_THIS_DIR))
|
sys.path.insert(0, str(_THIS_DIR))
|
||||||
|
|
||||||
+1
-1
@@ -27,7 +27,7 @@ from pathlib import Path
|
|||||||
from typing import Any, Dict, List
|
from typing import Any, Dict, List
|
||||||
|
|
||||||
_THIS_DIR = Path(__file__).resolve().parent
|
_THIS_DIR = Path(__file__).resolve().parent
|
||||||
_WORKTREE_ROOT = _THIS_DIR.parent
|
_WORKTREE_ROOT = _THIS_DIR.parents[1]
|
||||||
sys.path.insert(0, str(_WORKTREE_ROOT))
|
sys.path.insert(0, str(_WORKTREE_ROOT))
|
||||||
sys.path.insert(0, str(_THIS_DIR))
|
sys.path.insert(0, str(_THIS_DIR))
|
||||||
|
|
||||||
@@ -55,7 +55,7 @@ in real production traffic.
|
|||||||
## Run
|
## Run
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd scripts/toolperf_abeval
|
cd evals/toolperf_abeval
|
||||||
export ABEVAL_ROOT=/tmp/abeval-workspace # results + sandboxes land here
|
export ABEVAL_ROOT=/tmp/abeval-workspace # results + sandboxes land here
|
||||||
export ABEVAL_HOME=/tmp/abeval-home
|
export ABEVAL_HOME=/tmp/abeval-home
|
||||||
./run_all.sh /tmp/abeval-baseline /path/to/fixes-tree 3 \
|
./run_all.sh /tmp/abeval-baseline /path/to/fixes-tree 3 \
|
||||||
Reference in New Issue
Block a user