Files
hermes-agent/evals/postmortem/live_ab/subagent_context_cap.py
T
Teknium a8ca904922 feat(evals): post-mortem harness — forensics lanes + live A/B + review probes for the #102117 run fixes
evals/postmortem/ turns the one-off audit behind tracking issue #103563 into
something anyone with a Hermes state.db copy (and optionally rotated
agent.log*) can run on their own fan-out:

  forensics/   common.py discovers the run tree (root = most descendants,
               compression-rollover children excluded so cost buckets stay
               disjoint), fits pricing from estimated_cost_usd, and five lanes
               recompute the OBSERVED figures: tokens (buckets, depth/duration
               shares, context reconstruction, excess-cache-write proxy, cap
               replay), logcalls (per-call cache behaviour from agent.log with
               coverage printed first; strict and loose plateau definitions
               reported separately), delegation (timeouts, orphaned children,
               polling hours, batch-join withheld child-hours, truncated
               summaries), tools (hardline blocks, foreground refusals,
               whole-file rewrites), goal_loop (nudges, parked barrier), rework
               (public-surface drop at PR open + post-open commit inventory).
               Every figure is labeled OBSERVED or MODELED.
  live_ab/     the per-PR A/Bs (real code paths, fake providers, temp
               HERMES_HOME), paths from argv.
  review_probes/ the independent /review's probes, credited and adapted; each
               reproduced a round-1 defect and the fixed head must pass it.
  run.py       runs the offline probes against one or two checkouts and prints
               PASS/FAIL side by side (--live adds the ones that spend cents).
  tests/       synthetic-DB smoke test for the lanes and runner.

On the run's DB the lanes reproduce the tracking issue's population exactly
(1,394 sessions, 93,284 calls, $19,302.59; cache_write $11,159.76) and on
main vs an integration checkout of the 13 PRs the runner shows every probe
FAIL -> PASS (two guard-only probes pass on both, noted in run.py).

The trajectories are deliberately not shipped: the DB holds 51,956 home
paths, 5,341 e-mails, private IPs, chat ids and real-shaped credentials in
tool output. The lane reports and recomputed JSON are in a secret gist
linked from #103563.
2026-09-05 09:13:10 -07:00

33 lines
1.7 KiB
Python

"""Live: build a real child through delegate_tool's spawn path (real imports, temp HERMES_HOME) and read the
trigger it resolves on a 1M-window model. Run against main and the branch."""
import os, sys, tempfile, shutil
root = sys.argv[1]
sys.path.insert(0, root)
home = tempfile.mkdtemp(prefix="hh-")
os.environ["HERMES_HOME"] = home
os.environ["HERMES_STREAM_RETRIES"] = "0"
try:
from run_agent import AIAgent
import tools.delegate_tool as dt
parent = AIAgent(api_key="k", base_url="https://example.com/v1", provider="test-provider",
model="anthropic/claude-fable-5.1", quiet_mode=True, skip_context_files=True, skip_memory=True)
# find the child-construction function by name
fn = [getattr(dt, n) for n in dir(dt) if n.startswith("_") and "child" in n.lower() and callable(getattr(dt, n)) and "spawn" in (getattr(dt, n).__doc__ or "").lower() + n.lower()]
import inspect
cands = [n for n, f in inspect.getmembers(dt, inspect.isfunction) if "AIAgent(" in inspect.getsource(f)]
print("constructor fn:", cands)
f = getattr(dt, cands[0])
sig = inspect.signature(f); print("sig:", sig)
kwargs = {}
for name, p in sig.parameters.items():
if name == "parent_agent": kwargs[name] = parent
elif name == "goal": kwargs[name] = "hi"
elif p.default is inspect._empty: kwargs[name] = None
child = f(**kwargs)
child = child[0] if isinstance(child, tuple) else child
cc = child.context_compressor
print(f"window={cc.context_length:,} threshold_percent={cc.threshold_percent} trigger={cc.threshold_tokens:,} cap={cc.threshold_tokens_cap}")
print(f"parent trigger={parent.context_compressor.threshold_tokens:,}")
finally:
shutil.rmtree(home, ignore_errors=True)