Files
hermes-agent/evals/postmortem/live_ab/hardline_scanner_matrix.py
T
Teknium a8ca904922 feat(evals): post-mortem harness — forensics lanes + live A/B + review probes for the #102117 run fixes
evals/postmortem/ turns the one-off audit behind tracking issue #103563 into
something anyone with a Hermes state.db copy (and optionally rotated
agent.log*) can run on their own fan-out:

  forensics/   common.py discovers the run tree (root = most descendants,
               compression-rollover children excluded so cost buckets stay
               disjoint), fits pricing from estimated_cost_usd, and five lanes
               recompute the OBSERVED figures: tokens (buckets, depth/duration
               shares, context reconstruction, excess-cache-write proxy, cap
               replay), logcalls (per-call cache behaviour from agent.log with
               coverage printed first; strict and loose plateau definitions
               reported separately), delegation (timeouts, orphaned children,
               polling hours, batch-join withheld child-hours, truncated
               summaries), tools (hardline blocks, foreground refusals,
               whole-file rewrites), goal_loop (nudges, parked barrier), rework
               (public-surface drop at PR open + post-open commit inventory).
               Every figure is labeled OBSERVED or MODELED.
  live_ab/     the per-PR A/Bs (real code paths, fake providers, temp
               HERMES_HOME), paths from argv.
  review_probes/ the independent /review's probes, credited and adapted; each
               reproduced a round-1 defect and the fixed head must pass it.
  run.py       runs the offline probes against one or two checkouts and prints
               PASS/FAIL side by side (--live adds the ones that spend cents).
  tests/       synthetic-DB smoke test for the lanes and runner.

On the run's DB the lanes reproduce the tracking issue's population exactly
(1,394 sessions, 93,284 calls, $19,302.59; cache_write $11,159.76) and on
main vs an integration checkout of the 13 PRs the runner shows every probe
FAIL -> PASS (two guard-only probes pass on both, noted in run.py).

The trajectories are deliberately not shipped: the DB holds 51,956 home
paths, 5,341 e-mails, private IPs, chat ids and real-shaped credentials in
tool output. The lane reports and recomputed JSON are in a secret gist
linked from #103563.
2026-09-05 09:13:10 -07:00

27 lines
1.6 KiB
Python

import sys; sys.path.insert(0, sys.argv[1] if len(sys.argv) > 1 else ".") # usage: python hardline_scanner_matrix.py <repo_root>
from tools.approval_detection import detect_hardline_command as d
cases = {
# the reviewer's witnesses
"newline-hidden reboot in quoted $(grep)": ('echo "$(grep -P \'safe\' /dev/null\nreboot)"', True),
"grep with backtick operand": ("grep -e `echo needle` file", False),
# the original class (must stay fixed)
"canonical sed/grep/cut": ('sed -n "$(grep -n X f | cut -d: -f1),+3p" f', False),
"grep -c inside $()": ('echo "$(grep -c x f)"', False),
# controls: other separators inside the substitution
"; hidden reboot": ('echo "$(grep x f; reboot)"', True),
"&& hidden reboot": ('echo "$(grep x f && reboot)"', True),
"| hidden shutdown": ('echo "$(grep x f | shutdown -h now)"', True),
"nested $( ) newline reboot": ('echo "$(echo $(grep x f)\nreboot)"', True),
"backtick newline reboot": ('echo "`grep x f\nreboot`"', True),
# data that must stay allowed
"reboot as quoted grep pattern": ("grep -F 'sudo reboot' notes.md", False),
"commit msg with reboot on a data line": ('git commit -m "fix\nsudo reboot handling"', False),
}
bad = 0
for name, (cmd, want_block) in cases.items():
got, why = d(cmd)
ok = got == want_block
bad += not ok
print(("OK " if ok else "FAIL") + f" {name}: blocked={got} ({why})")
print("ALL OK" if not bad else f"{bad} FAILURES")