a8ca904922
evals/postmortem/ turns the one-off audit behind tracking issue #103563 into something anyone with a Hermes state.db copy (and optionally rotated agent.log*) can run on their own fan-out: forensics/ common.py discovers the run tree (root = most descendants, compression-rollover children excluded so cost buckets stay disjoint), fits pricing from estimated_cost_usd, and five lanes recompute the OBSERVED figures: tokens (buckets, depth/duration shares, context reconstruction, excess-cache-write proxy, cap replay), logcalls (per-call cache behaviour from agent.log with coverage printed first; strict and loose plateau definitions reported separately), delegation (timeouts, orphaned children, polling hours, batch-join withheld child-hours, truncated summaries), tools (hardline blocks, foreground refusals, whole-file rewrites), goal_loop (nudges, parked barrier), rework (public-surface drop at PR open + post-open commit inventory). Every figure is labeled OBSERVED or MODELED. live_ab/ the per-PR A/Bs (real code paths, fake providers, temp HERMES_HOME), paths from argv. review_probes/ the independent /review's probes, credited and adapted; each reproduced a round-1 defect and the fixed head must pass it. run.py runs the offline probes against one or two checkouts and prints PASS/FAIL side by side (--live adds the ones that spend cents). tests/ synthetic-DB smoke test for the lanes and runner. On the run's DB the lanes reproduce the tracking issue's population exactly (1,394 sessions, 93,284 calls, $19,302.59; cache_write $11,159.76) and on main vs an integration checkout of the 13 PRs the runner shows every probe FAIL -> PASS (two guard-only probes pass on both, noted in run.py). The trajectories are deliberately not shipped: the DB holds 51,956 home paths, 5,341 e-mails, private IPs, chat ids and real-shaped credentials in tool output. The lane reports and recomputed JSON are in a secret gist linked from #103563.
27 lines
1.6 KiB
Python
27 lines
1.6 KiB
Python
import sys; sys.path.insert(0, sys.argv[1] if len(sys.argv) > 1 else ".") # usage: python hardline_scanner_matrix.py <repo_root>
|
|
from tools.approval_detection import detect_hardline_command as d
|
|
cases = {
|
|
# the reviewer's witnesses
|
|
"newline-hidden reboot in quoted $(grep)": ('echo "$(grep -P \'safe\' /dev/null\nreboot)"', True),
|
|
"grep with backtick operand": ("grep -e `echo needle` file", False),
|
|
# the original class (must stay fixed)
|
|
"canonical sed/grep/cut": ('sed -n "$(grep -n X f | cut -d: -f1),+3p" f', False),
|
|
"grep -c inside $()": ('echo "$(grep -c x f)"', False),
|
|
# controls: other separators inside the substitution
|
|
"; hidden reboot": ('echo "$(grep x f; reboot)"', True),
|
|
"&& hidden reboot": ('echo "$(grep x f && reboot)"', True),
|
|
"| hidden shutdown": ('echo "$(grep x f | shutdown -h now)"', True),
|
|
"nested $( ) newline reboot": ('echo "$(echo $(grep x f)\nreboot)"', True),
|
|
"backtick newline reboot": ('echo "`grep x f\nreboot`"', True),
|
|
# data that must stay allowed
|
|
"reboot as quoted grep pattern": ("grep -F 'sudo reboot' notes.md", False),
|
|
"commit msg with reboot on a data line": ('git commit -m "fix\nsudo reboot handling"', False),
|
|
}
|
|
bad = 0
|
|
for name, (cmd, want_block) in cases.items():
|
|
got, why = d(cmd)
|
|
ok = got == want_block
|
|
bad += not ok
|
|
print(("OK " if ok else "FAIL") + f" {name}: blocked={got} ({why})")
|
|
print("ALL OK" if not bad else f"{bad} FAILURES")
|