a8ca904922
evals/postmortem/ turns the one-off audit behind tracking issue #103563 into something anyone with a Hermes state.db copy (and optionally rotated agent.log*) can run on their own fan-out: forensics/ common.py discovers the run tree (root = most descendants, compression-rollover children excluded so cost buckets stay disjoint), fits pricing from estimated_cost_usd, and five lanes recompute the OBSERVED figures: tokens (buckets, depth/duration shares, context reconstruction, excess-cache-write proxy, cap replay), logcalls (per-call cache behaviour from agent.log with coverage printed first; strict and loose plateau definitions reported separately), delegation (timeouts, orphaned children, polling hours, batch-join withheld child-hours, truncated summaries), tools (hardline blocks, foreground refusals, whole-file rewrites), goal_loop (nudges, parked barrier), rework (public-surface drop at PR open + post-open commit inventory). Every figure is labeled OBSERVED or MODELED. live_ab/ the per-PR A/Bs (real code paths, fake providers, temp HERMES_HOME), paths from argv. review_probes/ the independent /review's probes, credited and adapted; each reproduced a round-1 defect and the fixed head must pass it. run.py runs the offline probes against one or two checkouts and prints PASS/FAIL side by side (--live adds the ones that spend cents). tests/ synthetic-DB smoke test for the lanes and runner. On the run's DB the lanes reproduce the tracking issue's population exactly (1,394 sessions, 93,284 calls, $19,302.59; cache_write $11,159.76) and on main vs an integration checkout of the 13 PRs the runner shows every probe FAIL -> PASS (two guard-only probes pass on both, noted in run.py). The trajectories are deliberately not shipped: the DB holds 51,956 home paths, 5,341 e-mails, private IPs, chat ids and real-shaped credentials in tool output. The lane reports and recomputed JSON are in a secret gist linked from #103563.
46 lines
3.4 KiB
Python
46 lines
3.4 KiB
Python
"""Live A/B for the nested-delegate deadline. A depth-1 orchestrator child dispatches one leaf that runs a
|
|
~460 s task (longer than the 420 s sequential deadline). On main the orchestrator's delegate_task call returns
|
|
'timed out after 420.0s' and the leaf runs on as an orphan; on the branch the call blocks and returns the
|
|
real result. Uses glm-5.3 via Nous for cost. Deadline shortened via config to keep the run short."""
|
|
import os, sys, json, time, re, subprocess
|
|
# Usage: python nested_delegate_deadline.py <repo_root> (run once per ref; LIVE: a couple of real child calls)
|
|
root = sys.argv[1]; arm = os.path.basename(os.path.normpath(root))
|
|
sys.path.insert(0, root)
|
|
# Temp HERMES_HOME with the real auth + a config that shortens the generic sequential deadline to 40 s, so the
|
|
# run takes ~1.5 min instead of 8. The fix exempts delegate_task from this deadline entirely, so the shortened
|
|
# value is exactly what main will hit.
|
|
import shutil, tempfile, yaml
|
|
home = tempfile.mkdtemp(prefix="dl_home_"); os.environ["HERMES_HOME"] = home
|
|
real_home = os.environ.get("HERMES_HOME_SOURCE", os.path.expanduser("~/.hermes")) # credentials are copied from here into a temp home
|
|
shutil.copy(f"{real_home}/auth.json", f"{home}/auth.json")
|
|
cfg = yaml.safe_load(open(f"{real_home}/config.yaml", encoding="utf-8")) or {}
|
|
cfg.setdefault("timeouts", {}).setdefault("tools", {})["sequential_call"] = 40
|
|
cfg.setdefault("delegation", {})["orchestrator_enabled"] = True
|
|
yaml.safe_dump(cfg, open(f"{home}/config.yaml", "w", encoding="utf-8"))
|
|
import agent.tool_executor as te
|
|
assert te.__file__.startswith(root)
|
|
from agent.deadline import resolve_timeout
|
|
print("effective sequential deadline:", resolve_timeout("tools.sequential_call", default=te._resolve_concurrent_tool_timeout()))
|
|
from run_agent import AIAgent
|
|
from hermes_cli.runtime_provider import resolve_runtime_provider
|
|
MODEL = "z-ai/glm-5.3-flash"
|
|
rt = resolve_runtime_provider(requested="nous", target_model=MODEL)
|
|
sid = f"dl_{arm}_{int(time.time())}"
|
|
ag = AIAgent(model=MODEL, provider="nous", base_url=rt.get("base_url"), api_key=rt.get("api_key"), api_mode=rt.get("api_mode"),
|
|
session_id=sid, quiet_mode=True, enabled_toolsets=["terminal", "delegation"], platform="cli", max_iterations=8,
|
|
skip_context_files=True, skip_memory=True)
|
|
# Make this agent a depth-1 orchestrator exactly as delegate_tool_child_run does for a real nested orchestrator:
|
|
# at depth>0 delegate_task runs synchronously inside the tool call, which is the path under the deadline.
|
|
ag._delegate_depth = 1
|
|
ag._delegate_role = "orchestrator"
|
|
task = ("Use delegate_task exactly once (not background) with a single task whose goal is: "
|
|
"\"Run the shell command `sleep 75 && echo LEAF_DONE_MARKER` with the terminal tool (background=false is fine, it is under the tool timeout), "
|
|
"then reply with the exact text the command printed.\" "
|
|
"When delegate_task returns, reply with ONE line: RESULT: followed by the child's summary text verbatim (or the error text if it errored).")
|
|
t0 = time.time(); r = ag.run_conversation(task); wall = time.time() - t0
|
|
final = (r.get("final_response") or "").strip()
|
|
print(f"ARM {arm}: wall={wall:.0f}s final={final[:300]!r}")
|
|
timed_out = "timed out" in final.lower()
|
|
got_marker = "LEAF_DONE_MARKER" in final
|
|
print(json.dumps({"arm": arm, "delegate_timed_out": timed_out, "orchestrator_got_leaf_result": got_marker, "wall_s": round(wall)}))
|