0ee5ae61e2
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same transcripts: chunk-file reads until its REAL auto-compaction fires (verified via compacted events in the rollout jsonl; peak 455-483K vs its 258K window), then quizzes post-compaction with the identical question banks and judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access over its rollout history — the session_search differentiator, measured.