Commit Graph

8 Commits

Author SHA1 Message Date
Teknium 0ee5ae61e2 docs(evals): real Codex CLI head-to-head arm + results
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same
transcripts: chunk-file reads until its REAL auto-compaction fires (verified
via compacted events in the rollout jsonl; peak 455-483K vs its 258K
window), then quizzes post-compaction with the identical question banks and
judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean
closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access
over its rollout history — the session_search differentiator, measured.
2026-08-15 18:41:24 -07:00
Teknium e2990428e7 docs(evals): ship transcript-building scripts + full eval detail in-repo
- scripts/reconstruct_lineage.py: rebuild full uncompacted lineage
  transcripts from a state.db COPY (descendant-tree walk, content-hash
  dedupe, synthetic-artifact strip, system_prompts hash resolution)
- scripts/replay_lineage.py + scripts/build_html_report.py: replay a 500K
  prefix through any checkout's compressor and render before/after
  side-by-side with compaction artifacts color-coded
- README: transcript-building workflow, scoping-tripwire section
- results/SCORECARD-2026-08-15.md: full per-transcript scorecards, all 4
  exam question banks, survival analysis, methodology + caveats
2026-08-15 17:40:43 -07:00
Teknium 9146f4c851 fix(evals): explicit encoding on all file I/O (ruff PLW1514) 2026-08-15 17:07:46 -07:00
Teknium 44536bd41b docs(evals): 4-transcript compaction scorecard
lean+recovery 68.3% avg recall @ 49K retained vs current 45.8% @ 162K —
+22.5pts at 0.30x tokens. Anchor index moved GUI needle-fact recall
23.3->60.0 closed-book, 46.7->80.0 with recovery.
2026-08-15 17:01:22 -07:00
Teknium c4bbb14e52 feat(compression): mechanical anchor index + region-scoping tripwire
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
  file paths, error strings, handles, URLs from the compacted region into a
  bounded indexed summary section. LLM-free, so needle identifiers cannot be
  paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
  golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
  summarizer input carries ONLY the compacted region (head/tail sentinels
  never reach the serialized turns body) in both legacy and lean modes.
2026-08-15 17:01:22 -07:00
Teknium 7a82457ede feat(compression): digest noise filter + FTS5 recovery sim + digest-aware query hints
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
  digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
  session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
  mine anchor identifiers
2026-08-15 17:01:22 -07:00
Teknium 8fe9025abd feat(compression): lean tail mode + recovery-aware eval arm
Lean mode (tail_mode='lean', default stays 'legacy'):
- tail budget = clamp(2.5% of window, 10K, 25K) instead of 0.20*window
- stale tail tool results demoted to session_search recovery stubs
- chunked identifier-preserving digests of the compacted region (map-reduce,
  pristine pre-prune tool contents)
- verbatim user messages embedded in summary (codex retention-by-role rule)
- deterministic session_search recovery footer

Eval: policies matrix gains lean + a '+recovery' arm giving the answerer one
simulated session_search round-trip against the archived region.
2026-08-15 17:01:22 -07:00
Teknium 33242d5ee0 feat(evals): compaction recall eval harness
Measures recall accuracy vs tokens retained across compaction policies.
Real transcripts in, LLM-generated recall exam from the summarized region,
per-policy answer+judge passes, scorecard out.
2026-08-15 17:01:21 -07:00