A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.
Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
main-model fallback (a larger output budget may finish the summary), and
on terminal failure ABORTS compression preserving the session unchanged
(new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
existing retry/backoff loop.
_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.
Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.
Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).