test_7100 replicated the phrase list inline, so it tested its own copy rather
than production; point it at is_context_overflow_failure_result. Drop a dead
`error=` parameter from the normalizer test helper.
Moving the `if response` guard below the failed branch exposed the normalizer's
loose overflow predicate (bare "token"/"exceed"/"context"/"payload", or any 400
on a long session) to failed turns that carry real text: billing, rate-limit,
auth and content-policy replies were rewritten to "Session too large / /compact".
Hoist run_turn's stricter classifier (compression_exhausted, multi-word phrases,
400 on history > 50) into a module-level `is_context_overflow_failure_result` and
use it for both the #1630 transcript skip and the user-facing rewrite, so the two
can never disagree. Populated text is only rewritten when it is the bare provider
envelope (`_looks_like_gateway_provider_error`) on an overflow turn; curated agent
text (compression-timeout guidance, /compress hint) survives. Replaces the
sanitizer-wording test with the passthrough invariants that catch the regression.
Two defects in _normalize_empty_agent_response surfaced together during a
state.db lock-contention incident on an enterprise Slack deployment:
- the error lookup used dict.get's default, which an explicit
'error': None value bypasses, rendering 'The request failed: None' /
'unknown error';
- persistence failures fell through to the generic branch, whose 'use
/reset' advice is harmful for this failure mode (destroys conversation
context, fixes nothing).
Persistence-failed turns (failure_reason session_persistence_failed:*,
with a legacy fallback on the error text) now get a dedicated message:
storage was temporarily unavailable, the message was recorded, send it
again — with a disk-specific variant. No /reset suggestion. All other
branches unchanged.