8aeb3f6ee3
The permanent-brick class in #69078: xAI returns 'Invalid PNG image' when a re-serialized image part in replayed history becomes undecodable. The existing image-error patterns cover only Anthropic 'exceeds max dimension' wordings and 'model does not support images' strings, so the classifier lands on a generic non-retryable 400 and neither the shrink path nor the strip path fires. Every subsequent turn (even bare text) fails identically because the poison stays in history — the session is permanently wedged until deleted. Two recovery layers, deliberately separate: - Semantic split: new FailoverReason.image_corrupt with _IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'), checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and _classify_by_message. Corrupt bytes route to strip-and-retry, never to the shrink path (shrinking corrupt bytes cannot help). - Generic fallback: any non-retryable 400 whose outgoing messages still contain image parts gets one strip-and-retry via the existing _strip_images_from_messages helper, guarded by a new stripped_images_this_turn one-shot flag on TurnRetryState. This un-bricks the session for any current or future provider wording without adding another pattern list to maintain. Item 3 from the report (multimodal-part integrity across FTS persistence + compaction handoff) is a separate investigation and remains follow-up work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: paultaki <paultaki@users.noreply.github.com>