cf64ca20c5
The rotation path flushes its un-persisted transcript to the parent (#47202) and only then calls publish_compression_child. The abort handler rolls back the in-memory transcript and keeps agent.session_id on the parent - its own comment says "keep the parent live and discard the stale compacted snapshot" - but the rows the flush just wrote are not part of what it discards. Every failed rotation therefore leaves the parent transcript longer than it found it, whatever the failure was. That is survivable for a one-off failure and pathological for a sticky one. A parent row carrying ended_at fails the publish on every attempt and nothing in this path clears it, so each auto-compaction appends another copy of the current turn to the transcript it was supposed to shrink. Worse, the growth then satisfies conversation_compression's own len(durable_parent) > len(messages) check, so the next attempt adopts the inflated snapshot as if it were genuine concurrent activity and the in-memory transcript doubles too. Check that one precondition before writing. It is a plain read of the row the publish is about to read anyway, and it raises the publish's own message, so split_status=aborted, failure_class=session_split_failed and the rollback path are all unchanged; a live parent reaches the flush exactly as before. Deliberately not extended to the compression lease, which is re-acquirable - a transient miss there would abort a rotation that would otherwise have committed. old_session_id moves above the flush so a failure raised from here takes the same in-memory rollback as any other pre-publish failure. Scope: this fixes the amplification for every abort cause. It does not fix what marks a live session as ended in the first place (#88197 Bug 1), which needs a maintainer decision on end-reason taxonomy and is tracked on the issue; an affected session still aborts every attempt, it just stops making itself larger while it does. Refs #88197