feat(compression): add opt-in idle-triggered context compaction

Long-lived sessions (e.g. a Telegram thread resumed over hours/days)
accumulate a large context that the existing size-based threshold only
trims once it crosses `threshold × context_window`. Until then every
turn re-reads the full history, which on large-context models can mean
hundreds of K of cache-read tokens per call even across long idle gaps.

Add a time-based trigger that complements (does not replace) the size
threshold: when a session resumes after `compression.idle_compact_after_seconds`
of inactivity, compact the accumulated history up front, before the first
reply. Disabled by default (0), so existing behaviour is unchanged.

The trigger reuses `_last_activity_ts` (the last time the turn loop did
work) to measure the idle gap at turn start, gates the token estimate
behind a cheap gap pre-check, and skips compaction when the context is
already at/below the post-compression target (threshold × target_ratio)
so a short idle thread never pays for a summarization that saves nothing.
It also defers to an active compression-failure cooldown.

The decision is factored into a pure predicate, `_should_idle_compact`,
which is unit-tested without a live agent.
This commit is contained in:
Muthuvel
2026-07-01 01:48:46 +08:00
committed by Teknium
parent 76e17bc32d
commit 72056faf8f
4 changed files with 165 additions and 0 deletions
+10
View File
@@ -474,6 +474,16 @@ compression:
# head messages, matching the pre-feature behaviour.
protect_first_n: 3
# Idle compaction (default: 0 = disabled). When > 0, a session that resumes
# after at least this many seconds of inactivity compacts its accumulated
# history up front, before the first reply, so a long-lived thread you come
# back to later doesn't re-read its full stale context on every turn.
# Time-based, so it complements (does not replace) the size-based `threshold`
# above. It is skipped when the context is already small (at or below the
# post-compression target = threshold × target_ratio), so it never wastes a
# summarization on a short idle thread. Example: 1800 = compact after 30 min idle.
idle_compact_after_seconds: 0
# To pin a specific model/provider for compression summaries, use the
# auxiliary section below (auxiliary.compression.provider / model).