Files
hermes-agent/agent
Teknium c49e2a496e fix(compression): route-aware pruned estimate at remaining pressure sibling sites
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.

- agent/turn_context.py idle-triggered compaction: use
  _preflight_request_tokens (anchor -> native pruned -> generic) instead
  of the raw generic request estimate, so resuming a compacted codex
  session after an idle gap does not fire a compaction the next request
  never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
  the warn site's route-aware figure so the dedup re-arms correctly on
  native sessions.
- agent/conversation_loop.py post-response should_compress fallback
  (last_prompt_tokens==0, i.e. no provider usage after a disconnect or
  gateway restart — the unanchored case in #97602's repro): route through
  _midturn_request_pressure_tokens instead of the generic figure.

Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
2026-08-30 05:15:46 -07:00
..