refactor(compression): scope Codex-native compaction to the app-server runtime

Drop the Responses-API native compaction path and its opt-in umbrella
flag from the salvaged feature. On the Codex OAuth chat route Hermes
owns the message list and the summary compressor works (and stays
provider-portable — encrypted compaction items would lock the session
history to chatgpt.com and break /model switches and provider
fallback). On the app-server runtime (codex CLI/agent) the codex agent
owns the real thread context, so thread/compact/start is the only
mechanism that can actually shrink it (#36801) — that path is now the
default behavior for codex_app_server sessions, controlled by
compression.codex_app_server_auto (native|hermes|off), no umbrella
flag.

Removed: responses.compact() call path, codex_compaction_items replay/
persistence plumbing, codex_native_compaction + codex_responses_threshold
config keys, desktop settings fields, and their tests. Kept: everything
app-server (compact_thread(), compaction notifications, bookkeeping,
docs, tests) plus cache-busting keys for the surviving knobs.
This commit is contained in:
teknium1
2026-07-07 02:11:23 -07:00
committed by Teknium
parent d1c8c03416
commit 87b65e24a7
5 changed files with 13 additions and 51 deletions
+3 -13
View File
@@ -415,17 +415,6 @@ compression:
# for the ChatGPT Codex OAuth route. Set false to opt back down to threshold.
codex_gpt55_autoraise: true
# Codex-native compaction paths are opt-in (default: false) to preserve
# Hermes' existing auxiliary summarizer behavior for existing installs.
# When true, Codex OAuth uses Responses API compact and Codex app-server
# compaction uses the app-server thread compact API.
codex_native_compaction: false
# Codex OAuth / Responses API compaction trigger (default: 0.85 = 85%).
# Used only when codex_native_compaction is true. The recent tail still uses
# protect_last_n below.
codex_responses_threshold: 0.85
# Fraction of the threshold to preserve as recent tail (default: 0.20 = 20%)
# e.g. 20% of 50% threshold = 10% of total context kept as recent messages.
# Summary output is separately capped at 12K tokens (Gemini output limit).
@@ -437,8 +426,9 @@ compression:
# compression of older turns.
protect_last_n: 20
# Codex app-server auto-compaction mode:
# Used only when codex_native_compaction is true.
# Codex app-server (codex CLI runtime) thread-compaction mode. The codex
# agent owns the real thread context on this runtime, so Hermes' summarizer
# cannot shrink it — compaction goes through the app server instead.
# native = let Codex decide when to compact its own thread (default)
# hermes = let Hermes threshold trigger Codex thread/compact/start
# off = Hermes will not auto-trigger compaction; Codex may still compact natively