The first compaction of a session runs check_compression_model_feasibility()
inside compress_context() BEFORE _announce_compression_start(). That probe is
network-bound (live model catalog / provider lookups; connect timeouts stack up
through proxies and slow remote gateways), so every automatic entrypoint that
reaches compress_context without its own pre-emit — the post-tool gate in
agent/turn_preflight.py::compress_after_tool_results, overflow recovery,
manual /compress — left the client with no `kind="compacting"` status for the
whole probe. On Desktop that is a bare working-row spinner with no
"Summarizing thread" label (#111294).
Move the announcement ahead of the probe in the one choke point so every
caller is covered, and retire the announced phase (force_terminal) when the
probe's hard rejection propagates so the client never stays "compacting" for
an attempt that never started.
Live repro: /tmp probe driving compress_after_tool_results with a 2 s
feasibility stand-in — before: first compacting status at t=2.29 s (after the
block); after: t=0.13 s.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
Nine f-string sites minted `YYYYMMDD_HHMMSS_<hex>` independently with the hex width already
drifted (6 on CLI/TUI/agent/import, 8 in the gateway store, 12 in portability imports).
hermes_cli/session_lost_and_found.py classifies schema-less salvage rows by that shape, so a
site drifting the prefix would silently change recovery. hermes_state_ids.new_session_id(now,
hex_len=) is now the only writer and owns SESSION_ID_PATTERN; stdlib-only so agent/, cli.py and
gateway/ can import it without the SessionDB graph.
Widths are kept per site on purpose: the Desktop's session-id candidate regex is pinned to 6 hex
chars for interactive ids; the gateway store and portability importer keep 8/12 (more rows per
second). Not a bug, so not "fixed".
gateway/platforms/qqbot/adapter.py hard-coded `agent:main:qqbot:<scene>:<chat>` for the
update-prompt authz key, ignoring the profile namespace build_session_key applies; a secondary
bot in a multiplexed gateway got `agent:<profile>:...` keys and its clicks were rejected. The key
now comes from the one builder via BasePlatformAdapter._source_session_key.
Behavior change: QQ update-prompt clicks are authorized under the profile-namespaced key
(byte-identical `agent:main:` for the default profile).
The publish-side taxonomy (_end_stamp_class, _compression_parent_obstacle,
compression_parent_deliberately_ended) and the agent-guard delegation produced
exactly main's verdict -- every non-automatic stamp still fails closed -- so
they only changed an error string. Inline the "explicit close with no
continuation" test into reopen_if_explicitly_closed(), restore main's publish
branch and agent guard untouched, and keep two tests: the field shape rotates
after the host clears the stamp; boundary/compression/automatic stamps and a
session already claimed for teardown are never cleared.
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).
- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
(:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).
Closes#97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
Both steer sites now build the row through one helper, prompt_builder.steer_user_row:
a role:user row with display_kind="steer" and no leading blank lines. The alternation
repair (_merge_consecutive_users) skips a steer-typed prev row, so a run that ended
right after a steered batch (Ctrl-C, interrupt) does not get the next real prompt
merged INTO the already-persisted steer row — which would have rewritten it in place
and re-broken live≠replay parity, the exact class this PR fixes.
TUI/desktop history projects the steer row as the user's own words instead of the
model-facing marker wrapper; 'steer' joins the display_kind union. The compression
anchor scan keeps its tool-row branch for transcripts persisted before this change and
its docstring says so.
Lowering the session trigger must not replace the window-relative lean
selection budget with threshold times target_ratio. Invalidate the lean
cache through the existing property while preserving explicit legacy and
external-engine fallback behavior.
Narrow adaptation of the aux-sync diagnosis and invariants in #93576,
without adding a required recalibration method to context engines.
Related: #95681, #93576
Co-authored-by: Turgut Kural <58116817+TurgutKural@users.noreply.github.com>
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).
- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
resumed turn while the durable transcript still matches, cleared with the row on
compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).
Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
The codex_app_server runtime bypasses the conversation loop, so the
usage-anchored context accounting captured there never ran: agent._usage_anchor
stayed None forever. Every preflight estimate therefore fell back to the rough
mirror-transcript heuristic, which is deliberately never compacted on this
runtime (_record_codex_app_server_compaction preserves the mirror), so it grows
monotonically while the real thread may be tiny or freshly compacted. With
compression.codex_app_server_auto=hermes that estimate alone tripped the
threshold and fired thread/compact/start on nearly every turn of a long-lived
session (#100381).
Mirror the main loop's post-response capture: _record_codex_app_server_usage
now snapshots the reported thread usage into agent._usage_anchor (base prompt
+ completion exactly as the provider counted them, with estimation confined to
messages appended since). A usage-less turn keeps the previous anchor, and the
post-compaction invalidation site is unchanged.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Completes the constant + samples + noise-regex trio teknium named as the bar
for this heartbeat on #98371; the Telegram noise-filter and TUI retag
parametrised suites now iterate the heartbeat wording too.
Review follow-ups: the heartbeat opened a visible compacting phase even when
the context engine suppressed the routine start status (no terminal edge would
ever close it), and its start() emitted a second start line milliseconds after
the routine one (two chat messages on adapters without send_or_update_status).
Gate the client-visible heartbeat on the start status having been emitted, drop
the start-time emit, and route ticks through agent._emit_status so CLI print
and gateway filtering match every other compaction status.
The contributor heartbeat called status_callback("compacting", <ad-hoc text>).
Every other compaction status uses the "lifecycle" key: the TUI gateway
re-tags lifecycle statuses to kind="compacting" via
is_compaction_progress_status, Telegram edits one bubble per status key, and
gateway/run.py's chat-platform filter only recognises the registered
templates — so the ad-hoc text leaked to Telegram/Discord on every tick with
compression.progress_notices off (verified: _prepare_gateway_status_message
passed it through). Register COMPACTION_HEARTBEAT_STATUS next to
COMPACTION_STATUS, emit it under "lifecycle", and cover the filter.
Context compression can stream for minutes with no deltas, tool events,
or status lines reaching remote transports. Idle-progress watchdogs on
those clients treat the silence as a dead turn and interrupt it — the
Android relay app's 180s turn watchdog fires session.interrupt, killing
a healthy compression mid-flight and rolling back its work. On sessions
near the context ceiling this loops forever: every new prompt retriggers
preflight compression, which dies at exactly +180s again (observed
telemetry: attempts aborted at 180469ms/180233ms/180219ms with
failure_class=explicit_interrupt).
Fix: the existing _CompressionActivityHeartbeat (which today only
refreshes the SessionDB activity tracker) now also emits a 'compacting'
status through agent.status_callback — once at compression start and on
every heartbeat tick. The gateway already routes status_callback to
status.update events, and clients already reset their watchdogs on any
received event, so each heartbeat re-arms them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6d2e0e860d6fdb44ce549d1fa11a06662b1a4b0b)
Review follow-ups on the in-flight replay:
- Run _sanitize_tool_pairs BEFORE the re-append. Its trailing-in-flight
exemption (#79278) walks back from the list end; with the replay user row
there, a genuinely pending assistant(tool_calls) looked orphaned and had its
calls stripped, so the executor's late tool result was dropped.
- When the restatement is merged onto a user-pinned summary carrier, flag the
carrier (_inflight_replay_merged). The carrier's metadata marks it synthetic,
so conversation_compression._ensure_compressed_has_user_turn inserted a second
copy of the same request; it now treats the flag as intent-present, and the
next cycle recognises the carrier as the task instead of losing it.
- Header idempotency: restate the text after the last header so a task that
survives several compactions carries one header and one copy.
- Exclude metadata-flagged scaffolding (_todo_snapshot_synthetic, recovery
nudges) from the in-flight scan via the shared _is_real_user_message.
Tests cover all four; each fails with its fix reverted.