Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).
Now there is one authority:
- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
over threshold waits ONE request for the provider's real count instead of compressing on a guess
(first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
(note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
past the whole window, and provider-proven overflow all compress immediately; the post-compaction
latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
that scripted whole-history estimates now state the fact they relied on (provider omits usage).
Fixes#103391 (closes#103397 by construction — the baseline it repaired no longer exists).
The salvaged helper was appended to the gateway/run.py facade; new behaviour
belongs in a topical sibling. agent/session_activity.py already owns the
activity-snapshot contract the three render sites read from, so the formatter
moves there as format_iteration_progress and the three call sites import it at
module level instead of late-importing the facade.
Tests trimmed to the salvage bar (<= 2 invariant tests per fix): one
parameterized contract on the helper (unbounded/None hide the ceiling, a real
budget keeps N/M) plus the contributor's end-to-end busy-ack test through the
real render path. The per-site heartbeat/timeout tests exercised the same
helper through mocks and are dropped; evals/gateway_status_render/
iteration_ceiling_ab.py drives all three real render sites with a real AIAgent
for before/after evidence (3 sentinel leaks on main -> 0).
Related: #103109 (same fix, same target module; credited in the PR),
#102845 (third render site, closed by its author in favour of #102817).
Rebase of the existing fix onto current main: the god-file split that
was in flight when this PR opened has landed, moving both original
call sites out of gateway/run.py into gateway/run_busy.py and
gateway/run_turn.py, which is why this PR showed a merge conflict.
Also folds in a third call site (see below) that a competing PR
(#102845, closed by its author in favor of this one) identified after
this PR first opened.
Three user-facing gateway status lines render "iteration N/M" from
AIAgent.get_activity_summary()'s api_call_count / max_iterations pair:
the long-running heartbeat (run_turn.GatewayTurnMixin.
_run_agent_notify_long_running), the busy-session acknowledgment
(run_busy.GatewayBusySessionMixin._compose_busy_ack_message), and the
gateway-timeout diagnostic message shown to the user when a run is
force-timed-out for inactivity (run_turn.GatewayTurnMixin.
_run_agent_timeout_result). AIAgent.max_iterations defaults to
sys.maxsize (unlimited tool-calling iterations for a top-level session
-- see run_agent.py), so all three printed the literal
9223372036854775807 as the denominator, e.g.:
⏳ Working — 3 min — iteration 2/9223372036854775807, receiving stream response
That reads as a bug rather than "unbounded" and is meaningless to a
user.
Add _format_iteration_progress() to gateway/run.py, a small shared
formatting helper: once the configured max_iterations is at or above
sys.maxsize, it renders "iteration N" alone and omits the denominator;
a genuinely finite budget (e.g. a subagent's delegation.
max_iterations: 250) still renders "iteration N/M" as before. All
three call sites now go through it. The gateway-timeout diagnostic
message's own operator-facing logger.error() call keeps the raw
resolved value for diagnostics -- only the two user-facing diag_lines
built from it are reformatted.
Added tests/gateway/test_format_iteration_progress.py (the helper's
own unit tests: unbounded default, above-sentinel values, finite
budgets, and graceful handling of a missing/malformed max_iterations),
tests/gateway/test_gateway_timeout_iteration_progress.py (both
diagnostic-message branches, unbounded and finite, plus the heartbeat
call site exercised end to end through its async polling loop), and a
new regression test in tests/gateway/test_busy_session_ack.py::
TestBusySessionAck::test_status_detail_omits_denominator_for_unbounded_max_iterations
that exercises the busy-ack call site end to end with the real
sys.maxsize default and asserts the sentinel never reaches the
rendered text.
Defect 1 in the original report (a direct_result tool's raw output
occasionally becoming the user-visible reply) is left for a separate
change -- the reporter frames it as an open core-level design question
("a per-turn suppression option... would fix the class for all
plugins"), not a drop-in fix, and it touches reply composition rather
than status-line formatting.
Part of #102806 (defect 2: the iteration-ceiling display; defect 1 is a separate change).
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
Re-applies the gateway compat removal byte-for-byte; see 92d0bd0d73 for the
full inventory (30 re-exports/aliases + 2 shim modules dropped, 3 shim-only
names re-removed, 24 callers + 34 test files repointed). No new changes.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
AST-driven, body-identical move of 359 GatewayRunner methods into cohesive
mixin modules (gateway/run_{voice,adapters,topics,turn,shutdown,busy,
config_loaders,startup,watchers,notifications,inbound,goals,agent_cache}.py)
plus TurnRunner -> gateway/run_turn_runner.py. run.py-internal symbols are
imported lazily inside method bodies so patch('gateway.run.X') keeps
intercepting; neutral deps are top-level; logger name stays 'gateway.run'.
_UNSET moved to leaf gateway/run_common.py (def-time default-arg sentinel).
Whole-module inspect.getsource(gateway_run) AST-walker tests repointed to
the module that now holds the walked code.