fix(agent): never floor an anchored pressure figure; keep the estimator total on lone surrogates

Follow-ups from review of the two salvaged #87490 commits:

- _pressure_with_real_floor now applies only on the rough fallback branch.
  A valid usage anchor is provider-exact and wins as-is: on MoA turns the
  anchor deliberately uses the pre-fold aggregator usage while
  last_real_prompt_tokens holds the folded figure, so flooring the anchored
  value would re-add fan-out tokens the anchor exists to exclude. Docstring
  rewritten to describe the real path split (anchor since d3a1c46510).
- estimate_tokens_rough: encode with errors="replace". main's estimator
  never raised; text.encode() on a lone surrogate (routine in tool output,
  see message_sanitization) raised UnicodeEncodeError and would abort a
  turn where main produced a slightly-off number.
- Record the cl100k/o200k/Qwen2.5 calibration for the bytes/4 rule.
- tests: accented Latin within +10% of the ASCII rule; mixed Cyrillic/ASCII
  counts ASCII at one byte; lone surrogates don't raise; anchored pressure
  is never floored (wiring shape).
This commit is contained in:
kshitijk4poor
2026-09-03 02:48:58 +05:30
committed by kshitij
parent f07b6ff426
commit a1d5a976b3
4 changed files with 79 additions and 13 deletions
+11 -2
View File
@@ -3662,10 +3662,19 @@ def estimate_tokens_rough(text: str) -> int:
# token) where chars/4 under-counted them ~2x and let sessions ride
# the provider's context ceiling below the compaction threshold.
# ASCII spans inside mixed text still count at 1 byte each.
return (len(text.encode("utf-8")) + 3) // 4
#
# Calibrated against cl100k/o200k/Qwen2.5 (estimate / mean real):
# Russian 0.67->1.24, Ukrainian 0.55->1.03, Arabic 0.53->0.96,
# Hindi 0.34->0.90, Greek 0.37->0.68, Polish 0.63->0.69; accented
# Latin barely moves (French 1.02->1.03, German 0.99->1.02,
# Spanish 1.04->1.07) because only the accented chars widen.
# Pure-ASCII prose already over-counts at ~1.4 on the same rule.
# errors="replace": lone surrogates (routine in tool output; see
# message_sanitization) must not turn an estimate into a raise.
return (len(text.encode("utf-8", "replace")) + 3) // 4
# Mixed CJK + other: dense chars stay ~1 token each; the sparse
# remainder is byte-counted for the same corrective.
return dense + ((len(stripped.encode("utf-8")) + 3) // 4)
return dense + ((len(stripped.encode("utf-8", "replace")) + 3) // 4)
def estimate_messages_tokens_rough(