fix(agent): never floor an anchored pressure figure; keep the estimator total on lone surrogates
Follow-ups from review of the two salvaged #87490 commits:
- _pressure_with_real_floor now applies only on the rough fallback branch.
A valid usage anchor is provider-exact and wins as-is: on MoA turns the
anchor deliberately uses the pre-fold aggregator usage while
last_real_prompt_tokens holds the folded figure, so flooring the anchored
value would re-add fan-out tokens the anchor exists to exclude. Docstring
rewritten to describe the real path split (anchor since d3a1c46510).
- estimate_tokens_rough: encode with errors="replace". main's estimator
never raised; text.encode() on a lone surrogate (routine in tool output,
see message_sanitization) raised UnicodeEncodeError and would abort a
turn where main produced a slightly-off number.
- Record the cl100k/o200k/Qwen2.5 calibration for the bytes/4 rule.
- tests: accented Latin within +10% of the ASCII rule; mixed Cyrillic/ASCII
counts ASCII at one byte; lone surrogates don't raise; anchored pressure
is never floored (wiring shape).
This commit is contained in:
+11
-2
@@ -3662,10 +3662,19 @@ def estimate_tokens_rough(text: str) -> int:
|
||||
# token) where chars/4 under-counted them ~2x and let sessions ride
|
||||
# the provider's context ceiling below the compaction threshold.
|
||||
# ASCII spans inside mixed text still count at 1 byte each.
|
||||
return (len(text.encode("utf-8")) + 3) // 4
|
||||
#
|
||||
# Calibrated against cl100k/o200k/Qwen2.5 (estimate / mean real):
|
||||
# Russian 0.67->1.24, Ukrainian 0.55->1.03, Arabic 0.53->0.96,
|
||||
# Hindi 0.34->0.90, Greek 0.37->0.68, Polish 0.63->0.69; accented
|
||||
# Latin barely moves (French 1.02->1.03, German 0.99->1.02,
|
||||
# Spanish 1.04->1.07) because only the accented chars widen.
|
||||
# Pure-ASCII prose already over-counts at ~1.4 on the same rule.
|
||||
# errors="replace": lone surrogates (routine in tool output; see
|
||||
# message_sanitization) must not turn an estimate into a raise.
|
||||
return (len(text.encode("utf-8", "replace")) + 3) // 4
|
||||
# Mixed CJK + other: dense chars stay ~1 token each; the sparse
|
||||
# remainder is byte-counted for the same corrective.
|
||||
return dense + ((len(stripped.encode("utf-8")) + 3) // 4)
|
||||
return dense + ((len(stripped.encode("utf-8", "replace")) + 3) // 4)
|
||||
|
||||
|
||||
def estimate_messages_tokens_rough(
|
||||
|
||||
Reference in New Issue
Block a user