Files
hermes-agent/agent
Darafei Praliaskouski f07b6ff426 fix(agent): count non-CJK sparse text by UTF-8 bytes in the rough estimator
The ~4 chars/token rule is calibrated for ASCII; Cyrillic, Greek, Arabic and
similar 2-byte scripts tokenize at ~2-3 chars/token, so chars/4 under-counts
them ~2x and the pre-flight pressure figure trails real usage by tens of
percent on non-English sessions. Counting UTF-8 BYTES at ~4/token uses the
encoding width itself as the corrective: ASCII is unchanged (1 byte/char),
2-byte scripts count at chars/2, and the CJK dense path keeps its explicit
~1 token/char rule with the sparse remainder byte-counted. The ASCII
isascii() O(1) fast path is preserved; the non-ASCII paths add a single
C-level encode over text that was already being regex-scanned.

Complements the last-real-prompt floor: the floor catches sessions that are
already at the ceiling, this keeps the estimate from lagging in the first
place.
2026-09-03 03:09:06 +05:30
..