68bc4b216e
finish_reason='length' has two causes: the answer was long (max_tokens reached), or the prompt itself left no room to generate. _continue_text treated both the same: append the fragment + a continuation nudge and retry, up to 4 times. In the second case every retry sends a strictly longer prompt, so each attempt is worse (Ollama n_ctx=32768: 32,638 -> 32,685 -> 32,732 prompt tokens, all truncated), the user is told "model hit max output tokens", and max_tokens is not the lever. The response's usage already carries prompt_tokens and the compressor already resolves the model's context window; compare them once per truncation. Under _MIN_CONTINUATION_HEADROOM (512) free tokens the turn ends on the first truncation, keeps the partial text, names the context window as the cause and points at /compress or a larger window. Unknown usage or window keeps today's behaviour. max_tokens semantics untouched. Co-authored-by: gaoanze888 <214786078+gaoanze888@users.noreply.github.com> Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
3 lines
34 B
Plaintext
3 lines
34 B
Plaintext
gaoanze888
|
|
# PR #106223 co-author
|