Commit Graph

406 Commits

Author SHA1 Message Date
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Justin Wilson 9f12121206 fix(compression): do not let prune rearm lock out over-threshold sessions
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.

Fixes #101889
2026-09-03 12:23:55 +05:30
Teknium 22ae371664 refactor(agent/conversation_loop): single-line helper signatures and guard expressions 2026-09-02 19:37:48 -07:00
Teknium ac77ef36df refactor(agent/conversation_loop): seed _LoopState from TurnContext by name, compact result dict builders and guard ladders 2026-09-02 19:23:46 -07:00
Teknium 762ff705e1 refactor(agent/conversation_loop): compact billing message assembly, _LoopState comments, prompt-identity helpers 2026-09-02 19:18:17 -07:00
Teknium f76fe598a3 refactor(agent/conversation_loop): lift the API retry loop into _run_api_retry_loop; drop blank lines after lazy imports 2026-09-02 19:11:11 -07:00
Teknium 83f8de7eab refactor(agent): flatten interrupt() claim closures, group lazy-origin imports, reflow literals/log calls byte-identically 2026-09-02 19:05:00 -07:00
Teknium b3d75f78f3 refactor(agent/conversation_loop): compact per-turn agent resets and finalize_turn call 2026-09-02 18:57:46 -07:00
Teknium 7fa7b58b3c refactor(agent): compact conversation_loop docstrings/comments, unify partial-turn result shape, tidy empty_response_guard/interrupt_compat/iteration_budget 2026-09-02 18:53:53 -07:00
Teknium 351d0444dc refactor(agent/conversation_loop): collapse boolean ladders, unify context-engine hook gate and tool-call copy-on-write 2026-09-02 18:42:39 -07:00
Teknium 004cd68e20 refactor(agent/conversation_loop,deadline): extract bot-chat staleness + prompt persist helpers, unify guidance printing, compact deadline docstrings 2026-09-02 18:34:28 -07:00
Teknium 31025f210c refactor(agent/conversation_loop): drive phase helpers through a _LoopState + _run_phase (run_conversation 597 -> ~250 LOC) 2026-09-02 18:18:15 -07:00
Teknium 2769937936 refactor(turn): lift iteration entry/announce, Nous rate guard, API interrupt, retry-restart consumer and preflight-timeout result out of run_conversation 2026-09-02 16:37:14 -07:00
Teknium a56f731ac6 refactor(turn): extract preflight gate + per-iteration transcript prep into agent/turn_preflight_gate.py, agent/turn_iteration_prep.py 2026-09-02 16:27:44 -07:00
Christopher cf86c47624 fix(agent): evict stale screenshot payloads before send
Call the existing keep-newest vision retirement on the per-call
api_messages clone after sanitization so OpenAI-style tool-result
screenshots are not re-uploaded on every later turn.
2026-09-03 04:51:51 +05:30
Teknium 9780739e12 refactor(turn): extract per-iteration request assembly (api_messages/MoA/cache plan/pressure) into agent/turn_request_assembly.py 2026-09-02 16:10:07 -07:00
Teknium 9fda4e5bac refactor(turn): extract retry-loop API error handler, request build, provider call and response check into agent/turn_api_*.py + agent/turn_response_check.py 2026-09-02 16:07:51 -07:00
Teknium 0dc36e0934 refactor(turn): extract tool round + final text response branches into agent/turn_tool_round.py, agent/turn_final_response.py 2026-09-02 15:40:17 -07:00
Teknium 67ef2e50fe refactor(turn): extract response intake (normalize/hooks/scratchpad/codex-incomplete) into agent/turn_response_intake.py 2026-09-02 15:37:47 -07:00
Teknium d42282ae9e refactor(turn): extract outer-loop exception handler into agent/turn_loop_errors.py 2026-09-02 15:29:00 -07:00
kshitijk4poor a1d5a976b3 fix(agent): never floor an anchored pressure figure; keep the estimator total on lone surrogates
Follow-ups from review of the two salvaged #87490 commits:

- _pressure_with_real_floor now applies only on the rough fallback branch.
  A valid usage anchor is provider-exact and wins as-is: on MoA turns the
  anchor deliberately uses the pre-fold aggregator usage while
  last_real_prompt_tokens holds the folded figure, so flooring the anchored
  value would re-add fan-out tokens the anchor exists to exclude. Docstring
  rewritten to describe the real path split (anchor since d3a1c46510).
- estimate_tokens_rough: encode with errors="replace". main's estimator
  never raised; text.encode() on a lone surrogate (routine in tool output,
  see message_sanitization) raised UnicodeEncodeError and would abort a
  turn where main produced a slightly-off number.
- Record the cl100k/o200k/Qwen2.5 calibration for the bytes/4 rule.
- tests: accented Latin within +10% of the ASCII rule; mixed Cyrillic/ASCII
  counts ASCII at one byte; lone surrogates don't raise; anchored pressure
  is never floored (wiring shape).
2026-09-03 03:09:06 +05:30
Darafei Praliaskouski 73b8ec3cea fix(agent): floor pre-API compaction pressure at the last real prompt size
The chars/4 rough estimate under-counts Cyrillic and other non-ASCII scripts
by up to ~2x, so a session can ride the provider's real context ceiling while
the rough pressure stays under the compaction threshold. On providers that
silently clip over-window prompts (ollama /v1) the reactive overflow handler
never fires either, and the length-continuation retry path re-enters the API
call without passing the post-response gate — reproducing the truncation
death spiral this branch already addresses (observed live after the first
commit: real prompts 64,842 -> 64,995 against a 55,705 threshold, output
room shrinking 694 -> 541 tokens).

Floor the pre-API pressure figure at the provider's last reported
prompt_tokens — authoritative, script-independent — except for the one turn
after a compaction when that value is known-stale (#36718's
awaiting_real_usage_after_compression window).
2026-09-03 03:09:06 +05:30
Teknium ab48b1ddc5 refactor(agent): extract per-call API message build into turn_context.build_api_messages 2026-09-02 13:30:23 -07:00
Teknium 30b0759bb3 refactor(agent): extract content_filter refusal handling into turn_truncation.handle_content_policy_refusal 2026-09-02 13:30:23 -07:00
Teknium 4f53d7bd28 refactor(agent): extract post-tool-call compression decision into turn_preflight.compress_after_tool_results 2026-09-02 13:30:23 -07:00
Teknium f0d2355e26 refactor(agent): extract text-response stop gates (verify-on-stop, pre_verify hook, kanban guard) into agent/turn_stop_gates.py 2026-09-02 13:30:23 -07:00
Teknium 22cba7dd41 refactor(agent): extract classified-error routing (compaction gate, long-context tier, eager/auth fallback, Nous 429) into turn_recovery.route_classified_error 2026-09-02 13:30:23 -07:00
Teknium ef9dc88348 refactor(agent): drop write-only _bot_capability_refreshed attribute (zero readers repo-wide) 2026-09-02 13:30:22 -07:00
Teknium 7df9d3be5c refactor(agent): extract pre-API preflight compression gate into agent/turn_preflight.py 2026-09-02 13:30:21 -07:00
Teknium 0028951f35 refactor(agent): extract tool-call name/JSON validation into agent/turn_tool_validation.py 2026-09-02 13:30:21 -07:00
Teknium d398d91529 refactor(agent): extract empty/thinking-only response recovery ladder into agent/turn_empty_response.py 2026-09-02 13:30:20 -07:00
Teknium df42f37e09 refactor(agent): extract response-shape validation, invalid-response diagnostics and Codex incomplete continuation 2026-09-02 13:30:19 -07:00
Teknium 8b958e1b0f refactor(agent): unify interruptible backoff sleeps and fallback-restart arming; extract compute_error_backoff 2026-09-02 13:30:18 -07:00
Teknium 661c9af826 refactor(agent): extract finish_reason=length truncation recovery into agent/turn_truncation.py 2026-09-02 13:30:18 -07:00
Teknium 645db06053 refactor(agent): extract 413/context-overflow compression recovery into agent/turn_overflow.py 2026-09-02 13:30:17 -07:00
Teknium 5543e7ae0f refactor(agent): resume — extract API-error attempt logging into turn_recovery.log_api_error_attempt (verified partial work) 2026-09-02 13:30:17 -07:00
Teknium 25b165add5 refactor(agent): extract terminal API-failure result builders into agent/turn_recovery.py 2026-09-02 13:30:16 -07:00
Teknium 649c227e9e refactor(agent): extract per-response usage accounting into agent/turn_usage.py 2026-09-02 13:30:16 -07:00
Teknium fefc471deb refactor(agent): extract post-classification one-shot recovery chain into agent/turn_recovery.py 2026-09-02 13:30:15 -07:00
Teknium 79469656e7 refactor(agent): extract pre-classification API-error recovery into agent/turn_recovery.py 2026-09-02 13:29:47 -07:00
Teknium 4e548ce7a0 refactor(agent): compact turn-loop comments/docstrings to invariant statements (AST-identical) 2026-09-02 13:29:46 -07:00
globalvet2025 b36489be37 fix: gate Nous auth-refresh message behind verbose mode
The 'Nous agent key refreshed after 401' message used a bare print(),
making it always visible. The equivalent xAI/Codex, Copilot, and Anthropic
auth-refresh messages all use _buffer_vprint() (verbose-gated). This brings
the Nous path in line with the others so the routine ~15min OAuth key
refresh no longer prints noise on every retry.
2026-09-02 10:55:46 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 2a0605a807 fix(agent): reasoning-off continuation reaches the wire on the legacy chat path; reset one-shot flag per turn
Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
  chat_completions path always re-emitted extra_body.reasoning with
  enabled=True, so both reasoning_effort: none and the one-shot
  length-continuation override went out as {enabled: true, effort: none}.
  Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
  turn start so a flag armed by an interrupted/errored turn can never
  strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
  /thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
  reasoning.enabled=false) and a stale-flag turn-scope test.
2026-09-02 00:55:42 -07:00
AlexGabbia fb76fb0526 fix(agent): thinking-only length truncations no longer wedge continuations
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).

The length-continuation path handled that shape badly:
  1. the empty response was appended as an interim assistant fragment,
     poisoning the transcript until the pre-call sanitizer healed it
     (observed 3+ healings per turn on the reporting user's session);
  2. every continuation re-ran with thinking ON, re-deriving the whole
     thinking budget against a growing context, so 4 attempts still produced
     nothing and the turn died with 'Response remains truncated after 4
     continuation attempts'.

Now:
  - interim assistant fragments with no visible content are never appended
    (whichever way they got empty);
  - a thinking-only truncation sets a one-shot reasoning-off override that
    build_api_kwargs consumes for the next request, so the continuation
    writes the answer instead of re-thinking it;
  - the ceiling exit clears a pending override and, when every fragment was
    empty, returns an actionable final_response instead of an invisible None.
2026-09-02 00:55:42 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 0ebe70d574 fix(agent): long-context tier recovery also rechecks the rebuilt request
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.

The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
2026-09-01 21:36:09 -07:00