Commit Graph

390 Commits

Author SHA1 Message Date
Teknium 2769937936 refactor(turn): lift iteration entry/announce, Nous rate guard, API interrupt, retry-restart consumer and preflight-timeout result out of run_conversation 2026-09-02 16:37:14 -07:00
Teknium a56f731ac6 refactor(turn): extract preflight gate + per-iteration transcript prep into agent/turn_preflight_gate.py, agent/turn_iteration_prep.py 2026-09-02 16:27:44 -07:00
Teknium 9780739e12 refactor(turn): extract per-iteration request assembly (api_messages/MoA/cache plan/pressure) into agent/turn_request_assembly.py 2026-09-02 16:10:07 -07:00
Teknium 9fda4e5bac refactor(turn): extract retry-loop API error handler, request build, provider call and response check into agent/turn_api_*.py + agent/turn_response_check.py 2026-09-02 16:07:51 -07:00
Teknium 0dc36e0934 refactor(turn): extract tool round + final text response branches into agent/turn_tool_round.py, agent/turn_final_response.py 2026-09-02 15:40:17 -07:00
Teknium 67ef2e50fe refactor(turn): extract response intake (normalize/hooks/scratchpad/codex-incomplete) into agent/turn_response_intake.py 2026-09-02 15:37:47 -07:00
Teknium d42282ae9e refactor(turn): extract outer-loop exception handler into agent/turn_loop_errors.py 2026-09-02 15:29:00 -07:00
Teknium ab48b1ddc5 refactor(agent): extract per-call API message build into turn_context.build_api_messages 2026-09-02 13:30:23 -07:00
Teknium 30b0759bb3 refactor(agent): extract content_filter refusal handling into turn_truncation.handle_content_policy_refusal 2026-09-02 13:30:23 -07:00
Teknium 4f53d7bd28 refactor(agent): extract post-tool-call compression decision into turn_preflight.compress_after_tool_results 2026-09-02 13:30:23 -07:00
Teknium f0d2355e26 refactor(agent): extract text-response stop gates (verify-on-stop, pre_verify hook, kanban guard) into agent/turn_stop_gates.py 2026-09-02 13:30:23 -07:00
Teknium 22cba7dd41 refactor(agent): extract classified-error routing (compaction gate, long-context tier, eager/auth fallback, Nous 429) into turn_recovery.route_classified_error 2026-09-02 13:30:23 -07:00
Teknium ef9dc88348 refactor(agent): drop write-only _bot_capability_refreshed attribute (zero readers repo-wide) 2026-09-02 13:30:22 -07:00
Teknium 7df9d3be5c refactor(agent): extract pre-API preflight compression gate into agent/turn_preflight.py 2026-09-02 13:30:21 -07:00
Teknium 0028951f35 refactor(agent): extract tool-call name/JSON validation into agent/turn_tool_validation.py 2026-09-02 13:30:21 -07:00
Teknium d398d91529 refactor(agent): extract empty/thinking-only response recovery ladder into agent/turn_empty_response.py 2026-09-02 13:30:20 -07:00
Teknium df42f37e09 refactor(agent): extract response-shape validation, invalid-response diagnostics and Codex incomplete continuation 2026-09-02 13:30:19 -07:00
Teknium 8b958e1b0f refactor(agent): unify interruptible backoff sleeps and fallback-restart arming; extract compute_error_backoff 2026-09-02 13:30:18 -07:00
Teknium 661c9af826 refactor(agent): extract finish_reason=length truncation recovery into agent/turn_truncation.py 2026-09-02 13:30:18 -07:00
Teknium 645db06053 refactor(agent): extract 413/context-overflow compression recovery into agent/turn_overflow.py 2026-09-02 13:30:17 -07:00
Teknium 5543e7ae0f refactor(agent): resume — extract API-error attempt logging into turn_recovery.log_api_error_attempt (verified partial work) 2026-09-02 13:30:17 -07:00
Teknium 25b165add5 refactor(agent): extract terminal API-failure result builders into agent/turn_recovery.py 2026-09-02 13:30:16 -07:00
Teknium 649c227e9e refactor(agent): extract per-response usage accounting into agent/turn_usage.py 2026-09-02 13:30:16 -07:00
Teknium fefc471deb refactor(agent): extract post-classification one-shot recovery chain into agent/turn_recovery.py 2026-09-02 13:30:15 -07:00
Teknium 79469656e7 refactor(agent): extract pre-classification API-error recovery into agent/turn_recovery.py 2026-09-02 13:29:47 -07:00
Teknium 4e548ce7a0 refactor(agent): compact turn-loop comments/docstrings to invariant statements (AST-identical) 2026-09-02 13:29:46 -07:00
globalvet2025 b36489be37 fix: gate Nous auth-refresh message behind verbose mode
The 'Nous agent key refreshed after 401' message used a bare print(),
making it always visible. The equivalent xAI/Codex, Copilot, and Anthropic
auth-refresh messages all use _buffer_vprint() (verbose-gated). This brings
the Nous path in line with the others so the routine ~15min OAuth key
refresh no longer prints noise on every retry.
2026-09-02 10:55:46 -07:00
Teknium 8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium 2a0605a807 fix(agent): reasoning-off continuation reaches the wire on the legacy chat path; reset one-shot flag per turn
Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
  chat_completions path always re-emitted extra_body.reasoning with
  enabled=True, so both reasoning_effort: none and the one-shot
  length-continuation override went out as {enabled: true, effort: none}.
  Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
  turn start so a flag armed by an interrupted/errored turn can never
  strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
  /thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
  reasoning.enabled=false) and a stale-flag turn-scope test.
2026-09-02 00:55:42 -07:00
AlexGabbia fb76fb0526 fix(agent): thinking-only length truncations no longer wedge continuations
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).

The length-continuation path handled that shape badly:
  1. the empty response was appended as an interim assistant fragment,
     poisoning the transcript until the pre-call sanitizer healed it
     (observed 3+ healings per turn on the reporting user's session);
  2. every continuation re-ran with thinking ON, re-deriving the whole
     thinking budget against a growing context, so 4 attempts still produced
     nothing and the turn died with 'Response remains truncated after 4
     continuation attempts'.

Now:
  - interim assistant fragments with no visible content are never appended
    (whichever way they got empty);
  - a thinking-only truncation sets a one-shot reasoning-off override that
    build_api_kwargs consumes for the next request, so the continuation
    writes the answer instead of re-thinking it;
  - the ceiling exit clears a pending override and, when every fragment was
    empty, returns an actionable final_response instead of an invisible None.
2026-09-02 00:55:42 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium 0ebe70d574 fix(agent): long-context tier recovery also rechecks the rebuilt request
The Anthropic long-context 429 handler restarts on row count alone,
the same shape #100614 fixed in the generic overflow handler. Arm the
same provider-overflow recovery flag there so the rebuilt request is
measured against the reduced window before the provider is retried.

The 413 (byte-scored) and output-cap (max_tokens) handlers are a
different yardstick and are left as-is.
2026-09-01 21:36:09 -07:00
Gille bdc46f5c09 fix(agent): recheck compressed requests after overflow 2026-09-01 21:36:09 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
overtoneblue e17276c7b4 agent: forward first stream chunk timestamp to post_api_request hook
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.

Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.

Backward compatible: existing hook subscribers ignore unknown kwargs.
2026-09-01 08:30:45 -07:00
kshitijk4poor b20cc5f787 docs(agent): explain intentional preflight vs in-loop message divergence
The preflight handler surfaces the boundary exception's per-request text
(token count, 'provider call was not sent') rather than the in-loop
_COMPRESSION_TIMEOUT_FINAL_RESPONSE constant, which describes a
different state (compression ran and could not reduce). Document the
divergence so it is not 'fixed' into a single message later.
2026-09-01 03:48:40 +05:30
kshitijk4poor 1e8f6a0491 fix(agent): surface preflight compression timeout as typed result, not generic error
When the turn-start fail-closed boundary (#98424) raises
PreflightCompressionTimedOut, the exception escaped run_conversation to
the surfaces' generic exception handlers. The gateway deliberately never
exposes raw exception text, so users saw 'Sorry, I encountered an
unexpected error... Try again or use /reset' instead of the boundary's
actionable guidance, and the compression_exhausted clean-session
recovery contract (#9893/#35809) never engaged.

Catch it at the build_turn_context callsite and convert it into the
same typed recovery dict the in-loop timeout consumers return
(salvaged #98741 / PR #99710): failed=True, partial=True,
compression_exhausted=True, turn_exit_reason=context_compression_timeout,
with the actionable message in final_response and error.

Regression test proves the exception no longer escapes and the typed
contract fields survive to the caller (mutation-checked: test fails on
main without the handler).
2026-09-01 03:48:40 +05:30
Teknium fb9b2c893f feat(agent): escalate repeated transcript-sanitiser heals with a one-time user notice (#96870)
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):

- Per-session heal counters (heal events + messages healed) tracked by the
  repair path in agent_runtime_helpers.py, session totals preserved across
  10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
  configurable via agent.sanitizer_heal_escalation_threshold in
  config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
  (events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
  the conversation loop through _emit_warning (status callback -> gateway
  status message / CLI print). Never injected into conversation context or
  the wire copy: prompt caching, role alternation, and durable history are
  untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
  the /debug share // hermes debug report, and the config key surfaced in
  hermes dump overrides. errors.log carries the ERROR line for `hermes logs
  errors`.
2026-08-31 13:11:41 -07:00
HexLab98 20fd5d0b25 fix(agent): stop empty-transcript sanitizer from warning on every send (#96870)
Fill empty non-final user/assistant turns on the wire copy during send-time projection so the sanitizer does not re-heal the same poisoned row every call. When a caller still hits the owner, log WARNING then one ERROR per session window instead of flooding errors.log.
2026-08-31 13:11:41 -07:00
Kyzcreig c1bd0511cf fix(gateway): persist the platform message id on every user turn 2026-08-31 12:46:32 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Teknium 64cc87e668 fix(compression): keep estimate seam positional-compatible for monkeypatched estimators
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
2026-08-30 20:40:43 -07:00
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Teknium 19a59e9c93 fix(compression): transiently-blocked no-op is a soft defer, never exhaustion (#97488)
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
2026-08-30 19:46:21 -07:00
james47kjv 80764b6d39 fix(codex): nudge the second continuation of a compaction-only turn
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.

But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.

Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
2026-08-30 05:16:02 -07:00
Teknium be92703788 fix(compression): keep tool-schema tokens in the unanchored fallback estimate
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
2026-08-30 05:15:46 -07:00