Commit Graph

237 Commits

Author SHA1 Message Date
dhruv kejriwal 71516214c3 fix(compressor): thread custom_providers into context-length resolution
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
2026-09-08 02:27:01 +05:30
Teknium a7198a8855 fix: keep budget checkpoints out of cancelled tool results
Skip checkpoint evaluation when a turn is interrupted so cancellation rows
remain durable without urging continued execution. The existing minimal-agent
interrupt regression also avoids dereferencing an absent iteration budget.

Consolidate the warning coverage into two invariants, including real SQLite
readback and dispatcher/child scope controls. Cold-start tool availability
between construction cases to model independent worker processes. Place ratio
normalization beside the existing iteration budget instead of growing init.

Real cancelled-tool A/B against current main, draft, and fix: three cancelled
rows and zero writes on all arms; persisted checkpoint notices 0 / 1 / 0.
Repeated scripted HTTP/SQLite loop A/B preserves completion opportunity,
ordinary default-off behavior, and blocked/two-failure exhaustion behavior.

Local targeted run initially passed 15 cases with one fixture cache-isolation
failure; corrected target and inherited affected suites remain queued behind
the campaign lock. This commit is not a CI-green or merge-ready claim.
2026-09-07 08:28:43 -07:00
Teknium 93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium 7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 2ed33fb38e refactor(memory): keep the spill, drop registration-trust rework
The prefetch spill is the fix; the is_builtin registration flag and duplicate-
instance rejection defended against a provider naming itself "builtin", which no
live path can do (only a test fixture does). Restore the name-based check and the
four test files that only changed for the new kwarg; keep two spill invariants.
2026-09-06 13:25:48 -07:00
Joey d932fa5929 fix(memory): spill oversized external prefetch 2026-09-06 13:25:48 -07:00
Teknium 0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
686f6c61 c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
Teknium d97f318717 Merge pull request #103526 from NousResearch/fix/nous-auth-stampede
fix(nous): adopt a same-account fresh key before expiry and start the keepalive in every process (620 hourly 401s → 0; review-found account takeover closed)
2026-09-06 12:02:58 -07:00
joaomarcos 0ed7acb051 perf(prefix-cache): pin session-start workspace snapshot across compaction rebuilds
- pin session-start workspace snapshot (_frozen_workspace_snapshot) on first build so dynamic git probes don't churn Tier 2 during compaction rebuilds in active coding sessions
- replay pinned snapshot across rebuilds when cwd matches; re-probe only on cwd switch
- honor coding_context invariant that workspace is a session-start snapshot, preventing prefix-cache divergence at offset ~4,662
- add invariant tests covering workspace snapshot pinning across git mutations and cwd transitions
- addresses upstream prompt divergence identified in #103326
2026-09-06 22:45:31 +05:30
Teknium 058ad0329e fix(nous): adopt a fresh agent key before the one in hand expires, and start the keepalive in every process that routes to Nous
The Nous agent key lives 3,599 s. In the 1,393-agent refactor run every
in-process agent learned about the hourly expiry from its own 401: 620
authentication_error 401s in the logged window (177 in one hour), each a
failed attempt the model never saw, and the credential pool benched the
sole credential for all of them at once. At 08:30 the storm took the
parent process down.

Two gaps. The proactive refresher (hermes_cli/nous_auth_keepalive.py) is
started only by the gateway and the web server; the CLI process, and every
subagent built inside it, never started it. And even with a fresh key in
the store, nothing adopted it before a request: a request went out with
whatever key the agent was constructed with until it 401'd.

Now _finalize_routing starts the keepalive (idempotent, process-wide,
daemon) whenever an agent resolves onto provider "nous", and
prepare_iteration calls _adopt_nous_key_before_expiry(): the agent key is
a JWT, its exp is read locally, and inside a 180 s skew the store is
re-read under the auth-store lock with force_refresh=False, so the
keepalive's (or a peer's) fresh key is adopted without a POST; when none
exists, ONE refresh runs there instead of N reactive ones after N 401s.
_try_refresh_nous_client_credentials no longer rebuilds the client when
the store returns the key already in hand.

Live A/B (local server: 401s any bearer but FRESH; store patched to hold
FRESH; 40 agents holding a JWT that expires in 30 s fire concurrently):
main 40 x 401 then recover, branch 0 x 401.

Tests (4): far from expiry the store is not touched; inside the skew the
store's fresh key is adopted with force_refresh=False; the same key back
from the store is not re-adopted; a real AIAgent routed to nous starts
the keepalive and one routed to openrouter does not.
2026-09-05 01:27:23 -07:00
Teknium d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium 2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium 7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00
Teknium c4b8485877 review-fix(suppress-audit): agent_init/browser_supervisor/skills_sync_client/write_approval — restore BASE exception semantics 2026-09-03 09:48:44 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 3a8e3a2e88 review-fix(moa): restore _relay_moa_reference_event/_moa_reference_output_allowed + quiet-reference test
BASE (63279301bc) exported agent.agent_init._moa_reference_output_allowed and
_relay_moa_reference_event (salvaged in 3dfe712384 from #67334, plugin-importable,
zero in-tree callers) and covered them with
tests/agent/test_moa_quiet_reference_output.py. The simplification PR deleted both
the helpers and the test. Restore them byte-identical to BASE.

Note: the live build_moa_facade relay in agent/moa_loop.py is unchanged. A/B
harness (/tmp/rf/rev/moa_quiet_ab.py) shows BASE's facade relay already fired
moa.reference under platform=cli/tool_progress_mode=off — the guard only ever
lived in the orphan helper. -Q is protected on both trees by cli.py nulling
agent.tool_progress_callback (_configure_quiet_agent / BASE cli.py:22308).
2026-09-03 09:35:34 -07:00
Teknium 3635f9aa4a refactor(agent/agent_init): reuse agent_runtime_helpers._ra instead of a duplicate lazy run_agent shim 2026-09-02 22:22:03 -07:00
Teknium 17412822d9 refactor(agent): try/except/pass -> contextlib.suppress (28 sites, AST round-trip verified) 2026-09-02 22:00:23 -07:00
Teknium 395327c16f refactor(agent/agent_init): fold fallback-chain banner; drop redundant aux-context pre-init 2026-09-02 20:23:31 -07:00
Teknium 3728e76093 refactor(agent/agent_init): compact compression summary + context-engine tool injection 2026-09-02 20:20:49 -07:00
Teknium 336fd952d6 refactor(agent/agent_init): fold custom-provider model-match + key-filter ladders 2026-09-02 20:17:33 -07:00
Teknium 2c85a124f2 refactor(agent/agent_init): fold route normalizer + memory-provider warn branch 2026-09-02 20:16:07 -07:00
Teknium 963b0e2eef refactor(agent/agent_init): tighten route-url, run-budget and extra_body helpers 2026-09-02 20:13:57 -07:00
Teknium dc5bd74cfb refactor(agent/agent_init): try/except-pass → contextlib.suppress (26 sites) 2026-09-02 20:05:54 -07:00
Teknium 45e204a48e refactor(agent/agent_init): extract _apply_openai_header_policy; reuse identity table; fold display-config ladders 2026-09-02 20:00:03 -07:00
Teknium 9065fb15fa refactor(agent/agent_init): pack init_agent signature (AST-neutral) 2026-09-02 19:53:20 -07:00
Teknium 1defa76992 refactor(agent/agent_init): flatten tool-loading banner branches; single model section lookup 2026-09-02 19:49:23 -07:00
Teknium 574d04e3cc refactor(agent/agent_init): CompressionSettings as SimpleNamespace; bool-gate loop; drop dead _emit_warning hasattr; tighten lazy-import blanks 2026-09-02 19:44:33 -07:00
Teknium 63f77f9f34 refactor(agent/agent_init): compact function docstrings (AST-neutral) 2026-09-02 19:41:35 -07:00
Teknium a7b48544da refactor(agent/agent_init): session + usage state as default tables 2026-09-02 19:38:50 -07:00
Teknium f7e89504dd refactor(agent/agent_init): control/turn/stream state as declarative default tables (_set_defaults) 2026-09-02 19:30:33 -07:00
Teknium 51cec303ad refactor(agent/agent_init): collapse getattr on always-present provider; simplify anthropic key fallback 2026-09-02 19:21:40 -07:00
Teknium 8a0be6af2a refactor(agent/agent_init): compact inline comments to WHY-only (AST-neutral) 2026-09-02 19:17:01 -07:00
Teknium a5e537cfa9 refactor(agent/agent_init): wip checkpoint — phase-helper split of init, param tables, header factory table, compression parse helpers; mixin compaction 2026-09-02 19:07:13 -07:00
Teknium 76062b1d6b refactor(agent): decompose init_agent into ordered phase helpers
init_agent (2711 LOC) becomes a ~280-line ordered orchestrator over
_resolve_api_mode / _finalize_routing / _init_* / _build_client /
_load_tools / _parse_compression_config -> CompressionSettings /
_resolve_context_length / _build_context_engine / ... phase helpers.
_build_client is further split per wire mode (_init_anthropic_client,
_init_moa_client, _init_bedrock_client, _init_openai_client with
_explicit_client_kwargs / _routed_client_kwargs). Statement order and
every side effect on the agent are preserved (AST body-parity checked
against origin/main).

Dedupe/dead code: drop _relay_moa_reference_event/_moa_reference_output_allowed
(zero callers; only their own test) and their test file; alias
_normalize_route_base_url; _parse_config_int replaces three copies of the
strict int parser; _cfg_flag replaces four inline truthy-set checks;
_client_kwargs_from_routed + _fallback_entries replace duplicated
routed-client/fallback-entry blocks; _warn_invalid_config_int unifies the
three log+stderr invalid-int warnings (byte-identical text);
_bedrock_region_from_url; _memory_provider_init_kwargs; the
host->default_headers if/elif chain becomes the _HOST_DEFAULT_HEADERS
dispatch table; callback params assigned from _CALLBACK_PARAMS.

Comments/docstrings hand-compacted to their rationale (invariants,
ordering, failure modes kept; issue numbers and narrative dropped).
test_pre_compress_checkpoint_contract source-check repointed at the
CompressionSettings field names.

Verified: tests/run_agent (2066 passed) + all agent_init-referencing tests
(1134 passed), get_tool_definitions() byte-identical vs origin/main, import
smokes for cli/run_agent/gateway.run/hermes_cli.main/agent.conversation_loop/
tui_gateway.server.
2026-09-02 13:29:34 -07:00
muhifni 1cd736ff63 fix(terminal): scope terminal config per turn under profile multiplexing
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.

Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:

- tools/terminal_scope.py: ContextVar holding the routed profile's
  COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
  TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
  resolves ONLY from it - an omitted key yields the defined default,
  never os.environ. Unreadable/malformed policy installs a refusal
  scope; terminal_tool / execute_code refuse instead of running under
  ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
  _profile_runtime_scope, tui_gateway session/build/turn scopes, cron
  per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
  (_get_env_config, _resolve_container_task_id shared key, orphan
  reaper lifetime, degraded mode), gateway/platforms/base.py docker
  media translation (volumes, shared key, persistence), runtime_cwd /
  agent_init / skill_utils / code_execution_tool / file_tools cwd
  anchors, prompt_builder / browser_tool / env_probe backend checks,
  gateway footer, @-refs and slash-command cwd. env_probe resolves the
  backend in the caller's context, since the probe worker thread does
  not inherit the ContextVar.

Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.

Fixes #68559
Fixes #94200
Fixes #101132
Fixes #95470

Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
2026-09-02 05:34:28 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
João Vitor Cunha ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
David Metcalfe c905c2b4b5 fix(agent): honor model.streaming: false as a non-streaming escape hatch (#72901)
The conversation loop has forced stream=True for every turn — subagents
included — since #3120 (always-prefer-streaming for liveness health
checking). Self-hosted OpenAI-compatible backends with broken streaming
tool-call paths (e.g. vLLM --tool-call-parser qwen3_xml + reasoning
parser + MTP) can leak tool-call markup into plain text and return zero
tool_calls, so delegated tasks silently no-op instead of executing.

model.streaming was never a real config key, so users could not opt out.
Seed agent._disable_streaming from model.streaming: false at init; the
loop already routes that flag to the non-streaming path (the same path
used when a provider rejects streaming at runtime). Default stays
streaming-on, preserving #3120's behavior for everyone else. Orthogonal
to display.streaming (token rendering).

Tests: config->flag seeding (patched loader + real config.yaml E2E),
legacy string model section, multi-agent config propagation.
2026-09-01 22:14:06 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
VJ Pixel 3571118218 fix(agent): run init-time fallback for ANY exhausted primary provider
The #17929 init-time fallback block was nested inside the
'_explicit not in {auto, openrouter, custom}' guard, so an exhausted
openrouter credential pool skipped fallback_providers entirely and
AIAgent.__init__ raised 'No LLM provider configured' — surfacing on
Telegram only as the generic 'unexpected error' message.

2026-08-23 outage (~22:09-23:50, 60 occurrences incl. a cron job):
single-entry openrouter pool hit daily free-tier quota; chain had a
healthy local Ollama entry that never got tried because provider was
the default 'openrouter'.

Un-nest the block so any primary without usable credentials walks the
chain before failing; providers explicitly chosen by name keep their
dedicated missing-key diagnostic when both primary AND chain fail.
Regression tests cover the openrouter-exhausted path both with and
without a usable chain entry.
2026-09-01 02:00:05 -07:00
Teknium 7cefa87ea7 fix(agent_init): reserve Gemini's default maxOutputTokens in the compressor when max_tokens is unset
The native generateContent adapter never runs uncapped: when
model.max_tokens is unset it sends maxOutputTokens=65,535
(GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) because Gemini treats an omitted cap
as a low internal default. The context compressor's trigger is
pct×(window − max_tokens), and constructing it with max_tokens=None
reserved 0 — so on a 128K Gemma window the trigger landed at 98,304
while the real safe input budget was 65,537, and the provider 400'd
before compaction fired.

Live repro (real imports, temp HERMES_HOME, native Gemini base_url,
window=131072, max_tokens unset):
  before: compressor.max_tokens=None, threshold_tokens=98304,
          wire maxOutputTokens=65535 → trigger ABOVE the safe budget
  after:  compressor.max_tokens=65535, threshold_tokens=64000 → below it

Scoped to the native Gemini wiring (provider names + native base_url via
is_native_gemini_base_url; the /openai compat endpoint is excluded). The
generic provider-default reservation gap remains tracked in #63839.

Reported by @Artemonim in #57275 (residual claim 4).
2026-08-31 12:22:55 -07:00
Teknium cb71d5f1b1 fix(agent_init): clamp compressor window to Ollama num_ctx resolved after construction
model.ollama_num_ctx is resolved AFTER the context compressor is
constructed, so a config that sets only ollama_num_ctx (without
model.context_length) ran every request at the smaller served num_ctx
while the compressor still targeted the probed GGUF window (e.g. 256K
Gemma metadata). The compaction trigger then sat several times above the
window the server actually serves and never fired — reproducing the
original #57275 'blows past the limit' symptom on current main.

Live repro (real imports, temp HERMES_HOME, config = {model:
{ollama_num_ctx: 65536}}, probed window 262144):
  before: _ollama_num_ctx=65536, compressor.context_length=262144,
          threshold_tokens=196608 (300% of the served window)
  after:  compressor.context_length=65536, threshold below the window

The clamp is one-directional (a num_ctx larger than the resolved window
never inflates the compressor) and reuses update_model() so every
threshold-derived budget recalibrates. Overlaps #60103 (silent-clamp
dead zone) — this is the init-order half.

Reported by @Artemonim in #57275 (residual claim 3).
2026-08-31 12:22:17 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00