Merge origin/main into feat/hermes-relay-shared-metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
This commit is contained in:
@@ -407,6 +407,18 @@ compression:
|
||||
# Enable automatic context compression (default: true)
|
||||
# Set to false if you prefer to manage context manually or want errors on overflow
|
||||
enabled: true
|
||||
|
||||
# Opt-in compression progress notices on chat platforms (default: false).
|
||||
# By design, routine automatic compression is SILENT on human-facing chat
|
||||
# gateways (Telegram, Discord, Slack, ...) — it happens in the background
|
||||
# with server-side logging only. Set true to also deliver the routine
|
||||
# progress statuses (compacting started, preflight/pre-API compression,
|
||||
# idle compaction, retry progress, and the compaction-complete notice) to
|
||||
# chat platforms. Unrelated operational noise (auxiliary model failures,
|
||||
# provider retry chatter) stays suppressed either way, and compression
|
||||
# FAILURE notices + manual /compress feedback are always visible
|
||||
# regardless of this setting. (#52995)
|
||||
progress_notices: false
|
||||
|
||||
# Trigger compression at this % of model's context limit (default: 0.50 = 50%)
|
||||
# Lower values = more aggressive compression, higher values = compress later
|
||||
@@ -448,6 +460,15 @@ compression:
|
||||
# compression of older turns.
|
||||
protect_last_n: 20
|
||||
|
||||
# Minimum number of REAL (actionable) user messages guaranteed to survive in
|
||||
# the uncompressed tail (default: 1 = the existing single last-user anchor,
|
||||
# behavior-preserving). Raise to e.g. 3 to keep the last 3 real user turns
|
||||
# verbatim even when bulky tool outputs fill the tail token budget — blank
|
||||
# platform echoes, compaction handoffs, and synthetic continuation rows never
|
||||
# count toward N. The tail can exceed the token budget when this pulls the
|
||||
# cut back; the guarantee wins over the budget by design.
|
||||
min_tail_user_messages: 1
|
||||
|
||||
# Compression retry rounds before a turn gives up with "max compression
|
||||
# attempts reached" (default: 3, same as the previous hardcoded value).
|
||||
# Raise (e.g. 6) for tool-schema-heavy sessions where 3 rounds cannot bring
|
||||
@@ -484,6 +505,34 @@ compression:
|
||||
# summarization on a short idle thread. Example: 1800 = compact after 30 min idle.
|
||||
idle_compact_after_seconds: 0
|
||||
|
||||
# Proactive tool-result prune (default: 0 = disabled). Opt-in token trigger
|
||||
# for a deterministic, no-LLM prune of OLD tool-result payloads, run
|
||||
# independently of `threshold` above. On large-window models (512K/1M) the
|
||||
# ratio threshold rarely fires, so bulky tool outputs (terminal dumps, file
|
||||
# reads, web extracts) ride along in history and get re-billed every turn.
|
||||
# When re-sent history exceeds this many tokens, the prune dedupes identical
|
||||
# results, summarizes older oversized ones, and truncates large tool-call
|
||||
# arguments — protecting the most recent `protect_last_n` messages and never
|
||||
# calling the model. Try 48000 to enable. Built-in compressor engine only;
|
||||
# other context engines inherit a safe no-op.
|
||||
# NOTE: a committed prune rewrites already-sent history, which invalidates
|
||||
# the provider's prompt-cache prefix — the min_reclaim gate below keeps
|
||||
# those cache breaks episodic (like a compression boundary) instead of
|
||||
# per-turn.
|
||||
proactive_prune_tokens: 0
|
||||
|
||||
# The prune's summarize pass only touches tool results larger than this many
|
||||
# characters (clamped to >= 200 so a generated summary can't be
|
||||
# re-summarized). Default 8000.
|
||||
proactive_prune_min_result_chars: 8000
|
||||
|
||||
# A proactive prune only COMMITS when it reclaims at least this many tokens
|
||||
# (measured on the pruned output). This is the prompt-cache hysteresis gate:
|
||||
# one meaningful, amortized cache break per batch of stale tool output
|
||||
# instead of a tiny break on every tool iteration. 0 = commit any non-zero
|
||||
# prune. Default 4096.
|
||||
proactive_prune_min_reclaim_tokens: 4096
|
||||
|
||||
# To pin a specific model/provider for compression summaries, use the
|
||||
# auxiliary section below (auxiliary.compression.provider / model).
|
||||
|
||||
@@ -750,6 +799,14 @@ agent:
|
||||
# window on /restart, and keep it well under systemd's TimeoutStopSec.
|
||||
# restart_drain_timeout: 0
|
||||
|
||||
# Upper bound (seconds) a submitted prompt waits for the deferred agent
|
||||
# build (MCP discovery, model metadata, skills scan) before failing with a
|
||||
# visible error. The wait is patient — the message is delivered as soon as
|
||||
# the build completes, and a progress notice is shown past 30s — so this cap
|
||||
# only fires on a genuinely hung build. Raise it for deployments with many
|
||||
# slow or unreachable MCP servers. Default 600.
|
||||
# build_wait_timeout: 600
|
||||
|
||||
# Max app-level retry attempts for API errors (connection drops, provider
|
||||
# timeouts, 5xx, etc.) before the agent surfaces the failure. Lower this
|
||||
# to 1 if you use fallback providers and want fast failover on flaky
|
||||
|
||||
Reference in New Issue
Block a user