Commit Graph

2003 Commits

Author SHA1 Message Date
Teknium 4032a15ad0 refactor(prompt): remove the ~1.2K-token Nous Subscription block from the system prompt (#95005)
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-08-25 12:59:29 -07:00
Teknium 1ee524f77d fix(memory): bind the checkpoint gate to post-turn micro-compaction too
Independent review caught a compaction authority the gate missed:
post-turn micro-compaction (turn_finalizer -> _micro_compact) absorbs the
oldest exchanges into a rolling summary with no pre-compress checkpoint
hook in its path, and both compression.checkpoint_required and
compression.micro_compact could be enabled together — assistant evidence
could vanish into a summary the checkpoint filter later excludes, without
ever reaching the durable provider.

- agent_init: checkpoint_required forces micro-compaction off (warned),
  mirroring the native-compaction suppression
- turn_finalizer: defense-in-depth guard at the call site (attribute is
  plain mutable state a future path could flip on a live agent)
- behavioral regression test with a sabotage control (gate off proves the
  harness reaches the call site; gate armed proves zero calls)
- docs + config example mention the suppression; stale v1 test header fixed
2026-08-25 03:55:55 -07:00
Teknium 9e551d2931 refactor(memory): renumber checkpoint API — v1 is the implicit historical contract, v2 opts into fail-closed checkpoints
Per review: existing providers should not be retroactively re-versioned or
handed a changed payload. Version 1 is now the implicit historical
on_pre_compress() contract (best-effort, raw message list) that every
pre-existing provider is already on; the fail-closed checkpoint contract
becomes version 2. MemoryManager routes the raw transcript to v1 providers
unchanged and hands the host-normalized evidence list only to v2+ checkpoint
providers, so the plugin surface contract for shipped providers is
byte-identical with the gate off.
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 8cc379b528 fix(memory): bind the checkpoint gate to every compaction authority
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):

- codex app-server: in "native"/"off" auto-compaction mode (native is the
  default) Hermes preflight is skipped and the codex agent compacts its
  own thread inside run_turn() — the compress_context() rejection was
  unreachable. init_agent now refuses checkpoint_required together with
  api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
  testable guard), and run_codex_app_server_turn() fails closed as
  defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
  native_compaction_context_management() now returns None while the gate
  is armed, so context_management never goes on the wire and the
  checkpoint-aware Hermes compressor stays authoritative. The suppression
  is logged once per process, not silently applied.

Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 70d0b1fffb fix(memory): review follow-ups for the pre-compress checkpoint contract
Addresses the review on #93996:

- gateway: hygiene and manual /compress load the memory provider only when
  compression.checkpoint_required is enabled (skip_memory=not required).
  The historical fast path — no provider init, no best-effort hook — is
  back for everyone who did not opt in, so default behavior is truly
  unchanged.
- conversation_compression: assistant messages carrying both prose and
  tool_calls keep their prose in the checkpoint evidence (the tool_calls
  payload is stripped, the original message is not mutated); pure
  tool-call wrappers without prose are still dropped.
- tests: legacy-database regression proving the _compressed_summary column
  is added by the declarative _reconcile_columns() path on a plain reopen
  (no version-gated migration needed — append_message works right after),
  plus coverage for the prose-preserving filter.
- docs: providers must implement idempotent, content-keyed checkpoint
  writes — a fail-closed block means the next attempt re-runs
  on_pre_compress over largely the same transcript.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 1104ffe0b9 feat(memory): opt-in fail-closed pre-compress checkpoint contract (API v1)
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.

This adds an opt-in, provider-agnostic checkpoint contract:

- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
  by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
  historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
  on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
  failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
  (default false, documented in cli-config.yaml.example). When enabled,
  compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
  uncompressed transcript is preserved) unless a checkpoint-capable provider
  confirms the durable checkpoint. Providers receive normalized direct
  user/assistant evidence: tool rows, system messages, tool-call wrappers,
  and prior compaction summaries are filtered host-side into one stable
  contract. codex_app_server compaction is rejected under the gate because
  it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
  migration via _reconcile_columns) so summary provenance survives process
  restarts; only the resume model history carries the marker, keeping
  get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
  (skip_memory=False) so a required checkpoint also guards those rewrites.

The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.

Refs #93986
2026-08-25 03:55:55 -07:00
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Guilherme Artiles 8ad20d065a test(curator): assert the instruction through the delivered prompt, not the source
The first version of this test used inspect.getsource() on skill_manager_tool
and regex-parsed the guarded action literals. AGENTS.md bans source-text tests,
and the ban is right here: that test would pass against a guard wired to the
wrong call site and fail on a pure rename, neither of which is the thing worth
guarding.

Replaced with a behavioral assertion in the shape of the neighbouring
dry-run-banner test: stub _run_llm_review, run run_curator_review, and assert
the prompt that actually reached the model names skill_view and all four
guarded actions (edit, patch, write_file, remove_file).

It still discriminates: with the prompt block removed, action=edit and
action=remove_file no longer appear anywhere in the assembled prompt (the
toolset list only mentions patch, create, write_file and delete), so the test
fails. Runtime guard behavior stays where it belongs, in
tests/tools/test_skill_manager_tool.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:31:14 -07:00
Guilherme Artiles 610e2e02fb fix(curator): tell the background reviewer to read before it writes
`_background_review_read_before_write_guard` refuses a background-review
`skill_manage` write when the target file was not loaded via `skill_view` in
the same review turn (patch, edit, write_file over an existing file,
remove_file).

`CURATOR_REVIEW_PROMPT` never says so. It lists `skill_view` only under "read
the current landscape", so the reviewer goes straight to the write and every
mutation is refused. The failure is silent from the outside: the curator run
completes, writes nothing, and reads like a pass that simply found nothing to
consolidate. On our deployment that was 32 of 32 attempted writes rejected over
48h before anyone read the logs.

This adds the missing instruction to the toolset block, plus a test that fails
if a future guarded action is added to `skill_manager_tool` without being named
in the prompt — the guard and the prompt have to drift together or not at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:31:14 -07:00
Nathan Shan d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
Chen Jin b45b028573 fix(agent): gate memory provider system_prompt_block on toolset config (#81014)
The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provider_tools_enabled()` via platform_toolsets or
disabled_toolsets. Result: the agent received instructions to call
`mnemosyne_remember`, `mnemosyne_recall`, etc., that did not exist in
its tool surface.

Centralize the gating into `memory_provider_tools_exposed(agent)`, use
it from both `inject_memory_provider_tools` and the system prompt
assembly path, and add regression tests covering:
* memory toolset enabled -> both tools and prompt block exposed,
* memory in disabled_toolsets -> neither exposed,
* memory not in enabled_toolsets and not built-in -> neither exposed,
* the built-in "memory" tool present as an opt-in -> both exposed,
* parity between `inject_memory_provider_tools` and
  `memory_provider_tools_exposed`.
2026-08-24 21:45:30 -07:00
Teknium 3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
fangliquanflq 1420176393 fix(lsp): abort diagnostics waits after transport death 2026-08-25 01:48:43 +05:30
fangliquanflq 2f506c2023 fix(lsp): retire clients when the protocol reader exits 2026-08-25 01:48:43 +05:30
kshitijk4poor 547f985286 refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
Per-site decisions:

1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
   MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
   via a function-local import; keeps the swallow-everything fail-open
   contract and the (proc) -> None signature (agent/shell_hooks.py imports
   it by name; _kill_git_process_tree alias preserved). The old body is kept
   verbatim as _legacy_kill_process_tree and used as fallback when the
   delegation import/call fails. A final proc.kill() is retained on the
   happy path so Popen bookkeeping sees the exit (matches old behavior).

2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
   (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
   body sent SIGTERM then SIGKILL with zero grace between them; the shared
   primitive sends SIGKILL only. With no grace period the observable effect
   is identical, and the psutil descendant sweep now also reaches
   agent-browser's setsid'd daemon grandchild, which killpg alone missed.
   tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
   the legacy fallback (its assertions describe the fallback's internals).

3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
   MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
   kill when escalate=True) — expressed as two delegated calls:
   kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
   kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
   proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
   body terminated children before the parent; the shared primitive
   signals the group atomically (child is a session leader via
   start_new_session=True) plus an identity-aware descendant sweep —
   strictly wider coverage, same signals.

4. gateway/status.py: KEPT BOTH SITES.
   - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
     incompatible with the shared primitive — it must RAISE OSError with
     taskkill's stderr on non-zero exit (callers branch on that), falls back
     to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
     single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
     fail-soft primitive would invert the error contract.
   - reap_gateway_children (~l2029): NOT migrated. It operates on a
     pre-snapshotted child list from a parent that is already dead
     (psutil.Process(pid) on the parent would fail), and every signal is
     wrapped in identity/ownership checks the primitive lacks: is_running()
     identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
     plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
     The coupling is the feature; migrating would delete the safety logic.

5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
   Dev tooling that intentionally kills by CAPTURED pgid because the direct
   child is usually already reaped (psutil/pid-based primitive cannot find
   it), and it avoids the psutil import on the test-runner hot path. Its
   docstring already documents why psutil is the wrong tool there.

New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
2026-08-25 01:34:56 +05:30
Vignesh Ramesh a0795acc83 fix(codex): identify Hermes requests 2026-08-24 11:25:04 -07:00
kshitij 6e534df114 Merge pull request #93773 from kshitijk4poor/fix/codex-sdk-transform-bypass-93650-v2
fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
2026-08-24 16:11:18 +05:30
Teknium ec5e369fe6 Merge pull request #93784 from NousResearch/salv/81234-retry-carrier
fix: /retry and /undo no longer replay an older message after compaction (#81233, salvage #81234)
2026-08-24 03:31:29 -07:00
Teknium c9e2a46df6 fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469)
The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
2026-08-24 03:23:07 -07:00
Teknium 4b622bbc4b test(model_metadata): lock in max_tokens last-resort fallback + cache self-heal
Adjust the #93423 max_tokens-only regression test to the merged policy:
max_tokens stays as an explicit LAST-RESORT fallback (some local servers
report nothing else) instead of being dropped entirely, and add coverage
that _reconcile_local_cached_context_length rewrites a cache entry
poisoned by the old probe (393216) upward to the real window (1048576)
once the probe is fixed.

Co-authored-by: pju-hoge <grkt@ppmz.com>
Co-authored-by: re-ITRT <1940428933@qq.com>
2026-08-24 03:20:57 -07:00
re-ITRT a0c802c02c fix(model_metadata): stop misreading max_tokens as context length in local probe
The local-endpoint context probe (_query_local_context_length_uncached)
treated max_tokens — an output-completion cap — as a candidate for the
model's context window. For OpenAI-compatible gateways that advertise a
1M context via context_size / max_input_tokens alongside a smaller
max_tokens output cap (e.g. TokenHub serving deepseek-v4-flash:
context_size=1048576, max_input_tokens=1048576, max_tokens=393216),
Hermes mis-detected the window as 393,216 and — because loopback
endpoints are reconciled against a live probe — actively overwrote a
previously-correct 1M cache entry.

- Add context_size and max_input_tokens to both /v1/models probe
  candidate lists (single-model detail and list branches).
- Remove max_tokens from the context-length candidates; it remains
  handled separately as an output cap (_MAX_COMPLETION_KEYS).

Adds regression tests covering context_size/max_input_tokens priority
over max_tokens and the max_tokens-only (no real context key) case.
2026-08-24 03:20:57 -07:00
Kyzcreig 4d729e4b31 fix(model-metadata): local ctx probe must not read max_tokens as the context window
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.

Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
2026-08-24 03:20:57 -07:00
carlotestor 394f0f0902 fix(context): prefer max_input_tokens over max_tokens for Anthropic proxies
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:

  max_input_tokens = context window (e.g. 1M for claude-fable-5)
  max_tokens       = max output     (e.g. 128k)

That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).

Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
2026-08-24 03:20:57 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
kchernev 10a070bd49 fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:29:45 +05:30
Teknium 0c3a507535 fix(classifier): 429 quota walls route to billing across providers; reset signals stay rate-limited
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:

- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
  the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
  exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
  credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
  exclusion so an explicit 'Rate limit exceeded' never promotes to
  non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
  reached. Resets in 6hr 29min.' wrongly read as terminal billing because
  main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
  signals. (credit @jtstothard, #74785)

Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.

Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.

Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
2026-08-23 20:02:07 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
Adolanium 2912c36aa4 fix(gateway): stop multiplex allowlist leak and bot-relay python -c injection
_auth_env fell through to os.environ on a scoped miss, so one profile
could inherit another profile's allowlists and allow-all flags.

bot_relay.waiter_command put connection_id into python -c source. A
quote in the id broke the waiter. A crafted id could run extra Python
in the sender gateway.
2026-08-23 20:00:30 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 23fb949f2c feat: /review briefing embeds the workspace's project context files
load_workspace_context() resolves the parent's workspace via the same
_resolve_workspace_hint used for child prompts (explicit sources only —
TERMINAL_CWD / agent cwd hints, never a bare getcwd fallback, so the
#64590 install-tree-leak guard concern doesn't apply) and runs it
through agent.prompt_builder.build_context_files_prompt — the exact
discovery/priority/cap logic the main system prompt uses (.hermes.md >
AGENTS.md chain > CLAUDE.md > .cursorrules; SOUL.md skipped). The
result is embedded in the reviewer briefing as binding review
standards. Subagents are built with skip_context_files=True, so without
this the reviewer judged repo work without the repo's own conventions.

5 new tests incl. real-filesystem AGENTS.md discovery through the real
loader. Docs updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 22381edc11 feat: /review briefing carries the parent's loaded skills
The reviewer subagent now inherits the primary agent's working skill
context: collect_parent_loaded_skills() gathers launch-preloaded skills
(from the activation notes in ephemeral_system_prompt) and mid-session
skill_view loads (from assistant tool_calls in history), deduped and
capped at 8, and the briefing instructs the reviewer to skill_view each
and treat their conventions as binding for the assessment.

Reference-file reads (file_path=...) don't count as loads; full-skill
injection was rejected as too costly (a single dev skill can be 40KB+).

Docs: delegation.md /review flow updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 580060ffd8 fix: reuse first-observed sequence when announced items land via output_item.done
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.

Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
2026-08-23 19:02:30 -07:00
cxxCoolStar 4f3ae189a3 fix(agent): harden pending Responses tool call settlement 2026-08-23 19:02:30 -07:00
cxxCoolStar 720344cfba fix(codex): settle pending Responses tool calls when output_item.done is omitted
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
2026-08-23 19:02:30 -07:00
fangliquanflq 654d537088 fix(agent): honor structured quota reset signals 2026-08-23 18:43:12 -07:00
fangliquanflq c2090ba6b4 fix(desktop): distinguish provider quota exhaustion 2026-08-23 18:43:12 -07:00
Teknium 04dd2bb233 test(agent): drain truncation warnings before and after each prompt-builder test
Follow-up to the ContextVar-leak fix: the autouse fixture now drains on
both sides (drain(); yield; drain()) so earlier files can't pollute this
file's assertions either.
2026-08-23 18:27:12 -07:00
Aniruddha Adak f168d857c3 test(agent): stop truncation-warning ContextVar leaking between test files
Running `pytest tests/agent/test_prompt_builder.py
tests/agent/test_system_prompt.py` failed
test_build_system_prompt_records_stable_prefix with AttributeError:
'...SimpleNamespace' object has no attribute '_emit_status'
(#93018). A truncation warning recorded by test_prompt_builder.py stays
in the shared thread context under plain pytest, so the later file's
build_system_prompt call drains a warning and forwards it to
agent._emit_status - which the test stub lacked.

Harden both sides:

- tests/agent/test_system_prompt.py: _make_agent() stub gains a no-op
  _emit_status, so draining a stray warning is harmless.
- tests/agent/test_prompt_builder.py: autouse fixture drains pending
  truncation warnings after every test, leaving the ContextVar clean.

The order-dependent failure no longer reproduces in either ordering.
2026-08-23 18:27:12 -07:00
aniruddhaadak80 cd6c088928 test(compression): align no-op strike tests with structural backoff (#93022)
Two suites still encoded the pre-#93093 contract that the three
structural no-op branches (insufficient_messages, no_compressible_window,
empty_post_handoff_window) increment _ineffective_compression_count:

- tests/agent/test_compaction_anti_thrash.py::
  TestMinimumMessagesBranch::test_too_few_messages_records_an_ineffective_pass
- tests/run_agent/test_infinite_compaction_loop.py::
  TestCompressNoOpRegistersIneffective::{test_no_op_increments_counter,
  test_two_no_ops_block_should_compress}

Structural no-ops are transcript-shape facts, not evidence of an
incompressible floor, so they now arm _structural_no_op_backoff_until
and leave the strike counter untouched. Update the tests to pin the new
contract (count unchanged, backoff armed via time.monotonic(),
should_compress blocked while it holds) and rename accordingly. The
outcome contract of test_two_no_ops_block_should_compress is preserved:
repeated no-ops still block further automatic compression.
2026-08-23 18:27:07 -07:00
Aniruddha Adak f778c0d941 fix(compression): structural no-ops defer retries instead of striking the breaker
Fixes #93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).

Distinguish "nothing eligible right now" from "attempted and
underperformed":

- New transient _structural_no_op_backoff_until (in-memory, 300s)
  armed by _record_structural_no_op() at the three structural no-op
  sites: insufficient_messages, no_compressible_window,
  empty_post_handoff_window. No strikes accumulate; auto-compaction
  resumes on its own once the backoff lapses or the transcript outgrows
  the protection window.
- The backoff gates should_compress via
  _automatic_compression_blocked_locally and surfaces in
  _compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
  never shrink retries at most once per backoff window instead of
  every turn.
- force=True (/compress) clears an active backoff before attempting;
  record_completed_compaction() lifts it - both prove the transcript
  is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
  durable ineffective counter unchanged.

Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
2026-08-23 18:27:07 -07:00
Aintworth ce51f535d3 test(gemini): cover nested/list/non-pointer ref cases; document false-positive tolerance
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
  (already wrapped, no 400 path), and $ref without '#/' prefix (stays
  structured).
- Document the deliberate structural (false-positive-tolerant) detection and
  its O(n) cost in the helper docstring.
2026-08-23 18:27:04 -07:00
Aintworth 03477166f9 fix(gemini): wrap schema-bearing tool results as opaque text
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).

tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.

Adds regression tests for the wrap path and the unchanged structured path.
2026-08-23 18:27:04 -07:00
Finn763 74e6885f0d fix(review): fail-closed compressor detachment + warm-cache first request (#93057 review)
Adversarial-review fixes for the #93057 snapshot-compaction PR:

- Fail-closed detachment: only re-enable compression after
  bind_session_state successfully severs the engine's parent binding.
  A failed rebind keeps the historical compression_enabled=False
  behavior and warns, instead of running compaction against a
  compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
  preflight + pre-API pressure check) until the fork's first provider
  response, so the first request replays the full snapshot as the
  intended cached read and compaction applies from the second request
  on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
  pre-fix code) and the existing threshold-crossing test reworked to a
  two-request review asserting the warm first request + compacted
  second request. 116 tests green across all touched suites; ruff
  clean.
2026-08-23 18:25:19 -07:00
Finn763 4202a508fd fix(review): bound same-model background review replay
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.

Closes #93057
2026-08-23 18:25:19 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium 65c58651b0 feat: review slot appears in every aux-model picker (desktop, dashboard, CLI)
Follow-up to #93339: the auxiliary.review slot existed in config but was
missing from every model-picker surface, so users could only set the
review model by hand-editing config.yaml.

- hermes_cli/web_server.py: review in _AUX_TASK_SLOTS (REST allowlist,
  stale-aux warning sweep)
- hermes_cli/main.py: review in _AUX_TASKS (hermes model aux picker)
- apps/desktop model-settings.tsx + all 5 i18n locales (en/ja/zh/
  zh-hant/ar): review slot with label/hint
- web/src/pages/ModelsPage.tsx: review row in dashboard Models page
- tests: registry-sync test pinning review across DEFAULT_CONFIG,
  _AUX_TASKS, and _AUX_TASK_SLOTS (curator pattern)
- docs: aux-task table in fallback-providers.md (en) + zh-Hans mirrors
  of fallback-providers and the delegation /review section missed in
  #93339
2026-08-23 18:22:39 -07:00
liuhao1024 51239e8e2a fix(vision): forward the API key to the server-type probe and cache failed verdicts
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).

Two changes:

- image_routing._should_probe_ollama_vision now takes the API key and
  forwards it; a new _resolve_inference_api_key mirrors
  _resolve_inference_base_url's resolution order (runtime value,
  model.api_key, providers blocks) so the key always matches the URL
  being probed.

- detect_local_server_type caches a None verdict in memory with a short
  failure TTL (5 min, vs 1h for positives) so the next turn is served
  from the negative entry instead of re-running the waterfall — while
  a transient failure (server starting, key being fixed) recovers in
  minutes. Negative verdicts are deliberately not written to the
  cross-process disk cache.
2026-08-23 17:47:50 -07:00