Commit Graph

2233 Commits

Author SHA1 Message Date
Alex Fournier 31402f630b fix(relay): complete native plugin cutover
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 93bec27f66 fix(relay): preserve shutdown scope cleanup
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 6ec2c0ba8b fix(relay): define process-wide profile policy
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski 3fad83df31 fix(relay): guard native plugin ownership cutover
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:03 -07:00
Bryan Bednarski ad9fb060c5 fix(relay): clarify opt-in plugin layering
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski c86fd74e28 refactor(relay): require 0.7 plugin APIs
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 4e34dfc932 refactor(relay): use canonical dynamic plugin config
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski 0b7288ebcb fix(relay): require explicit plugin configuration
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski e7da915f67 fix(relay): defer subscriber flush to shutdown
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski 918dd8a265 feat(relay): load standard dynamic plugin records
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski 88300217c2 feat(relay): activate configured dynamic plugins
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
Bryan Bednarski c4ae7f7a3b feat(relay): initialize discovered plugin components
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
ethernet 3c675019f1 fix(aux): retry once without response_format when a provider rejects it
Some providers reject the structured-output request field with a hard
400. The error classifier marks a 400 as non-retryable, so one rejected
field failed the whole auxiliary call. Session titles stayed derived
forever (#82816), and no fallback fired.

Three rejection shapes are covered, from live reports:
- vLLM gateways translate response_format into guided_grammar and fail
  when the grammar backend is absent (compile_grammar_error: No module
  named 'xgrammar').
- Some OpenAI-compatible endpoints answer "This response_format type
  is unavailable now".
- Anthropic-compatible gateways that predate structured outputs reject
  the translated field: "output_config: Extra inputs are not
  permitted". The documented case is the bedrock-mantle Messages
  endpoint.

The fix is reactive, the same pattern as the temperature and
max_tokens rungs: when the provider rejects the field, retry once
without it. Callers tolerate an unconstrained reply — the title prompt
demands bare JSON and _extract_title_text has a loose-JSON fallback —
so the call succeeds with prompt compliance instead of failing. The
retry only fires when the request carried the field, and both the sync
and async paths get the same rung.

Closes #82816
2026-08-18 20:34:46 -04:00
ethernet 8f2d61e3de fix(aux): translate top-level response_format kwarg on the Anthropic adapter
The adapter builds the Messages body from a fixed allow-list of kwargs.
A caller that passes response_format as a top-level kwarg (the OpenAI
SDK call shape) got it dropped on the floor. The request succeeded, but
the schema contract silently became prompt compliance. No in-tree
caller uses this shape today. The pin-test makes sure that a future
refactor cannot open this leak again.

The top-level kwarg gets the same output_config.format translation as
the extra_body shape. When a caller sends both shapes, the extra_body
value wins because every in-tree caller uses that shape.

Pin-test pattern from PR #85626 review follow-up.

Co-authored-by: Matt McClean <mmcclean@amazon.com>
2026-08-18 20:34:46 -04:00
Soju06 f709bd8445 fix(aux): translate response_format to output_config.format for anthropic transport
Plugin structured completions (plugin_llm.complete_structured) build an
OpenAI Chat Completions response_format payload in extra_body. The
anthropic_messages transport forwarded it verbatim, and strict
Anthropic-compatible gateways reject it with HTTP 400:

  response_format: OpenAI Chat Completions structured-output shape is
  not supported. Use output_config.format = {"type": "json_schema", ...}

Observed live: every discord-thread-autotitle structured call failed
for 2+ days (1,600+ logged errors) once the main provider became an
anthropic_messages gateway.

Fix: _translate_anthropic_response_format converts
- json_schema  -> output_config.format = {type: json_schema, schema: S}
- json_object  -> permissive object schema (SDK 0.87.0 has no
  schema-less JSON mode)

merging into any existing output_config (adaptive-thinking effort
coexists) and excluding response_format from the raw extra_body
passthrough alongside the existing reasoning exclusion. The async
adapter delegates to the sync adapter via asyncio.to_thread and is
covered by a test. Non-Anthropic transports are unchanged.
2026-08-18 20:34:46 -04:00
Jeffrey Quesnelle 9664e386f6 Merge pull request #85581 from bbednarski9/codex/fix-openai-sparse-response-objects
fix(openai): tolerate sparse response objects
2026-08-18 11:42:24 -04:00
Jeffrey Quesnelle aa09c8fe8c Merge pull request #85580 from bbednarski9/codex/fix-relay-client-timeout-payload
fix(relay): keep client timeout off managed payloads
2026-08-18 11:33:51 -04:00
Teknium 7a0cdbcd79 fix(tests): pin workspace snapshot in plugin-prompt-sections byte-stability test
test_real_aiagent_builds_section_once_and_keeps_it_out_of_static_prefix
builds the system prompt twice and asserts byte equality, but the prompt
embeds build_coding_workspace_block() — live git status/log output. A git
call failing between the two builds (xdist contention in CI) makes the
Branch/Recent-commits lines differ and fails the test on unrelated PRs
(first seen on #89027's run: diff showed only '- Branch: (detached HEAD)'
and recent-commit lines).

Mechanism reproduced locally: with coding posture on and cwd inside the
checkout, failing git calls on the second build only => first != rebuilt;
with the snapshot pinned, identical git failure => byte-equal. The real
block's byte-stability is coding_context's own contract; this test is
about plugin sections.
2026-08-18 02:12:30 -07:00
Bryan Bednarski a51337ecee Merge main into fix-openai-sparse-response-objects
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-17 18:58:48 -07:00
Bryan Bednarski 0f5f5a5996 Merge main into fix-relay-client-timeout-payload
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-17 18:55:18 -07:00
Jeffrey Quesnelle e818025b4d Merge pull request #85582 from bbednarski9/codex/fix-relay-lazy-completed-streams
fix(relay): unwrap lazy completed streams
2026-08-17 21:39:52 -04:00
Jeffrey Quesnelle 75bbc055a7 Merge pull request #85579 from bbednarski9/codex/fix-relay-canonical-operation-names
fix(relay): use canonical managed operation names
2026-08-17 21:29:18 -04:00
Teknium 6e22d26583 feat: project-skill quarantine + non-interactive trust inheritance
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.

Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).

Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.

Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
2026-08-17 14:06:16 -07:00
Lavie 9cf553ca3e refactor(agent): document meta api_mode fallback intent + review cleanups
Implement Claude Opus review findings for Meta API support:
- Document in agent/agent_init.py that provider="meta" without an api.meta.ai URL falls through to chat_completions by design (URL-driven wire selection).
- Comment on suppression guard in hermes_cli/runtime_provider.py noting api.meta.ai is handled by _detect_api_mode_for_url.
- Replace inline __import__ with top-of-module import in tests/hermes_cli/test_model_switch_openai_api_mode.py.
- Rename test_meta_retention_not_sent_when_overridden -> test_meta_retention_override_wins in tests/agent/transports/test_meta_codex_cache.py.
- Add test in tests/agent/test_meta_agent_init.py for provider="meta" fallback without api.meta.ai URL.
- Add test in tests/agent/test_auxiliary_client.py for prompt_cache_retention: "24h" under _CodexCompletionsAdapter.

Source: Claude Opus review findings for feat/meta-api-support.
2026-08-17 12:58:51 -07:00
Lavie d4658ee6d5 fix(agent): preserve provider-slug rewrite when host mandate fires
Relocate host_mandated_api_mode check from top of api_mode cascade to
fallback else branch so URL-based provider-slug rewrites (e.g.
api.anthropic.com -> provider='anthropic') always run first. Previously
the mandate branch set api_mode for api.anthropic.com without rewriting
provider, leaving provider='' and causing credential_pool_matches_provider
to fail closed and discard anthropic-scoped pools (#63425 regression
introduced in 8f60e8263).

The mandate is now a true fallback for hosts without an elif branch
(api.meta.ai -> codex_responses for 93-99% prompt-cache hits vs 0% on
chat, plus future mandates) with lazy import + try/except preserved.

Add regression tests: provider=None + api.anthropic.com URL implies
provider='anthropic'/api_mode='anthropic_messages' and preserves an
anthropic credential pool; provider=None + api.meta.ai URL implies
codex_responses.
2026-08-17 12:58:51 -07:00
Lavie 24545418eb fix(providers): route api.meta.ai through Responses API for prompt caching
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
  api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
  93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
  api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
  (after explicit api_mode wins, before provider-name specials) via lazy
  import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
  24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
  override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
  auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
  Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
  content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
  usage cache reporting, model-switch override, and config roundtrip; extend
  test_model_switch_openai_api_mode with meta cases.
2026-08-17 12:58:51 -07:00
Teknium f891d702df feat: project-local skill discovery with per-repo trust gate
Sessions started inside a git checkout now source skills from
<root>/.hermes/skills/ and <root>/.agents/skills/ (the cross-tool
convention shared with other agent harnesses) as the highest-precedence
skill tier: project > local > external_dirs.

Loading is trust-gated per repo (skills.trusted_project_dirs, managed by
'hermes skills trust'/'untrust') because skills are executable procedure
documents — auto-sourcing them from any cloned repo is a prompt-injection
vector. Untrusted repos with skills get a one-line banner notice instead.

- agent/skill_utils.py: find_project_root, get_project_skills_dirs,
  get_untrusted_project_skills_root, get_scan_ordered_skills_dirs;
  project dirs join the curator read-only ownership boundary
- agent/prompt_builder.py: project tier scanned first, entries tagged
  [project], same-named local entries shadowed; cache key extended
- tools/skills_tool.py: skills_list scans project dirs first (first-wins);
  skill_view resolves cross-tier collisions in favor of the project tier
  (same-tier ambiguity still refuses); security warning recognizes the tier
- agent/skill_commands.py + hermes_cli/commands.py: /skill-name slash
  commands and gateway slash menus include project skills
- tools/credential_files.py: project dirs mounted into remote backends
- cli.py: banner notice (loaded count / trust hint)
- hermes_cli/main.py + subcommands/skills.py: hermes skills trust/untrust
- config: skills.project_discovery (default on), skills.trusted_project_dirs
- docs: Project-Local Skills section in skills.md
- tests: tests/agent/test_project_skills.py (18 cases)

Session cwd is fixed at agent build time, so the resolved tier is stable
for the conversation and the system prompt stays byte-stable (cache-safe).
2026-08-17 11:39:13 -07:00
Alex Fournier 7ee68cca45 test(relay): cover lazy completion with interceptor
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-17 10:21:32 -07:00
Alex Fournier cc3418e069 fix(openai): complete sparse response normalization
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-17 10:17:52 -07:00
Alex Fournier 256d7efbfd test(relay): document stream priming boundary
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-17 09:39:11 -07:00
Jack Lau cf64ca20c5 fix(compression): stop an aborted rotation from growing the parent it could not publish
The rotation path flushes its un-persisted transcript to the parent (#47202)
and only then calls publish_compression_child. The abort handler rolls back
the in-memory transcript and keeps agent.session_id on the parent - its own
comment says "keep the parent live and discard the stale compacted snapshot" -
but the rows the flush just wrote are not part of what it discards. Every
failed rotation therefore leaves the parent transcript longer than it found
it, whatever the failure was.

That is survivable for a one-off failure and pathological for a sticky one.
A parent row carrying ended_at fails the publish on every attempt and nothing
in this path clears it, so each auto-compaction appends another copy of the
current turn to the transcript it was supposed to shrink. Worse, the growth
then satisfies conversation_compression's own len(durable_parent) >
len(messages) check, so the next attempt adopts the inflated snapshot as if it
were genuine concurrent activity and the in-memory transcript doubles too.

Check that one precondition before writing. It is a plain read of the row the
publish is about to read anyway, and it raises the publish's own message, so
split_status=aborted, failure_class=session_split_failed and the rollback path
are all unchanged; a live parent reaches the flush exactly as before.
Deliberately not extended to the compression lease, which is re-acquirable - a
transient miss there would abort a rotation that would otherwise have
committed. old_session_id moves above the flush so a failure raised from here
takes the same in-memory rollback as any other pre-publish failure.

Scope: this fixes the amplification for every abort cause. It does not fix
what marks a live session as ended in the first place (#88197 Bug 1), which
needs a maintainer decision on end-reason taxonomy and is tracked on the
issue; an affected session still aborts every attempt, it just stops making
itself larger while it does.

Refs #88197
2026-08-17 18:15:46 +05:30
Teknium d5167831b8 Port from can1357/oh-my-pi#7306: reject answer-shaped auto-title output
A tiny title model that ignores the 3-7 word titling task and answers
the user's first message instead used to have its whole reply stored
(truncated at 80 chars) as the session title. Truncating an assistant
blob still leaves an assistant blob — generate_title now rejects output
over 12 words and returns None, letting maybe_auto_title retry on the
next exchange. The 80-char truncation remains for genuine-but-wordy
titles that pass the word bound.
2026-08-16 22:10:02 -07:00
Teknium 4be4b9866e Port from earendil-works/pi#7494: preserve Gemini 3 tool call IDs
Gemini 3+ models require explicit tool call IDs on functionCall /
functionResponse parts in replayed history; without them parallel tool
calls can be rejected or mispaired. The native adapter now:
- threads the model id into request building and includes ids for
  Gemini >= 3 (version-gated: 2.x rejects unexpected id fields)
- preserves provider-returned functionCall.id on both non-streaming
  and streaming responses instead of always minting a random one
2026-08-16 22:07:52 -07:00
Teknium a8d5e16ccf Port from earendil-works/pi#7681: support AGENTS.override.md context override
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
2026-08-16 22:07:43 -07:00
Teknium 7533630150 fix(error_classifier): classify connect/DNS failure messages on generic exception types
Port from anomalyco/opencode#40707: connection-establishment and DNS
failure messages wrapped in generic exceptions (RuntimeError from local
shims, MCP bridges, SDKs re-raising without chaining) fell through to
FailoverReason.unknown, which misses the retry loop's eager transport
fallback — the full retry budget burned against a dead endpoint before
provider fallback.

New _CONNECTION_MESSAGE_PATTERNS (connect refused, no route, network
unreachable, DNS phrasings across Python/glibc/macOS/Node, fetch failed,
Envoy upstream connect error) classify as retryable timeout via
_classify_by_message, mirroring _TIMEOUT_MESSAGE_PATTERNS. Mid-stream
disconnect strings are deliberately excluded — they keep their
_SERVER_DISCONNECT_PATTERNS routing (large-session compression).
2026-08-16 22:07:25 -07:00
Victor Iglesias 070c6a5f8d fix(curator): abort rollback when safety snapshot fails 2026-08-16 22:06:41 -07:00
Teknium c7727540f3 test: strengthen empty-response guard tests (salvage follow-up for #75115) 2026-08-16 22:06:08 -07:00
Shannon Sands d10f87245e refactor(agent): move empty-response guard settings from env vars to config.yaml
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:

  agent:
    empty_response_guard:
      enabled: true            # false = legacy fixed 3-retry behaviour
      cost_threshold_usd: 0.25 # per-attempt cost that halves the budget

- hermes_cli/config_defaults.py: new documented subsection under agent
  (additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
  section to (enabled, threshold) with fail-open tolerance for
  malformed values; guard_enabled()/_cost_threshold_usd() now read the
  init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
  agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
  following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
  covering malformed sections, YAML string booleans, bad thresholds,
  and a DEFAULT_CONFIG sync check; new integration test proving
  enabled:false restores the legacy 1+3-call behaviour.

Requested by isak-ialogics on PR #75115.
2026-08-16 22:06:08 -07:00
Shannon Sands ac06c2ff8b fix(agent): stop re-billing deterministic empty responses (NS-503)
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).

Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.

New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:

- Deterministic-empty detection: two consecutive empty attempts with
  usage present, output_tokens == 0 (reasoning tokens count as output),
  and identical (model, provider, finish_reason) skip the remaining
  retries and go straight to the fallback chain — a different model may
  well answer. Missing usage, nonzero output, or any signature change
  keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
  exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
  empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
  included/subscription routes are untouched.

At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.

Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.

Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.

Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.

Refs NS-503.
2026-08-16 22:06:08 -07:00
chelsealong a9b4ec3126 fix(skills): rescan skill commands cache when active profile changes
Switching Desktop profiles mid-session changes HERMES_HOME but not the
platform scope, so get_skill_commands() kept serving the previous
profile's skill list. A skill only available under the new profile then
looked like a cache miss to callers such as slash.exec, which fall
through to the slash_worker dead path (#88023).
2026-08-16 19:55:03 -07:00
fangliquan b454e4da76 fix(state): release abandoned session database handles 2026-08-16 19:53:07 -07:00
Teknium bab7be3ca7 feat: raise Codex OAuth context to 900K for gpt-5.6 family and gpt-5.4 (subscription 1M rollout)
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.

gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
2026-08-16 18:31:54 -07:00
Teknium 5229975438 feat: raise Codex OAuth context to live-verified 350K for gpt-5.6 family and gpt-5.4
The Codex /models catalog advertises 272K for the gpt-5.6 (sol/terra/luna)
and gpt-5.4 slugs, but the backend actually accepts ~371K input tokens
(verified live against chatgpt.com/backend-api/codex/responses, Aug 16 2026:
~371K completed OK on all four slugs; ~382K+ rejected with
context_length_exceeded). 350K keeps ~22K margin under the observed ~372K
enforcement.

The bump applies ONLY when the resolved value is exactly the known-stale
272,000 advertisement — any other advertised value (higher or lower) is
trusted as a real server-side change, so a future catalog correction
deactivates the override automatically. gpt-5.5 and gpt-5.4-mini both
genuinely enforce 272K (rejected 360K live) and are excluded.
2026-08-16 15:43:09 -07:00
Teknium c257e9196b fix: make every tool interruptible — sequential executor abandons on user interrupt
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.

Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
  daemon worker (timeout None no longer means inline blocking) and polls
  the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
  synthesize a cancelled tool result (_ToolCancelledResult), emit the
  terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
  exactly like _ToolTimeoutResult, so an abandoned worker finishing late
  cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
  it owns its own human wait.

Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
  upscale) replaced with _wait_fal_result(), which polls is_interrupted()
  in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
  the "upscale failed, use original" fallback.

Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
2026-08-16 11:32:02 -07:00
Teknium 587ad8748f Revert "fix(agent): preserve local reasoning timeout opt-out"
This reverts commit 26b2b47593.
2026-08-16 10:51:11 -07:00
Teknium d709d29f19 fix(agent): trim background_review to the enabled switch
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
2026-08-16 10:27:52 -07:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
Teknium 5b19c8ee55 fix(acp): make --acp probe tri-state, cached, and mock-safe
Salvage hardening on top of #87308 (thanks @Dudeman456):

- Tri-state verdict: inconclusive probes (binary missing, --help
  failed/timed out) return None and fall through to the normal spawn
  path, preserving the established 'Could not start Copilot ACP
  command' error instead of masking it. This also fixes the two
  test_copilot_acp_client HOME-env regressions that went red on the
  PR: their mocked-Popen path was intercepted by the new unmocked
  subprocess.run probe.
- Cache definitive verdicts per binary path so CLIs that DO support
  --acp pay the ~50ms --help cost once per process, not per prompt.
- Skip the probe entirely when custom ACP args don't include --acp.
- Fix the help-text regex: the old pattern never matched '[--acp]'
  (leading '[' is neither start-of-string nor whitespace) and \b
  after 'p' matched '--acpfoo'.
- Hermeticity: stub subprocess.run in the two HOME-env tests; add 6
  probe-specific tests (fast-fail, fall-through, caching, skip).
2026-08-16 06:31:46 -07:00
Merge_Conflict - Pasi 309cf2c5e2 fix: honor JSON-array string forms for skills.disabled and agent.disabled_toolsets
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.

Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.

Fixes #86661
2026-08-16 06:24:56 -07:00