Commit Graph

2612 Commits

Author SHA1 Message Date
ywatanabe 419a050427 fix(agent): preserve tool cache in iteration summary 2026-09-14 20:35:28 +05:30
kshitijk4poor 62e5f46656 test(streaming): one Anthropic event-stream fake for the three parse-error tests
Three inline context-manager classes shared the same __enter__/__exit__
boilerplate and differed only in the events yielded before the raise.
Also drop the unreachable 'or agent.base_url' fallback: every
anthropic_messages init path sets _anthropic_base_url, and the two
sibling call sites read it bare.
2026-09-14 20:01:19 +05:30
kshitijk4poor 982e504262 fix(anthropic): partial tool names are reset per stream attempt
A tool_use block name recorded by a stream attempt that died before any
visible text survived into the next attempt: only the deltas_were_sent
mid-tool branch cleared result["partial_tool_names"]. When the retry then
streamed plain text and dropped, the partial stub blamed the stale tool
("Stream stalled mid tool-call (old_tool)") and the stale name could make
a later attempt look mid-tool-call when deciding whether the drop is
retryable. Reset it in _start_stream_attempt alongside
provider_tool_in_flight, which already has attempt-local semantics.

Regression: two-attempt stream (tool_use start + parse error, then text +
drop) — stub content and emitted deltas carry no stale tool name.
2026-09-14 20:01:19 +05:30
kshitijk4poor 3d88259483 fix(anthropic): retry a malformed tool-JSON stream with buffered tool input
Follow-up to the cherry-picked fix: keep the classifier widening
("expected value at line" is now a transient stream parse error on the
main turn) but replace the messages.create() fallback with a retry on the
same stream wire.

Why not create(): the fallback ran outside Relay (lost request rewrites),
outside _handle_stream_error (could replace text already shown to the
user with a different generation), ticked no liveness events for the
whole buffered payload, and a bare identical retry still re-emits the same
malformed JSON.

Why not drop the fine-grained-tool-streaming beta (#108583/#109056): live
probe on claude-sonnet-4-5, ~500-line tool call - beta on: max inter-event
gap 1.6 s; beta off: 139 s zero-event gap while Anthropic buffers the
args, which the 180/240 s stale-stream detector kills on larger payloads
(the regression 80a899a8e2 fixed).

Instead, on a parse error the retry sets `eager_input_streaming: false`
on every tool for that request only (the SDK/API per-tool field overrides
the legacy beta header), so Anthropic returns buffered, server-validated
args while the happy path keeps fine-grained streaming. A tool_use that
started streaming is registered in partial_tool_names so the mid-tool
transient retry fires the same way it does on the chat_completions wire.
2026-09-14 20:01:19 +05:30
joaomarcos dfb4caf4b7 fix(anthropic): recover malformed streamed tool JSON 2026-09-14 20:01:19 +05:30
teknium1 d57c28a554 fix(tests): compression stall-fallback tests stop racing the 0.2s ceiling
Under a loaded runner the primary stall plus the fallback retry overran the
0.2s total ceiling, so the retry never started and attempts==1 failed
intermittently (seen once in a 40-worker tests/agent run). Idle stays 0.05s;
the ceiling moves to 2s per the >=2s wall-clock rule in AGENTS.md.
2026-09-14 07:21:18 -07:00
joaomarcos 6bc0e9e6df fix(agent): same-model review fork keeps the parent's affinity header and Portal conversation root (#109964)
Trimmed salvage of #110045 (deltas 1 + 2 only), stacked on the #110009 scope inheritance:

- `declared_conversation_scope` treats an inherited value as a DECLARED scope only when it
  carries the `gwk_` prefix. A rotated CLI parent publishes no affinity scope (None → sticky
  key falls back to the conversation root); the fork now publishes exactly the same instead
  of an explicit physical lineage root. `resolve_prompt_cache_scope` honors any inherited
  value directly, so the body `prompt_cache_key` still matches.
- `build_cache_parity_fork` snapshots `parent._conversation_root_id()` as
  `_cached_conversation_root`; with `_session_db=None` the fork's own walk fell back to the
  parent's PHYSICAL id, so after a compression rotation the review's Portal
  `conversation=` tag fragmented usage attribution across one logical conversation.

Dropped from the original: copying `_gateway_session_key` onto the persistence-detached
fork (no cache-identity consumer reads it there; the compression-boundary hooks were
deliberately severed by `_detach_fork_compression`), and the defensive
hasattr/callable/try wrapper around `_conversation_root_id()`.
2026-09-14 06:55:54 -07:00
salch-cred a4b620f17c fix(agent): same-model review fork inherits the parent's resolved cache scope (#109964)
build_cache_parity_fork gives the same-model fork the parent's session_id,
cached system prompt, tools[] and session_start — but with
_persist_disabled=True and _session_db=None, BOTH cache-identity resolvers
diverged from the parent on their own: declared_conversation_scope failed
closed on _persist_disabled, and the lineage walk skipped on the missing
DB. The fork's affinity header (set_affinity_scope) and body
prompt_cache_key (cache_scope_id on the OpenAI-wire transports) therefore
keyed a different bucket than the gateway parent, costing one cold
~full-context request per review. Not gateway-only: any parent whose
lineage root != current physical id diverges too (teknium1's triage table).

Fix, per the triage's suggested direction: on the not-routed branch only,
the fork stamps _inherited_cache_scope = resolve_prompt_cache_scope_safe
(parent) — the parent's ALREADY-RESOLVED scope, no DB access from the fork,
persistence fully detached. Both declared_conversation_scope and
resolve_prompt_cache_scope return the inherited scope first when set, so
the header path and the body path are fixed together (fixing only one
leaves the other divergent — Vivamisu's header/body split observation).
Routed (different-model) forks, /branch children, delegate/tool children
and fresh sessions set nothing; the fail-closed default stands untouched.
/btw shares build_cache_parity_fork and gets the repair for free.
2026-09-14 06:55:54 -07:00
teknium 743140cd82 feat(gemini): send full JSON Schema tool parameters via parametersJsonSchema
Clean-room port of the approach in zed-industries/zed#63342. The native
Gemini adapter previously down-translated every tool schema into the
restricted FunctionDeclaration.parameters subset, which was lossy: anyOf
unions without an outer type, bare arrays, $ref/$defs indirection and
additionalProperties had to be stripped or repaired, and one
unrepresentable construct could 400 the entire request (live repro:
INVALID_ARGUMENT ...properties[bare_array].items: missing field).

Google now accepts plain JSON Schema in parametersJsonSchema on all
current models. The adapter sends full schemas through that field; the
old subset translator is replaced by a light normalizer that deep-copies,
strips root $schema, inlines same-document $refs (MCP pydantic / zod
emit them; unresolvable or circular refs pass through untouched with the
reason logged), and guarantees an object root.

Live-verified against the real API: the union+bare-array+$ref schema
that 400s through the legacy parameters field is accepted with 200 via
parametersJsonSchema on gemini-3.7-flash and gemini-2.5-flash, and
gemini-2.5-flash returns a correct functionCall against it.
2026-09-14 06:44:54 -07:00
liuhao1024 47714f9402 fix(agent): routed background reviews honor auxiliary.background_review.reasoning_effort
The review fork is a full AIAgent, not an auxiliary_client call, and its
routed branch deliberately skips the parent's reasoning_config (the parent's
effort vocabulary may be invalid for the routed provider). It also never read
the per-task key, so an explicit `auxiliary.background_review.reasoning_effort`
was silently ignored and the routed fork ran at the provider default (#94825).

Routed forks now parse the task key through the shared parse_reasoning_effort
(same levels and `none` alias as every other aux task); unset keeps the
provider default, an unknown level warns and falls through. The same-model
path is untouched: it still inherits the parent's reasoning_config verbatim
for prompt-cache parity.

Salvaged from #94832 (liuhao1024), re-applied on the decomposed
_fork_init_kwargs seam.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-14 05:25:01 -07:00
Finn763 6da540d3ec fix(background_review): surface the ignored reasoning_effort on same-model forks (#104116)
auxiliary.background_review.reasoning_effort was a silent no-op on the
same-model path: the fork inherits the parent's reasoning_config verbatim to
keep prompt-cache parity (#30532), and nothing told the user. Emit a one-time
user-visible warning (parent-scoped, so a nudge-per-turn session warns once,
not per fork) when the key is actually set, document the no-op in the config
reference, and leave the fork-birth request bytes unchanged.
2026-09-14 05:25:01 -07:00
teknium1 5280dec0ee fix(redact): repr pass leaves already-masked values alone
The Python-repr pass ran after the MCP probe header scrub and collapsed
'Authorization': 'Digest ***' to '***', erasing the scheme word that scrub
deliberately keeps (tests/hermes_cli/test_mcp_probe_redaction.py went red
on rebase). Skip values that already carry a mask marker.
2026-09-13 21:30:40 -07:00
Teknium 92d4c0233e fix(redact): widen repr-field key class to mixed-case credential suffixes
Port from OpenHands/software-agent-sdk#4508: their dict-entry secret
redaction was uppercase-only and leaked mixed-case keys (UserPassword,
sessionToken). Apply the same case-insensitive treatment to the Python
mapping-repr pass: a casefolded credential suffix (apikey/token/secret/
password/passwd/credential) now qualifies a key, while metadata names
(TOKEN_COUNT, password_policy, tokenizer) stay untouched.
2026-09-13 21:30:40 -07:00
Demerzel (backend-eng) 47668280e7 fix(redact): mask secrets in Python mapping reprs 2026-09-13 21:30:40 -07:00
Teknium 0ddb62bce6 fix(agent): streamed reasoning_details survive for replay continuity (port of earendil-works/pi#8605)
The streaming accumulator dropped delta.reasoning_details entirely — only
non-streaming responses preserved the OpenRouter unified reasoning replay
data (signatures, encrypted reasoning blocks). Providers that require the
reasoning_details sequence passed back on the next turn (OpenRouter
reasoning models, Anthropic signed thinking via shims) lost continuity on
every streamed turn, which is nearly all turns.

- Accumulate delta.reasoning_details (attr or model_extra) during
  streaming and attach the merged list to the final mock message, where
  _build_assistant_message's existing passthrough persists it.
- _append_streamed_reasoning_detail merges consecutive reasoning.text /
  reasoning.summary fragments into one logical entry (OpenRouter streams
  them as word-level deltas; unmerged they bloat the replayed signature
  payload) while encrypted/opaque entries stay discrete. Later fragments
  backfill signature/id/format/index fields the first fragment omitted.

Ported from earendil-works/pi#8605 (commit c5ad7c1b0), credit
@cristinaponcela for the merge-fragments pattern.

Tests: tests/run_agent/test_streamed_reasoning_details.py (8 tests;
sabotage-verified: disabling accumulation fails the E2E test, disabling
merging fails 4 unit tests). Neighboring suites green (test_streaming.py
40 passed).
2026-09-13 21:27:35 -07:00
teknium1 8a5a66d0d8 test: trim the Aug-2026 conformance suites to behaviour invariants on main
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.

Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
  hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
  hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
  vocabulary since #100479, so the "full accounting" row names the serve
  unit instead of borrowing a gateway relaunch.

Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
  the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
  with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
  pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
  plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
  round-trips, absent-registry, sequential "race" re-enactment.
2026-09-13 21:13:23 -07:00
Teknium d15f4a7826 test: regression conformance suites for the 7 recurring Aug-2026 bug classes
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:

- tests/conformance/test_profile_write_tripwire.py — no writes to the
  default profile tree while a profile is active (#88532 #92662 #89190
  #89625 #92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
  table for the Deprecated-.env warning (#88829 #89016 #89389 #90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
  flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
  strippable injected param x real 400 shapes must strip-and-retry;
  unknown params must still fail (#90257 #89897 #91164 #89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
  law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
  enumerated maintenance ops refuse/degrade under a live writer; copies
  of corrupt DBs are refused or flagged (#91839 #90806 #90613 #88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
  Bot Chat resolution is idempotent, never mints, unique per profile,
  race-safe (#92040 #90705 #92692 #90005 #90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
  crash never claims success; success requires full fleet accounting;
  refusal != failure (#91283 #91439 #92902 #92780)

117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
2026-09-13 21:13:23 -07:00
teknium1 93b9880d84 test: drop the tautological prompt-rendering assertion
The prompt template is built from _PROMPT_GOOD_EXAMPLES, so asserting the
template contains them can never fail independently of the code it checks.
The behavioural invariants (echo rejected, greeting allowed, near-miss
passes) stay.
2026-09-13 21:08:32 -07:00
Teknium f0d194df81 Port from QwenLM/qwen-code#9709: reject session titles that echo the prompt's own examples
Small title models parroting a prompt example back verbatim produced
sessions named "Fix login button on mobile" with no relation to the
conversation. The example lines in _TITLE_PROMPT_TEMPLATE now render
from _PROMPT_GOOD_EXAMPLES so the guard set and prompt cannot drift,
and generate_title rejects exact (case-insensitive, wrapper-stripped)
echoes so the instant derived title survives instead. 'Friendly
greeting' stays allowed — it is prescribed output for bare greetings.
2026-09-13 21:08:32 -07:00
Teknium cc81e436ce fix(models): delist hy3-free and laguna-s-2.1-free — OpenCode relay dropped them (anon 401)
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).

- hermes_cli/models_catalog_static.py: remove both slugs from the
  opencode-free offline floor and the opencode-zen discovery floor;
  document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
  dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
  anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
  invariant to cover both.

The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
2026-09-13 21:06:39 -07:00
teknium1 bc86ff6c4a test: cover OpenAI spend-limit codes on the documented 429 path too
OpenAI documents credit_balance_exhausted / *_spend_limit_exceeded /
organization_usage_limit_exceeded as HTTP 429 responses, so the invariant
must hold through _status_429 (which short-circuits _by_error_code), not
only on the status-less body path.
2026-09-13 20:49:13 -07:00
teknium 3ed62fc5e8 fix(agent): classify OpenAI spend/usage-limit error codes as billing
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
2026-09-13 20:49:13 -07:00
teknium1 a6c4278084 test: exercise the streaming finish_reason fold end-to-end, trim to invariants
The streaming choke point was covered only by an identity check on the
imported alias (a change-detector that would pass with a no-op wiring).
Replace it with a real chunk-loop run asserting STOP -> stop and
MAX_TOKENS -> length on the assembled response, and collapse the
duplicate transport cases into one parametrized invariant.
2026-09-13 20:47:20 -07:00
Hermes Agent 2598247ff5 Port from can1357/oh-my-pi#9566: uppercase finish_reason (STOP/MAX_TOKENS) no longer bypasses stop/length handling
Some OpenAI-compatible gateways fronting Gemini backends emit the native
uppercase finish reasons (STOP, MAX_TOKENS) instead of the lowercase
OpenAI contract values. Every downstream comparison in Hermes uses
lowercase literals, so an uppercase reason silently fell through: a
clean STOP completion missed the stop handling and a MAX_TOKENS
truncation never entered the length-recovery path.

Adds normalize_finish_reason() as the single owner in
agent/message_sanitization.py (case fold + alias map: max_tokens->length,
end->stop, function_call->tool_calls) and wires it at both wire-intake
choke points: ChatCompletionsTransport.normalize_response and the
streaming chunk-capture loop in chat_completion_helpers. Non-string and
empty values pass through unchanged so existing 'or "stop"' defaults
and the Poolside int-reason path keep their behavior.
2026-09-13 20:47:20 -07:00
teknium1 f80ea9987b fix: grep/awk/sed gate on file operands only; strip $HOME prefixes
Adding grep/awk/sed to _FILE_READ_COMMANDS made the PATTERN operand
participate in the secret-file predicate, so `grep .bashrc app.py` or
`grep -n .env src/settings.py` — reads of SOURCE files — ran the
ENV/YAML assignment pass and masked opaque values that main leaves
alone. Skip the first non-flag positional for the pattern-first
readers, as #109369 originally did, so only real file operands gate.

`cat $HOME/.hermes/config.yaml` was ungated because the `$` bail-out
fired before the `.hermes` segment was inspected; strip `$HOME/` and
`${HOME}/` like the HERMES_HOME prefixes.

Review follow-up on #110228.
2026-09-13 14:47:12 -07:00
teknium1 6e7ea9cbdc refactor(redact): one secret-file predicate for .env, shell rc and Hermes config.yaml
Fold KoNit-K's `_command_reads_secret_bearing_file` and the pre-existing
`_command_reads_env_file` into a single `_command_reads_secret_file` so the
`code_file` gate in `redact_terminal_output` has one owner: `.env`-style
basenames and shell rc/profile files anywhere, `config.yaml` only under a
`.hermes` directory or `$HERMES_HOME` (arbitrary project YAML stays on the
code_file path). `grep`/`awk`/`sed` join the reader set instead of a second
table with a positional-argument special case: on a file read, any non-flag
operand that names a secret-bearing file is enough — the pattern/program
operand never matches a basename, so the extra rule bought nothing.

Tests: the negative parametrization now uses an opaque credential-shaped value
(the placeholder it used before would never have been masked on either path,
so the "stays unredacted" half proved nothing) and asserts the same value IS
masked under `cat .env` in the same test.
2026-09-13 14:47:12 -07:00
KoNit-K 8ab9e79d0a fix(redact): recognize quoted HERMES_HOME config reads
Keep the narrow basename allowlist, but do not treat $HERMES_HOME as an
unresolved path, and split pipelines only on unquoted |;&.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 14:47:12 -07:00
KoNit-K 0997a23e57 fix(security): redact secrets from config file reads 2026-09-13 14:47:12 -07:00
teknium1 a8843ab985 fix: treat a reasoning-only stream drop as a drop, not a clean stop
The text-only drop guard in _finish_chat_stream required content_parts,
so a stream that died while still emitting delta.reasoning (no
finish_reason, no usage) fell through to the synthesized "stop". With
the reasoning-only clean-stop promotion in finish_text_response that
stamped "stop" turned the truncated thought into the final answer,
where main entered the continuation ladder. Extend the guard with
reasoning_parts so the drop yields the partial-stream stub and the
ladder still runs; a real clean stop carries finish_reason="stop" and
is unaffected.

Review follow-up on #110227.
2026-09-13 14:46:42 -07:00
teknium1 28931c5c04 fix(agent): persist promoted clean-stop reasoning; pin the length negative
Follow-up to KoNit-K's commit: rebuild the promotion on the existing
`agent._extract_reasoning` helper (the same reader the ladder terminal and
`build_assistant_message` use) and write the promoted text back onto
`assistant_message.content` so the persisted assistant row carries the answer
as ordinary content. Without that the transcript tail was an assistant row with
empty content and only `reasoning`, which `drop_thinking_only_and_merge_users`
strips from the next request — the model would see its own answer vanish on a
"continue" turn.

Tests: trim to the two invariants (clean stop → one API call, persisted as
content; `finish_reason == "length"` → never promoted, continuation still
owns it) and keep the truly-empty terminal case. The prefill wire-payload
regression test now drives a non-clean-stop reasoning-only reply, which is the
only shape that still reaches the prefill rung.
2026-09-13 14:46:42 -07:00
KoNit-K 88a9d9eac5 fix(agent): return reasoning-only clean stops 2026-09-13 14:46:42 -07:00
teknium1 53183d5016 fix(vision): advertise vision_analyze/browser_vision when the main model sees natively
check_vision_requirements only asked the auxiliary resolver, so a vision-capable
main model on a provider the resolver cannot serve (minimax-oauth, local vLLM,
anything uncatalogued) lost vision_analyze and browser_vision from the tool list
even though both handlers already route to the native fast path and work when
called. The image gate now accepts the native fast path OR an aux client; the
aux-only probe becomes check_video_requirements and stays on video_analyze,
whose handler has no native path.

Fixes #47149.
2026-09-13 14:46:38 -07:00
teknium1 6449a6c0e5 test(secrets): trim salvaged #108446 coverage to the two invariants (retry after failure; snapshot replaced on retry) 2026-09-13 14:41:26 -07:00
fangliquanflq 4a927e9a6e fix(gateway): revoke stale profile secret snapshots 2026-09-13 14:41:26 -07:00
fangliquanflq d461b27dcd fix(gateway): clear stale profile secret snapshots 2026-09-13 14:41:26 -07:00
fangliquanflq 917cfbd740 fix(gateway): retry failed profile secret hydration 2026-09-13 14:41:26 -07:00
teknium1 76e88cae2a test: repoint cross-test imports at the moved modules; opt one backoff test out of the merged agent fixture
Twenty-three tests imported helpers from sibling test modules by dotted
path (tests.run_agent.test_run_agent, tests.cli.test_cli_init,
tests.gateway.test_42039_duplicate_user_message); those follow the moves
and renames. The run_agent autouse fixture that zeroes jittered_backoff
now covers all of tests/agent, where one pre-existing test asserts the
real backoff is positive — it opts out via a real_retry_backoff marker,
the same pattern real_concurrent_gate and real_agent_prewarm use.
2026-09-13 09:18:02 -07:00
teknium1 d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
kshitijk4poor aa40c1d21b fix(auxiliary): wrap the Messages adapter on the profile's declared api_mode
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
2026-09-13 19:46:52 +05:30
kshitijk4poor 5dd8fa1daa fix(auxiliary): anthropic_messages profiles keep /reasoning reachable on aux calls
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.

Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.

Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
        -> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After:  -> _reasoning_config={"enabled": False}.
2026-09-13 19:46:52 +05:30
kshitijk4poor 5c4c31cf4d refactor(agent): fail loudly on a missing turn clock; single lowercase in is_dangerous_confirmation
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
  turn prologue now raises instead of silently falling back to per-request wall time, which
  would re-create the mid-turn drift the fix removes. Only production caller
  (assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
  tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
  untrustworthy-stamp contract is its own test.
2026-09-13 19:23:09 +05:30
kshitijk4poor f296652a66 fix(agent): interrupt detection needs the executor's result shape; future-stamped confirmations expire
Two replay-boundary findings from the #109320 review (gaoanze888, andrexibiza):

- is_interrupted_tool_result matched the bracketed marker anywhere in the payload, so a
  successful `read_file`/`search_files` result QUOTING "[Command interrupted]" (a doc, a
  grep hit) was classified as a killed run — the read-only block vanished from the
  model-facing history, a `terminal` result became the UNKNOWN-effect notice — while
  send/replay parity still held because both consumers made the same mistake. Require
  the shape every executor actually produces: the marker is the LAST line of `output`,
  inside a JSON envelope with a non-zero exit code (or a bare text result). Real
  interruptions from tools/environments (130), managed_modal (130) and
  code_execution_tool (-1) all keep matching.
- strip_stale_dangerous_confirmations only bounded age from above, so a finite FUTURE
  stamp (clock skew, corruption) had negative age and kept the confirmation plus its
  api_content sidecar live. Freshness is now 0 <= admission - ts <= expiry; a future
  stamp expires like a corrupt one. Missing stamps (legacy rows) stay untouched.

Test: the parity fixture carries both quoted-marker results (JSON grep hit, bare doc text)
and a genuine execute_code interruption envelope; the expiry test covers "nan" and two
future epochs. Red under: marker-anywhere, any-line, old loose heuristic, upper-bound-only.
2026-09-13 19:23:09 +05:30
kshitijk4poor 820d3ca65d fix(agent): send-path canonicalization is prefix-only, clock is admission time, corrupt stamps fail closed
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:

- Prefix-only on the send path. build_api_messages now canonicalizes only
  messages[:current_turn_user_idx]; rows the current turn appended (its own tool
  calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
  the previous iteration had already sent whenever a tool result matched the
  interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
  to remove, and it also made the dangling-tail transform order-dependent on when
  the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
  "exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
  `grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
  result to an orphan notice (or dropped a read-only block). That heuristic was
  tolerable at resume time only; it now runs per request. Match the executors'
  bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
  stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
  live on the send path while replay expired it. _reset_per_turn_agent_state stamps
  time.time() once at admission; the three other writes (bind identity, stage
  message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
  `"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
  and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
  treat an unknowable age as expired; missing stamps (legacy rows) are still left
  alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
  main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
  _reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
  ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
  the durable list untouched, live tail preserved; admission-clock freeze across
  iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
  (send path unpatched, whole-list canonicalization, per-request clock, fail-open,
  loose heuristic).
2026-09-13 19:23:09 +05:30
joaomarcos 401fef6e6f fix(agent): freeze turn confirmation expiry and verify wire parity 2026-09-13 19:23:09 +05:30
joaomarcos e5ca5207de fix(agent): unify replay history canonicalization 2026-09-13 19:23:09 +05:30
kshitijk4poor 158ab33518 refactor(copilot-acp): scan session configOptions once per selection; drop unused binding
Also widen the two real-subprocess list_models tests from a 2 s to a 30 s response deadline:
spawning a Python interpreter under 16 parallel test workers occasionally exceeded 2 s and the
file flaked (TimeoutError in _request). The deadline only bounds a failure; the happy path
returns as soon as the fake server answers.
2026-09-13 19:17:06 +05:30
Gille 0cd897286a fix(models): discover Copilot ACP session catalog 2026-09-13 19:17:06 +05:30
Erosika 58b117ed5b test(honcho): stub the three-tuple _get_or_create_honcho_session return
The session manager now returns observation flags as a third value.
The user-id scoping test still stubbed the old two-tuple and failed to unpack.
2026-09-13 19:05:39 +05:30
teknium1 de114b3af1 refactor(platforms): one scoped-secret reader and spec-driven enablement/YAML-bridge boilerplate across all adapters
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:

- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
  docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
  `external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
  Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
  keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
  (SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
  bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
  slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
  `_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.

Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:

- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
  irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
  matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
  nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
  `yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
  mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.

Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
  neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
  seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
  env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
  does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
  all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
  "Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
  reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
  as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
  value (was list/tuple/set only) — same env text for every real YAML shape.

Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.

Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
2026-09-13 05:32:38 -07:00
teknium1 b5f072e2ce fix(agent): manual /compress "nothing to compress" gate stays gateway-only
The shared core applied `has_content_to_compress(head) is False -> nothing_to_do`
on every surface, where origin/main only had it in the gateway handler. That
predicate only knows the local summarizer's window: on CLI/TUI/ACP,
`_compress_context(force=True)` still routes codex_app_server sessions to native
compaction before any local-compressor check, and `ContextCompressor.compress`
commits the phase-1 tool-result prune / blank-echo drop even when no summary
window exists -- so the gate wrongly skipped real work there. It is now an opt-in
`skip_without_window` that only the gateway passes, restoring each surface's
prior behavior.

Review follow-up on #109610.
2026-09-13 05:20:26 -07:00