Commit Graph

1176 Commits

Author SHA1 Message Date
Teknium b24a781b6f test(agent): trim the restart-bound tests to two invariants
Collapse the four class-based tests into two parametrized invariants over both
refunding restart flags and move them to tests/agent/ (the phase modules live in
agent/): a single restart still refunds-and-continues; a re-armed restart breaks
after max_retries refunds. The stub grows the redirect seam the follow-up commit
uses so the queued-correction contract is covered by the same test.
2026-09-09 09:51:31 -07:00
yoyodine-industries e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
kshitijk4poor 9e6c4100cb fix(agent): close the interrupted tool tail on the overflow terminal
The overflow-terminal path ends the turn without reaching finalize_turn, so
a transcript that overflowed right after a tool batch ended on a raw tool
result; strict providers reject the next user turn (tool -> user). Close it
with the same final text, mirroring the truncated-tool-call terminal above.

Also: classify once before either log so an overflow no longer emits a
"so the loop can continue" WARNING followed by the contradicting "NOT
seeding" one; reset the stale-streak breaker once for both branches; drop
the "compression could not recover it" wording (this path never reached
compression); trim the test file to the three tests that bind behaviour
(stream -> terminal stub; 413 stays non-terminal; the terminal ends the
turn, closes the tool tail, carries compression_exhausted).
2026-09-09 17:45:12 +05:30
ca-shrimp 3b0e81459b fix(agent): carry compression_exhausted bit; scope overflow terminal to context_overflow
Address review P1s (andrexibiza) on #106266:

1. The overflow-terminal exit in recover_from_truncation now forwards the
   #98722 typed compression_exhausted bit (partial_result/end_turn gained the
   flag) so the gateway resets/moves future input to a clean session instead
   of leaving the bloated durable session authoritative for the next turn.

2. _overflow_terminal is scoped to FailoverReason.context_overflow ONLY.
   payload_too_large (413) has its own byte-scored recovery owner
   (turn_overflow._recover_payload_too_large, #88960/#47339) that must not be
   bypassed; a post-delta 413 keeps its normal continuation stub. Regression
   covers both lanes.

Tests: unit asserts result compression_exhausted=True on the marker; a
real streamed partial hitting a 413 payload-too-large error keeps content and
is not terminal. 50 streaming/continuation/gateway regressions pass.
2026-09-09 17:45:12 +05:30
ca-shrimp ece584b8f5 fix(agent): don't seed continuation stub after a context-overflow stream death
When a stream delivered text and then died on a context-overflow /
payload-too-large error, the partial content (often tens of KB) was seeded as
a length-continuation stub, growing the transcript monotonically. In a
session whose transcript cannot be compressed back under budget
(protect_last_n covers everything -> no_progress, or the summary would
itself be larger -> would_grow), every later request is larger than the one
that just failed — an unrecoverable loop where the user sees a 30+ minute
fake hang and the only remedy is killing the session (#106260).

classify_api_error already labels these errors context_overflow /
payload_too_large (should_compress=True). _partial_stream_stub now returns
an EMPTY stub marked _overflow_terminal for that class instead of seeding
the recovered text, and recover_from_truncation treats the marker as
terminal: the turn ends via the recovery contract with a clear message
(start /new) and the transcript is not polluted with the partial.

Normal partials (network stall, output-cap truncation, tool-call drops) are
unchanged — only the overflow error class changes behavior.

Tests: stub marker + empty content; a real streamed partial hitting a
'maximum context length' error returns the terminal stub; recover_from_
truncation ends the turn (no fragment/nudge appended) on the marker while a
normal stub still runs the continuation path. 67 streaming/continuation
regressions pass.
2026-09-09 17:45:12 +05:30
ericmaddox bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
kshitijk4poor 26f4a674e0 fix(agent): a /steer row is human input for every user-turn predicate
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.

Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
2026-09-09 13:08:25 +05:30
kshitijk4poor 91433c8466 fix(loop): the turn-boundary export skips preflight-timeout envelopes and stops re-anchoring the persist index
Follow-up to #106312. _preflight_timeout_result carries the prior history without this turn's
user row (#7100); with a repeated prompt ("continue") the verbatim scan resolved to the
historical copy and exported it as this turn's proven boundary — the exact relabeling the export
exists to prevent. Nothing is exported for that envelope now.

The trailing `agent._persist_user_message_idx = idx` ran after finalize_turn had already flushed
the transcript, so it never influenced a persist and the next turn reset it: dead state, removed.
2026-09-09 12:55:43 +05:30
kshitijk4poor 7dc796463d fix(agent): a persisted /steer row survives the next prompt's alternation repair; typed for history
Both steer sites now build the row through one helper, prompt_builder.steer_user_row:
a role:user row with display_kind="steer" and no leading blank lines. The alternation
repair (_merge_consecutive_users) skips a steer-typed prev row, so a run that ended
right after a steered batch (Ctrl-C, interrupt) does not get the next real prompt
merged INTO the already-persisted steer row — which would have rewritten it in place
and re-broken live≠replay parity, the exact class this PR fixes.

TUI/desktop history projects the steer row as the user's own words instead of the
model-facing marker wrapper; 'steer' joins the display_kind union. The compression
anchor scan keeps its tool-row branch for transcripts persisted before this change and
its docstring says so.
2026-09-09 12:21:28 +05:30
kshitijk4poor 4d0cec9a7d fix(agent): the pre-API-call /steer drain also stops smearing the persisted tool row
Second site of the same bug class #104444 fixes in apply_pending_steer_to_tool_results:
_inject_steer_into_newest_tool_result (the drain that runs when a /steer lands during an
API call) mutated the newest role:tool row in place. That row was already flushed
append-only, so the replayed history diverged from the live request bytes at the
injection point and broke the prompt cache exactly like the post-batch path.

Deliver it the same way: a standalone user row inserted right after the newest tool
result (not yet persisted, so the next flush writes it to the transcript). Restash when
there is no tool row yet, unchanged. Stale comments claiming steer lands "in the newest
tool result" and agent/AGENTS.md's alternation rule now describe the real shape.
2026-09-09 12:21:28 +05:30
kshitijk4poor 0d6e3637ef test: keep the steer suite on the canonical patch targets, not PLUGIN-COMPAT pointers
The cherry-picked commit carried an unrelated hunk repointing three patch()
targets back to run_agent.* — those are PLUGIN-COMPAT re-exports, off limits
in-tree (scripts/check_compat_pointers.py; removed 2026-09-14). Keep main's
model_tools.* / agent.process_bootstrap.OpenAI targets.
2026-09-09 12:21:28 +05:30
Albert.Zhou d24810483d fix(agent): persist /steer as a standalone user message
`apply_pending_steer_to_tool_results` used to smear the steer text onto
the last `role:tool` message's content. That tool row had already been
flushed to the session store and carries `_DB_PERSISTED_MARKER`; the
append-only persistence never rewrites it, so the replayable transcript
diverged from the live request bytes at the injection point — resumed
sessions (surface switch / process restart / background-review close)
missed the provider prompt cache (75-85% hit) and the user's mid-run
instructions were never part of the durable history.

The steer is now emitted as a standalone `role:user` message (marker
text preserved):
- role alternation stays legal: assistant(tool_calls) -> tool -> user is
  the documented 'user jumped in mid-run' pattern that
  `repair_message_sequence` deliberately keeps;
- the appended dict carries no `_DB_PERSISTED_MARKER`, so the next
  `_flush_messages_to_session_db` writes it to the session store —
  transcript bytes and replayed history finally agree, and the steer
  becomes searchable/retrievable like any other user message;
- the no-tool-result fallback (interrupt) still requeues the steer, which
  the caller then delivers as a normal next-turn user message.

Tests: TestSteerInjection updated for the new shape plus a persistability
assertion (no marker => flushable); tool-batch-segmentation malformed
scenario updated. steer + segmentation suites: 67 passed, 1 skipped.
2026-09-09 12:21:28 +05:30
kshitijk4poor 1f6718b0ff test(agent): prove the re-anchor through prepare_iteration; reuse the compaction _reanchor
The salvaged regression test exercised only repair_message_sequence and
reanchor_current_turn_user_idx — pre-existing helpers — so reverting the fix left it
green. It now drives prepare_iteration on a real AIAgent with adjacent user rows and
asserts the returned index addresses this turn's row and mirrors into
_persist_user_message_idx (red without the re-anchor: IndexError).

Both re-anchor sites (repair and compression restart) call
turn_context_compaction._reanchor instead of inlining "reanchor + mirror", so they
cannot drift. The export tests fold into one parametrized invariant plus the
run_conversation envelope test; the WHAT-restating comment shrinks to the WHY.
2026-09-09 12:20:04 +05:30
Felipe Portavales 37f42713ef feat(loop): export {turn_id, current_turn_user_idx} on every result envelope
Hosts that settle their own transcript by index (hermes-webui) cannot prove which
row of result["messages"] is the current user turn once this loop rewrote history
(alternation repair, compaction, post-turn micro-compaction): the instance-side
_persist_user_message_idx predates those rewrites, and a text match relabels an
identical historical prompt and claims its old answer. Only the producer can
assert the coordinate against the exact list it returns.

run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair
through export_current_turn_boundary on every envelope that leaves the loop
(success, partial/error, interrupt, retry-exhausted, tool-limit, preflight
timeout, codex runtime), computed on the final messages after finalize_turn and
micro-compaction. The pair is exported only when the addressed row is this turn's
user message verbatim (reanchor's last-match rule); a rewritten row exports
nothing so hosts fail closed. The final index is mirrored into
_persist_user_message_idx for the persist override.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
Felipe Portavales fe21a4d2f0 fix(loop): re-anchor current_turn_user_idx after the alternation repair merges rows
prepare_iteration() runs repair_message_sequence_with_cursor() before each API
call; the repair merges adjacent user rows in place (after a compaction, the
role=user summary sits next to the protected first user message). The loop's
current_turn_user_idx was recorded at turn start, so after a merge it points
past the current user row: the per-turn context injection (prefetch/plugin
context) silently misses it, and hosts that settle the transcript by this index
(hermes-webui) write the current user turn to the FRONT of the context —
rewriting the prompt's leading messages every turn (0% prefix-cache hits at
200K+ tokens, ~100 s re-prefill per turn) and duplicating the user's question.

The in-loop compression restart path already re-anchors; do the same after a
repair that changed the list: reanchor_current_turn_user_idx (last user row
carrying this turn's text), return the index through the IterationPrep verdict
so the loop state picks it up, and mirror it into agent._persist_user_message_idx,
which hosts read when the result carries no index. The new phase parameters
default to None so direct callers keep their signature.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
kshitijk4poor b7ac3ba1cd test(background-review): assert the behaviour, not the sentinel
The compaction-refresh and empty-surface tests pinned
_tool_snapshot_generation == _FROZEN_TOOL_SNAPSHOT_GENERATION next to
the behavioural assertion (refresh returns set(), tools unchanged). The
behaviour is the contract; the constant is the mechanism.
2026-09-09 10:32:49 +05:30
kshitijk4poor 586a7831d2 refactor(background-review): collapse the tool-surface copy to the agent_init shape; 2 tests
agent.tools is always a list (agent_init assigns it from
get_tool_definitions) and every entry is a well-formed function schema,
so the isinstance ladder over parent/entry/function/name guarded shapes
that cannot reach this helper. Use the same two lines agent_init uses;
`or []` keeps the empty-surface contract from the previous commit.
Docstring cut to the WHY (the between-turn refresh note described the
other guard). Tests trimmed to the two invariants: inherited tools
survive the compaction-boundary refresh (deep-copy isolation folded in),
and an empty parent surface is copied and frozen. Literal sentinel
asserts replaced with the constant.
2026-09-09 10:32:49 +05:30
0xAlyDev 06e59f2815 fix(background-review): inherit and freeze empty parent tools list for cache parity (#103579)
Copy and freeze review_agent._tool_snapshot_generation even when parent.tools is an empty list ([]). Previously, the truthiness check (\
ot parent_tools\) caused an empty parent tool surface to be skipped, allowing newly available late MCP or plugin tools to be retained on the review fork and leaving its snapshot generation unfrozen. This broke the byte-parity contract when no tools were active on the parent.

Returning early only when parent_tools is not an instance of list or tuple guarantees that an empty tool snapshot is faithfully inherited and frozen. Adds dedicated regression test test_unrouted_review_fork_inherits_empty_tool_surface.
2026-09-09 10:32:49 +05:30
0xAlyDev 0d72e07e67 fix(background-review): freeze review fork tool snapshot generation against compaction refresh
Freezes review_agent._tool_snapshot_generation to _FROZEN_TOOL_SNAPSHOT_GENERATION
(2_147_483_647) when inheriting the parent tool surface for same-model cache parity.

When in-place compaction boundaries trigger refresh_agent_mcp_tools(content_aware=True),
the staleness guard in _publish_tool_snapshot refuses the rebuild (snapshot_generation < published_gen),
preventing agent.tools from being reconstructed from the raw registry and preserving
inherited memory-provider and late tools across compaction boundaries (#103579).

Adds unit regression test verifying tool preservation across content_aware refresh.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-09 10:32:49 +05:30
0xAlyDev 6a97436e29 fix(agent): inherit parent's full tool surface on review fork for cache parity (#103579)
Ensure unrouted background_review forks inherit the parent's full advertised
tools[] surface. Without this, skip_memory=True caused memory-provider tools
(e.g. fact_store/fact_feedback) and dynamically injected plugin/late MCP tools
to be omitted from the fork's tools array, breaking byte-exact prefix-cache parity
and incurring full cold-read costs on providers where tools are part of the cache key.
Inheriting the full parent tools array preserves complete prefix cache parity
while execution dispatch remains strictly bounded by the thread tool whitelist.
2026-09-09 10:32:49 +05:30
Teknium 19cd839d54 fix(compression): keep lean tails lean after auxiliary feasibility
Lowering the session trigger must not replace the window-relative lean
selection budget with threshold times target_ratio. Invalidate the lean
cache through the existing property while preserving explicit legacy and
external-engine fallback behavior.

Narrow adaptation of the aux-sync diagnosis and invariants in #93576,
without adding a required recalibration method to context engines.
Related: #95681, #93576

Co-authored-by: Turgut Kural <58116817+TurgutKural@users.noreply.github.com>
2026-09-08 13:37:01 -07:00
Teknium 9d661c2c92 fix(prompt): keep memory guidance within available tools 2026-09-08 13:36:05 -07:00
Teknium 73ddf0672c test: keep attempt-cap probe on provider-confirmed fallback 2026-09-07 14:11:41 -07:00
Teknium ab98a92a45 fix(notifications): report applied skill batch operations
Use successful applied result records rather than requested operations, and keep staged writes silent. Include legacy delete/write messages.

Fixes #104506
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 08:14:14 -07:00
Teknium 7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium 0746484b14 test: assert skipped candidates do not compound cooldown 2026-09-07 07:08:25 -07:00
Teknium fb2f66f586 fix: describe remaining retry eligibility without promising recovery 2026-09-07 07:08:25 -07:00
fangliquanflq a4f2e42fbe test(agent): update custom runtime mock contract 2026-09-07 07:06:58 -07:00
Teknium 27f32bd50b test: exercise output-cap removal across native and child surfaces 2026-09-07 06:15:43 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 8d4b7414a0 test(recovery): distinguish the profile home from the full backup root 2026-09-07 06:06:20 -07:00
liuhao1024 342967058c fix(agent): interpolate the live backups dir into corruption recovery guidance
The corrupt-cause recovery guidance hardcoded `~/.hermes/backups/` while
every other path in the same message follows the active HERMES_HOME
(`{db_path}` is already interpolated). A custom-home or named-profile
deployment was told to restore from a directory that may not exist at all,
mid data-loss incident. Both sites (turn-completion explainer and gateway
startup broadcast) now interpolate `<hermes_root>/backups` via
get_default_hermes_root(), matching hermes_cli/backup.py's real backup
location.

Fixes #104250
2026-09-07 06:06:20 -07:00
Teknium ed3b9920ad test: model token flushes in the persistence race fixture 2026-09-07 06:02:22 -07:00
Teknium 7874ef9f62 test: keep title workers out of persistence fixtures 2026-09-07 02:30:04 -07:00
sylvainCDA 693641aa8b fix(chat-completions): strip name from tool-result messages for strict providers
The Chat Completions schema has no `name` field on `role: tool` (only on
the long-removed `role: function`), but Hermes carries the tool name over
onto the result message. Permissive providers ignore it; strict ones
(aki.io) reject the whole payload with `contains item with unknown key
name`, which breaks every tool call in the session.

Follow-up to the review on #51365:

- The strip now goes through the copy-on-write `mutable_msg()` path in
  `convert_messages()`, preserving the identity/copy-on-write contract
  instead of mutating `msg` in place.
- `handle_max_iterations()` hand-builds its summary payload and calls
  `chat.completions.create()` directly, bypassing the transport, so it
  leaked `name` even with the transport fixed. It now mirrors the same
  role-qualified removal, next to the existing tool_name/codex_*/timestamp
  strips.

The removal is role-qualified: `name` stays on user/assistant messages,
where it is schema-valid.

Regression coverage on both paths; both tests fail without the fix.
2026-09-07 02:50:45 +05:30
Teknium 0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
GodsBoy 5da6dcda5a test(compression): align checkpoint fixture and contributor mapping 2026-09-06 09:09:00 -07:00
Benjamin Brumbaugh cd71ee0708 fix(compression): defer local preflight after native checkpoint
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.

Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.

Fixes #100611
2026-09-06 09:09:00 -07:00
kshitijk4poor 8fd3e08ab9 test(agent): pin the emit-vs-buffer gate from both sides
The simplify reviewers flagged that every cooldown parametrize row was
> 60s, so the emit branch was never asserted against its buffered
complement. _drive_once now records which surface the status went
through, and a 30s row asserts short cooldowns keep the buffered line.
Mutation-checked: flipping the threshold fails the four emit rows.
2026-09-06 21:37:47 +05:30
kshitijk4poor 44e52e7524 fix(agent): normalize zero-cooldown semantics, dedupe test driver
/simplify-code pass on the Retry-After salvage stack:

- compute_error_backoff now decides "no usable cooldown" exactly once:
  a parsed 0.0 (retry-after: 0, or an HTTP-date in the past, which the
  shared parser clamps to 0) is treated as absent instead of falling
  through an accidental falsy check — prevents a hot-loop retry and
  keeps the sentinel semantics uniform (is None / is not None at all
  four sites).
- Comment corrected: the sibling unwrap lives in
  extract_api_error_context, not _extract_rate_limit_context.
- Test driver deduped: one _retryable_error factory + one _drive_once
  shared by the 429 and 524 tests; added the over-cap (3600 → 600) and
  no-cooldown fallback rows requested in the #103722 review.

Mutation checks: nested-unwrap neutralized → nested row red; cap
removed → over-cap row red; restored → all 298 green.
2026-09-06 21:37:47 +05:30
kshitijk4poor c6ef075613 fix(agent): parse nested retry_after bodies and emit long 5xx cooldowns
Salvage follow-ups on #88236 (krunkosaurus):

- Some providers nest the cooldown as body["error"]["retry_after"] (the
  same unwrap _extract_rate_limit_context already uses); only the top-level
  shape was read, so those errors silently fell back to jittered backoff.
- A 5xx Retry-After can reach the 600s cap; that wait was buffered
  (replayed only on terminal failure), leaving the user silent for
  minutes. Long provider cooldowns now emit immediately, mirroring the
  zai_coding_overload_long path. Jittered waits keep the old buffering.

Test widened with the nested-body parametrize row; proven red when the
unwrap is neutralized.
2026-09-06 21:37:47 +05:30
Mauvis Ledford a227484906 fix(agent): honor Retry-After on retryable 5xx
Retryable server errors can carry provider cooldowns just like 429 responses. Parse Retry-After from response headers or structured error bodies before falling back to jittered backoff, and cover HTTP 524 behavior with runtime regression tests.
2026-09-06 21:37:47 +05:30
Teknium 335ecf9f4a test: the two remaining native-wire contracts select nous.anthropic_wire=native explicitly
test_nous_anthropic_fallback_uses_the_messages_wire and
test_nous_child_rederives_api_mode_from_model describe the native wire, which
is now opt-in; select it in the test the same way the wire-contract suite
does, so both keep guarding the flip-back.
2026-09-06 05:55:18 -07:00
kshitijk4poor 0af92098f2 test(agent): split the Kimi lookalike-host case into its own test 2026-09-05 15:54:14 +05:30
kshitijk4poor 562383ad31 test(agent): Kimi Code fallback lands on anthropic_messages; lookalike host stays on chat_completions
Red on origin/main (chat_completions), green with the fix.
2026-09-05 15:54:14 +05:30
kshitijk4poor f5832f81ed test(streaming): drive the Relay finalizer through its collector
The chat_completions finalizer now reads collector-observed chunks, not the
consumer loop, so the test feeds the captured on_chunk before calling it —
the same ordering Relay guarantees.
2026-09-05 10:54:19 +05:30
Teknium 8924b3aed9 fix(skills): lead the lesson-layer contract with the primary purpose — how to do the task, to the user's specifications 2026-09-04 08:15:22 -07:00
Teknium c240e65399 fix(skills): self-improvement writes lessons, not incident logs
The skill review fork, the combined memory+skill review, and the curator's
consolidation pass all described references/ as the place for
"session-specific detail", and the curator's demote step said to move a
sibling's file under the umbrella. Followed literally over months that
produced one dev skill with a 100k SKILL.md dense in PR numbers and 443
one-per-session reference files, plus five sibling skills restating the
same rules and the repo's AGENTS.md.

The three prompts now share one shape contract: an entry is an imperative
rule plus one clause of why, stated once; no PR/issue numbers, dates, or
quoted chat as content; references/ is a small topical set extended in
place, never a per-session file; skills do not restate always-loaded
context. Consolidation is defined as distilling, and copying a sibling
verbatim under references/ is named as the failure. skill_manage's schema
carries the one-sentence version.

Two advisory linter rules make the shape visible in the tool result the
moment it starts to drift: incident-log-shape (PR/issue-number density in
prose) and references-sprawl (>60 reference files), the latter also run
on references/ writes. On the real before/after: the old skill trips both,
the consolidated one trips neither.
2026-09-04 08:15:22 -07:00
Teknium 342879f267 fix(compat-fallout): repoint 7 test files that still imported hermes_state/run_agent names from the old facade 2026-09-03 16:35:01 -07:00
Teknium da96762ef8 fix(compat-fallout): repoint 8 test files off dropped facade names (run_agent._DB_PERSISTED_MARKER/_EPHEMERAL_SCAFFOLDING_FLAGS/_is_ephemeral_scaffolding, cli.AIAgent, main._make_tui_argv/_print_curator_recent_run_notice/_format_time_ago, model_switch._collect_authed_provider_slugs, acp has_provider shim); arg_coercion test imports model_tools to populate registry 2026-09-03 15:50:55 -07:00