Commit Graph

1154 Commits

Author SHA1 Message Date
Teknium 73ddf0672c test: keep attempt-cap probe on provider-confirmed fallback 2026-09-07 14:11:41 -07:00
Teknium ab98a92a45 fix(notifications): report applied skill batch operations
Use successful applied result records rather than requested operations, and keep staged writes silent. Include legacy delete/write messages.

Fixes #104506
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 08:14:14 -07:00
Teknium 7a5fc1b2a9 fix: remove automatic session JSON snapshots 2026-09-07 08:08:41 -07:00
Teknium 0746484b14 test: assert skipped candidates do not compound cooldown 2026-09-07 07:08:25 -07:00
Teknium fb2f66f586 fix: describe remaining retry eligibility without promising recovery 2026-09-07 07:08:25 -07:00
fangliquanflq a4f2e42fbe test(agent): update custom runtime mock contract 2026-09-07 07:06:58 -07:00
Teknium 27f32bd50b test: exercise output-cap removal across native and child surfaces 2026-09-07 06:15:43 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium 8d4b7414a0 test(recovery): distinguish the profile home from the full backup root 2026-09-07 06:06:20 -07:00
liuhao1024 342967058c fix(agent): interpolate the live backups dir into corruption recovery guidance
The corrupt-cause recovery guidance hardcoded `~/.hermes/backups/` while
every other path in the same message follows the active HERMES_HOME
(`{db_path}` is already interpolated). A custom-home or named-profile
deployment was told to restore from a directory that may not exist at all,
mid data-loss incident. Both sites (turn-completion explainer and gateway
startup broadcast) now interpolate `<hermes_root>/backups` via
get_default_hermes_root(), matching hermes_cli/backup.py's real backup
location.

Fixes #104250
2026-09-07 06:06:20 -07:00
Teknium ed3b9920ad test: model token flushes in the persistence race fixture 2026-09-07 06:02:22 -07:00
Teknium 7874ef9f62 test: keep title workers out of persistence fixtures 2026-09-07 02:30:04 -07:00
sylvainCDA 693641aa8b fix(chat-completions): strip name from tool-result messages for strict providers
The Chat Completions schema has no `name` field on `role: tool` (only on
the long-removed `role: function`), but Hermes carries the tool name over
onto the result message. Permissive providers ignore it; strict ones
(aki.io) reject the whole payload with `contains item with unknown key
name`, which breaks every tool call in the session.

Follow-up to the review on #51365:

- The strip now goes through the copy-on-write `mutable_msg()` path in
  `convert_messages()`, preserving the identity/copy-on-write contract
  instead of mutating `msg` in place.
- `handle_max_iterations()` hand-builds its summary payload and calls
  `chat.completions.create()` directly, bypassing the transport, so it
  leaked `name` even with the transport fixed. It now mirrors the same
  role-qualified removal, next to the existing tool_name/codex_*/timestamp
  strips.

The removal is role-qualified: `name` stays on user/assistant messages,
where it is schema-valid.

Regression coverage on both paths; both tests fail without the fix.
2026-09-07 02:50:45 +05:30
Teknium 0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
GodsBoy 5da6dcda5a test(compression): align checkpoint fixture and contributor mapping 2026-09-06 09:09:00 -07:00
Benjamin Brumbaugh cd71ee0708 fix(compression): defer local preflight after native checkpoint
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.

Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.

Fixes #100611
2026-09-06 09:09:00 -07:00
kshitijk4poor 8fd3e08ab9 test(agent): pin the emit-vs-buffer gate from both sides
The simplify reviewers flagged that every cooldown parametrize row was
> 60s, so the emit branch was never asserted against its buffered
complement. _drive_once now records which surface the status went
through, and a 30s row asserts short cooldowns keep the buffered line.
Mutation-checked: flipping the threshold fails the four emit rows.
2026-09-06 21:37:47 +05:30
kshitijk4poor 44e52e7524 fix(agent): normalize zero-cooldown semantics, dedupe test driver
/simplify-code pass on the Retry-After salvage stack:

- compute_error_backoff now decides "no usable cooldown" exactly once:
  a parsed 0.0 (retry-after: 0, or an HTTP-date in the past, which the
  shared parser clamps to 0) is treated as absent instead of falling
  through an accidental falsy check — prevents a hot-loop retry and
  keeps the sentinel semantics uniform (is None / is not None at all
  four sites).
- Comment corrected: the sibling unwrap lives in
  extract_api_error_context, not _extract_rate_limit_context.
- Test driver deduped: one _retryable_error factory + one _drive_once
  shared by the 429 and 524 tests; added the over-cap (3600 → 600) and
  no-cooldown fallback rows requested in the #103722 review.

Mutation checks: nested-unwrap neutralized → nested row red; cap
removed → over-cap row red; restored → all 298 green.
2026-09-06 21:37:47 +05:30
kshitijk4poor c6ef075613 fix(agent): parse nested retry_after bodies and emit long 5xx cooldowns
Salvage follow-ups on #88236 (krunkosaurus):

- Some providers nest the cooldown as body["error"]["retry_after"] (the
  same unwrap _extract_rate_limit_context already uses); only the top-level
  shape was read, so those errors silently fell back to jittered backoff.
- A 5xx Retry-After can reach the 600s cap; that wait was buffered
  (replayed only on terminal failure), leaving the user silent for
  minutes. Long provider cooldowns now emit immediately, mirroring the
  zai_coding_overload_long path. Jittered waits keep the old buffering.

Test widened with the nested-body parametrize row; proven red when the
unwrap is neutralized.
2026-09-06 21:37:47 +05:30
Mauvis Ledford a227484906 fix(agent): honor Retry-After on retryable 5xx
Retryable server errors can carry provider cooldowns just like 429 responses. Parse Retry-After from response headers or structured error bodies before falling back to jittered backoff, and cover HTTP 524 behavior with runtime regression tests.
2026-09-06 21:37:47 +05:30
Teknium 335ecf9f4a test: the two remaining native-wire contracts select nous.anthropic_wire=native explicitly
test_nous_anthropic_fallback_uses_the_messages_wire and
test_nous_child_rederives_api_mode_from_model describe the native wire, which
is now opt-in; select it in the test the same way the wire-contract suite
does, so both keep guarding the flip-back.
2026-09-06 05:55:18 -07:00
kshitijk4poor 0af92098f2 test(agent): split the Kimi lookalike-host case into its own test 2026-09-05 15:54:14 +05:30
kshitijk4poor 562383ad31 test(agent): Kimi Code fallback lands on anthropic_messages; lookalike host stays on chat_completions
Red on origin/main (chat_completions), green with the fix.
2026-09-05 15:54:14 +05:30
kshitijk4poor f5832f81ed test(streaming): drive the Relay finalizer through its collector
The chat_completions finalizer now reads collector-observed chunks, not the
consumer loop, so the test feeds the captured on_chunk before calling it —
the same ordering Relay guarantees.
2026-09-05 10:54:19 +05:30
Teknium 8924b3aed9 fix(skills): lead the lesson-layer contract with the primary purpose — how to do the task, to the user's specifications 2026-09-04 08:15:22 -07:00
Teknium c240e65399 fix(skills): self-improvement writes lessons, not incident logs
The skill review fork, the combined memory+skill review, and the curator's
consolidation pass all described references/ as the place for
"session-specific detail", and the curator's demote step said to move a
sibling's file under the umbrella. Followed literally over months that
produced one dev skill with a 100k SKILL.md dense in PR numbers and 443
one-per-session reference files, plus five sibling skills restating the
same rules and the repo's AGENTS.md.

The three prompts now share one shape contract: an entry is an imperative
rule plus one clause of why, stated once; no PR/issue numbers, dates, or
quoted chat as content; references/ is a small topical set extended in
place, never a per-session file; skills do not restate always-loaded
context. Consolidation is defined as distilling, and copying a sibling
verbatim under references/ is named as the failure. skill_manage's schema
carries the one-sentence version.

Two advisory linter rules make the shape visible in the tool result the
moment it starts to drift: incident-log-shape (PR/issue-number density in
prose) and references-sprawl (>60 reference files), the latter also run
on references/ writes. On the real before/after: the old skill trips both,
the consolidated one trips neither.
2026-09-04 08:15:22 -07:00
Teknium 342879f267 fix(compat-fallout): repoint 7 test files that still imported hermes_state/run_agent names from the old facade 2026-09-03 16:35:01 -07:00
Teknium da96762ef8 fix(compat-fallout): repoint 8 test files off dropped facade names (run_agent._DB_PERSISTED_MARKER/_EPHEMERAL_SCAFFOLDING_FLAGS/_is_ephemeral_scaffolding, cli.AIAgent, main._make_tui_argv/_print_curator_recent_run_notice/_format_time_ago, model_switch._collect_authed_provider_slugs, acp has_provider shim); arg_coercion test imports model_tools to populate registry 2026-09-03 15:50:55 -07:00
Teknium b92308b1d2 simplify(compat): tools-A — repoint 4 stale docstring references (tools.approval.*, tools.transcription_tools.*) to the defining modules 2026-09-03 14:37:14 -07:00
Teknium 4fde117f4b simplify(compat): tests — repoint 638 web_server.<name> references (321 attr, 217 monkeypatch/patch.object, 60 from-imports, 40 patch() strings) across 65 test files to the owning modules 2026-09-03 14:21:52 -07:00
Teknium 7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium 14791b4d4e simplify(compat): approval — drop 43 facade re-exports + _command_detection_variants late-bind seam, repoint 30 callers + 46 test files
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.
2026-09-03 13:49:57 -07:00
Teknium 7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium 53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium a9b0dd6742 simplify(compat): cron scheduler — drop 82 re-exports, repoint 8 callers + 33 test files
cron/scheduler.py no longer re-exports the split modules (scheduler_delivery /
_script / _prompt / _preflight); it imports only the 19 names it calls itself
(bottom-of-file, E402 kept for the import cycle). Dropped the shim-only
`import shutil` and the F401 note on windows_hide_flags (still used by
scheduler.py). Split modules now call same-module helpers directly, reach
sibling split modules via late-bound module refs (_delivery/_script/_preflight)
next to _sched, and import windows_hide_flags themselves; origin-resident names
(load_config, Path, _SCRIPT_TIMEOUT, heartbeat_run_claim, ...) still go through
_sched. Callers/tests import + patch the defining module.
2026-09-03 13:38:36 -07:00
Teknium fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium eb8a30cc2f simplify(compat): skills_sync/skill_manager/kanban/computer_use — drop 37 re-exports/aliases, repoint 8 callers + 9 test files 2026-09-03 13:27:01 -07:00
Teknium d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00
Teknium 89fbd5d4d3 simplify(compat): conversation_loop — drop 9 re-exports, repoint 7 callers, 35 test sites 2026-09-03 13:12:23 -07:00
Teknium 022785a541 Merge origin/main (63279301bc): reasoning-mandatory 400 recovery folded into turn_recovery/error_classifier/models_reasoning_caps 2026-09-03 05:36:21 -07:00
Teknium 63279301bc fix(reasoning): retry after a mandatory-reasoning 400 resends the user's own effort
The retry must land on the same provider cache key as every prior request
in the session. Discard only the one-shot continuation disable and send
agent.reasoning_config verbatim; a config that is itself a disable is
omitted (that session never sent anything else, so nothing warm is lost).

Live: user effort=high, ephemeral disable → 400 → retry carries
{enabled: true, effort: high}.
2026-09-03 05:20:11 -07:00
Teknium f6bd1633f7 fix(reasoning): GLM-5.3 on Nous/OpenRouter no longer 400s when thinking is disabled
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.

- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
  no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
  queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
  (configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
  background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
  mandatory (parity with the Nous profile).

Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
2026-09-03 05:20:11 -07:00
Teknium 3bff2a581d Merge origin/main (562ee8ab76): fold usage-less empty-response guard, unterminated tool-call strip, skills-first memory guidance into the simplified modules 2026-09-03 04:31:46 -07:00
Teknium 7b72fd1247 fix(agent): stream cut mid tool-call markup no longer persists as assistant prose (#101899)
GLM-style models serialize tool calls as XML in the text channel; when the
stream drops mid-serialization with finish_reason=stop, the orphan
<arg_key>/<arg_value> fragment (or a bare unclosed <tool_call> opener)
matched neither the complete-block stripper nor the partial-stream guard
and was stored and displayed as ordinary assistant content.

strip_think_blocks (storage boundary) and the CLI display copy now strip
an unterminated block-boundary tool-call opener, or any line carrying
stray argument markup, to end of text. The response then reads as empty
and flows through the existing empty-retry path. Complete blocks and
inline prose mentions are unchanged.
2026-09-03 03:42:15 -07:00
Teknium c3e9defd18 fix(agent): trim usage-less empty fix to guard + observability
Drop the loop-side '(empty)' rewrite (the turn-completion explainer already
owns that at delivery, and gateway/desktop match on the sentinel) and the
extra token-count persistence. Keeps: usage-absent empty streaks arm the
deterministic fast-fail after two attempts with no content or reasoning,
and every completed API call logs even when the provider omits usage
(#101898).
2026-09-03 03:41:52 -07:00
fangliquan 3755dca7d6 fix(agent): handle usage-less empty responses 2026-09-03 03:41:52 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
kshitijk4poor 58a4a11727 fix(compression): release the reclamation no-op dedup key on every rearm reset
Follow-up to the salvaged #101894 (@jwilson411). The over-threshold
"reclamation did not run" warning is deduped on (reason, rearm mark), and
the key was only cleared when a prune committed. Every other path that
zeroes the rearm mark — compress(), on_session_reset/on_session_end,
bind_session_state, update_model — left the key in place, so a lockout
that warned at rearm=0, then a full compaction, then the same lockout
again was silent, contradicting the helper's own "warns again" contract
(and leaking the key across sessions on a rebound compressor).

- ContextCompressor._reset_proactive_prune_rearm(): one helper for the
  five rearm-to-zero sites; clears the dedup key alongside the mark.
- _warn_reclamation_no_op(): dropping back under threshold releases the
  key (mirrors _clear_context_overflow_warn semantics on the agent side).
- test_proactive_prune_loop_wiring: the attempts_exhausted fixture now
  models the only state the real engine can produce for that branch
  (should_compress() is should_compress_info()[0]) — budget spent
  (max_compression_attempts=0) with the engine saying RUN, instead of
  should_compress=False paired with (True, None).
- Two guards: lockout warns again after a rearm reset; dropping under
  threshold releases the key. Both fail with the clears removed.
2026-09-03 12:23:55 +05:30