Commit Graph

2958 Commits

Author SHA1 Message Date
Jack Lau f57209bc9f fix(agent): carry the ambiguity of Anthropic's 'out of extra usage' 400 through classification, cooldown, and terminal surfaces
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:

- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
  status-less paths now attach error_context {billing_unverified,
  possible_content_filter}. Reason stays FailoverReason.billing (rotation +
  fallback remain the right recovery either way); ClassifiedError grows a
  billing_unverified property.

- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
  unverified billing exhaustion gets the short transient cooldown instead of
  the one-hour bench, regardless of pool size: a content-filter rejection
  leaves the credential healthy and fails identically on every key, and the
  hour-long sole-credential latch is what replayed the stored error and made
  real fixes look ineffective. A true 402 keeps the full bench. The marker
  persists with the entry so a restart cannot upgrade it back to a bench.

- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
  threads billing_unverified and hands the pool 'billing_unverified' as the
  persisted failure_reason.

- agent/conversation_loop.py: the fallback-switch status, max-retries status,
  terminal label, and both structured terminal results hedge when the verdict
  is unverified. New _billing_terminal_label + _billing_failure_result build
  the returned terminal response in one place; the result dict now carries
  billing_unverified and the billing_block gains 'unverified': true. The
  confirmed-billing path (a real 402 or an API-key credit depletion) keeps
  the original assertive wording, so the caveat no longer dilutes it.

Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.

Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
2026-08-14 21:54:56 -07:00
Jack Lau 6fbbe18be8 fix(agent): reword SKILLS_GUIDANCE trigger and stop mislabelling its 400 as billing
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.

Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).

Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:

- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
  reporter verified returns 200. Meaning, the skill_manage reference, and the
  ## Skill Safety Rule block are all preserved. The reword is empirically
  validated rather than understood, so a comment records the bisect and warns
  that any rewrite must be re-verified against an OAuth token, not an API key.

- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
  longer asserts exhaustion as fact. It hedges the opening line, names the
  content-filter alternative, and gives the operator a way to tell the two apart
  (if the usage page still shows quota, suspect a content rejection). It also
  points at `hermes auth reset anthropic`, because the credential exhaustion
  latch replays the stored error for ~60 min without issuing a request — which
  makes a real fix look like it did not work.

- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
  not an API key, despite auth_type="api_key". It stays in api_key_env_vars
  because that tuple doubles as the credential-discovery list; removing it would
  stop Hermes finding a `claude setup-token` credential at all.

Docs updated to match the reworded prompt.

Fixes #82154
2026-08-14 21:54:56 -07:00
dhruv kejriwal ffaa63f887 fix(compressor): never commit a compression that grows the transcript (in-place path)
The gateway rotation guard (#83339) only protects the rotate path, but
in-place compaction commits inside compress_context() via
archive_and_compact — before the gateway can inspect the result. Add the
anti-growth check at the commit site so both paths are covered: a
compression whose rough output exceeds its input is a strict no-op
(original transcript kept durable, session identity untouched).

Covers the observed failure where session hygiene persisted 426 -> 426
messages and ~379K -> ~688K tokens.
2026-08-15 10:22:47 +05:30
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Teknium 8dc5608a78 fix(compression): adopt live continuation tip at flush across multi-hop chains
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes #82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-14 21:39:44 -07:00
Teknium 8369170fea fix(agent): only strip _moa_prepared_request when the live client is not the MoA facade
The unconditional pop from the previous commit also stripped the key when
agent.client was still the real MoA facade, forcing the facade to re-prepare
from scratch — a duplicate reference fan-out per turn (caught by
test_moa_virtual_provider_aggregator_is_actor). Gate the defensive strip on
the same prepare()-capability probe used at the injection site so the
handshake survives on the facade while a swapped-in native client is still
protected.
2026-08-14 21:36:41 -07:00
RelaxJonh 6bf3d39499 fix(agent): strip _moa_prepared_request before dispatching to native client
After a client replacement (credential rotation, dead-connection cleanup,
or fallback+restore), agent.client may become a native OpenAI client
while agent.provider stays "moa".  The _moa_prepared_request key was
passed through to the native SDK, causing TypeError on every turn.

Pop the key at the dispatch point (chat_completion_helpers.py:509).
The MoAClient facade already handles a missing key by falling through
to its normal resolution path.

Closes #78382
2026-08-14 21:36:41 -07:00
Drexuxux ab879a1e22 fix(agent): stop the MoA prepared request reaching a swapped-in native client
`_moa_prepared_request` is a private handshake between the conversation
loop and MoAChatCompletions.create. It is attached whenever
agent.provider == "moa", on the assumption that agent.client is still the
in-process MoA facade.

Credential rotation, provider fallback and dead-connection cleanup all
rebuild agent.client from _client_kwargs between attempts, and
pending_moa_prepared_request deliberately carries a prepared request
across exactly that boundary. The rebuilt client is a native OpenAI
client while provider stays "moa", so the key reaches an SDK that has
never heard of it:

    TypeError: Completions.create() got an unexpected keyword argument
    '_moa_prepared_request'

That error is non-retryable, so every remaining turn on the session
fails. Both dispatch paths are affected: the non-streaming one calls
agent.client directly, and _create_request_openai_client returns
agent.client unchanged for provider "moa".

Re-check the live client at the point the key is attached, which covers
both paths at once. When the facade is gone, send the prepared prompt
without the handshake and log the downgrade.
2026-08-14 21:36:41 -07:00
牧濑红莉栖(BOT) 24eae1fa15 fix(agent): preserve MoA facade when rebuilding primary client (stream retry, rotation, fallback+restore)
When agent.provider == "moa", the MoAClient facade *is* the client - there is
no real OpenAI wire endpoint behind the moa://local placeholder. Client rebuilds
(_replace_primary_openai_client: stream-retry pool cleanup, credential rotation,
dead-connection cleanup, fallback+restore) go through create_openai_client and
produce a native OpenAI client while provider stays "moa". The next primary
call then either raises a `_moa_prepared_request` TypeError (#78382) or, when
_client_kwargs carry an unrelated relay base_url, leaks the request to a foreign
gateway (observed as HTTP 503 "group ... no available channel" from an
unrelated new-api relay right after an aggregator empty-stream retry).

Fix: in create_openai_client, when provider is "moa", return
build_moa_facade(agent, model) instead of a native client. This covers every
rebuild entry point. The three already-fixed call sites
(restore_primary_runtime, try_recover_primary_transport, switch_model) assign
the facade directly and do not go through create_openai_client, so they are
unaffected.

Closes #78382
2026-08-14 21:36:41 -07:00
SHL0MS bec7df1bba fix(anthropic): coerce blank system text blocks at extraction (#70909)
Residual from PR #70910 after #77509 landed the message-list scrub: a
whitespace-only system content block carrying a cache_control marker
still reached the wire and 400'd the whole request ("text content
blocks must contain non-whitespace text"), wedging the session on every
retry. The block cannot be dropped (it carries the cache breakpoint),
so coerce its text to the shared non-whitespace placeholder when
extracting the system param, copying the block so caller message dicts
are never mutated.

Adds SHL0MS's request-level regression suite from #70910; four of its
five cases already pass on main via #77509 — the system-block case
fails without this fix.
2026-08-14 21:27:05 -07:00
Turgut Kural 4e60771ab7 fix(transport): scope empty tool_calls comment to transport-layer coverage 2026-08-14 21:25:48 -07:00
Turgut Kural c464001da8 fix(transport): strip empty/null tool_calls on assistant messages
Strict OpenAI-compatible providers (onerouter / Qwen, DeepSeek v4) reject
an assistant message carrying tool_calls: [] (or null) with HTTP 400
'Empty tool_calls is not supported in message.'

The pre-API sanitizer in agent_runtime_helpers.sanitize_api_messages already
drops these on the conversation_loop path, but auxiliary / custom-provider
routes that bypass that sanitizer can still reach the wire with an invalid
empty array and abort the whole session (non-retryable 400).

Normalize at the transport layer too: detect an empty-list / null
tool_calls on assistant messages, strip the key on the per-call copy (never
mutate the stored history), and keep real tool_calls untouched. Includes
unit tests covering empty-list, null, real-call preservation, mixed batches,
user-role non-mutation, copy-on-write, and cross-provider parity.

Follow-up to #58755.
2026-08-14 21:25:48 -07:00
webtecnica f316f7d086 fix(session): drop empty tool_calls in repair_message_sequence (#77921) 2026-08-14 21:25:48 -07:00
liuhao1024 b2453b5894 fix(sanitize): drop tool_calls key when dedup removes all calls
The dedup pass in sanitize_api_messages (introduced by #58327) can
produce an empty tool_calls array when all tool_call_ids in a message
are duplicates of earlier messages in a long conversation history.

DeepSeek v4 and newer OpenAI reject empty tool_calls with HTTP 400:
'Invalid messages[N].tool_calls: empty array'.

When kept_tcs is empty after dedup, drop the tool_calls key entirely
instead of writing tool_calls: [].

Fixes #64335
2026-08-14 21:25:48 -07:00
Ufonik 54aad0d670 fix(tools): refresh activity heartbeat while a tool call is in flight (#84491)
The gateway turn-inactivity watchdog (gateway/run.py::_watch_gateway_turn_inactivity)
abandons a turn once seconds_since_activity exceeds the inactivity timeout
(default 30 min). Activity was only stamped when a tool started and when it
completed, so a tool call that runs silently for 30+ minutes (quiet builds,
long pytest suites, large downloads, network waits with no output) froze the
clock and the watchdog hard-abandoned a turn that was still making progress,
reaping the tool's processes mid-execution (issue #84491).

Add a daemon-thread heartbeat inside _run_agent_tool_execution_middleware that
touches agent._touch_activity every 30s while the tool is in flight, until the
call returns. Both the sequential and concurrent execution paths funnel through
this single middleware, so one heartbeat covers every tool. The thread is
stopped in a try/finally so it always tears down even if execute() raises, and
is never started when a guardrail/authorization block short-circuits before
dispatch. A genuinely hung tool remains bounded by the tool layer's own
timeouts (terminal default 180s, concurrent batch deadline ~420s), so the
heartbeat only extends the turn's life while the call is legitimately running.

Verified by an independent reviewer (no security/logic defects); 5 unit/integration
tests pass on Python 3.12 (upstream CI). The 30-min gateway backstop remains
for turns whose agent loop itself stalls.
2026-08-14 21:20:03 -07:00
HexLab98 cc0d5ce7b7 fix(agent): bound hung inline API calls when abort cannot kill the socket
The keepalive httpx client uses read=None, and stranger-thread abort cannot close FDs, so a DeepSeek stall on the cron inline path waited until TCP died — hours past the 600s watchdog. Inject a per-call read timeout matching the stale budget, walk in-flight pool requests, and clear the socket timeout before shutdown without releasing the FD.
2026-08-14 21:19:55 -07:00
HexLab98 fc100f4b3b fix(compression): scan full window for handoffs before cross-session discard
A degenerate compress_end can hide an in-window handoff past the cut; the
#57835 guard then cleared a valid same-session _previous_summary (#83248).
2026-08-14 20:52:16 -07:00
Teknium bc5805c35f fix: compare base-URL hostnames, not substrings, in provider-identity checks
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:

- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
  URL containing the OpenRouter host substring (path segment, lookalike
  domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
  openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
  localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
  the nousresearch substring anywhere in the URL.

All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
2026-08-14 20:47:13 -07:00
Teknium 5d9e4aaaf2 fix: raise ProviderStreamError for choiceless error chunks + regression tests 2026-08-14 20:46:42 -07:00
AlexFucuson9 48cca664c6 fix: detect in-stream error chunks in SSE streaming path
Some OpenAI-compatible providers (DeepInfra, etc.) return validation
errors as in-stream SSE chunks: HTTP 200 with choices=None and
error_type/error_message in model_extra. The streaming loop silently
dropped these chunks, causing EmptyStreamError ("empty stream") and
pointless retries on the same bad request.

Fix: check for error_type/error_message on chunks with no choices
before skipping. When found, raise RuntimeError with the provider's
error message so the error classifier can properly handle it.

Fixes #65631
2026-08-14 20:46:42 -07:00
beiyesi b55dd04711 fix(streaming): adapt provider errors to relay 2026-08-14 16:40:02 -07:00
beiyesi 04bc5321c9 fix: preserve non-JSON provider stream errors 2026-08-14 16:40:02 -07:00
beiyesi 5cbc645d09 fix: require terminal signal for streamed errors 2026-08-14 16:40:02 -07:00
beiyesi 0fcebb29f4 fix streaming bare data error payloads 2026-08-14 16:40:02 -07:00
beiyesi 661a4c4f88 fix streaming provider error events 2026-08-14 16:40:02 -07:00
SHL0MS b21e0bd8c9 fix(providers): honor per-provider TLS on custom /models and pricing probes
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:

- the requests-based metadata/pricing probe
  (agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
  (hermes_cli/models.py::probe_api_models)

A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.

This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:

- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
  ssl_ca_cert before falling back to the env vars. Callers with no base_url
  keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
  passes it through open_credentialed_url, which gains an ssl_context seam on
  the cloned secure opener. Unmatched or public endpoints pass None and keep
  urllib's default policy.

Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
2026-08-14 16:37:36 -07:00
fangliquan f1025b2c00 fix(sessions): fence transcript writes with the turn-lease holder
Refresh-loss interrupt is cooperative, so a stalled writer could still flush after another process reclaimed the conversation. Carry the holder into append_message / append_messages_batch and reject the write in the same SQLite transaction when the lease row is missing, expired, or owned by someone else.
2026-08-15 03:22:16 +05:30
Teknium 518bc90e74 fix(agent): bound HERMES_HOME override wins over shared session-db home
The messaging gateway multiplexes profiles over ONE shared launch-home
state.db, binding the profile per turn via the HERMES_HOME ContextVar
(copy_context into the worker thread). _agent_home derived the home from
db_path unconditionally, so on that lane the launch home stomped the
correctly-bound profile — deterministically inverting the leak #86313
fixed (found by @kshitijk4poor's post-merge probe; @helix4u flagged the
plugin-metadata half).

- _agent_home: bound override wins; session_db home is the unbound-thread
  fallback
- _plugin_session_info: profile_name derives from _agent_home too
- full-prompt wiring regression (SOUL + skills + profile line on a bare
  thread with the bot's DB) — reverting any call-site wire fails it;
  multiplex, bare-thread, and plugin-metadata cases each pinned;
  sabotage-verified both new tests fail against the merged behavior
- skills LRU cap 8 -> 32 (key is now per-profile x platform)
2026-08-14 14:37:00 -07:00
Teknium 1f1b4d9947 fix(agent): scope SOUL.md load to the agent's own profile home (#50233)
load_soul_md resolved the home ambiently, so a build thread that lost the
HERMES_HOME ContextVar read the launch profile's SOUL.md into another
profile's prompt — same class as the skills-index leak fixed in #86313.
load_soul_md and build_context_files_prompt now accept home_override, and
build_system_prompt_parts passes the agent's own home (from session_db)
at both SOUL call sites. Ambient behavior unchanged when no override.
2026-08-14 14:15:42 -07:00
Jeff Mettel 126acbf219 fix(agent): stop doubling the profile path in the system-prompt hint
The named-profile branch of the "Active Hermes profile" hint built its
paths by appending `/profiles/{active_profile}` to `get_hermes_home()`.
But `_resolve_active_profile_name()` returns a non-default name *only*
when `get_hermes_home()` has already resolved under `<root>/profiles/`
— that is how it derives the name in the first place. Both scoping
mechanisms (a `HERMES_HOME=<root>/profiles/<name>` env var and the
multiplexer's `set_hermes_home_override` contextvar) satisfy that, so
the suffix always doubled.

The same branch used `get_hermes_home()` for the *default* profile's
data pointers, where the root was intended — placing them inside the
active profile.

On a real 4-profile install the hint rendered:

    reads and writes ~/.hermes/profiles/via/profiles/via/
    default profile's data lives at ~/.hermes/profiles/via/skills/

against actual paths of `~/.hermes/profiles/via/` and `~/.hermes/skills/`.

Use the session home directly as the profile home, and
`get_default_hermes_root()` for the root pointers. Every named-profile
session was shipping a prompt that named nonexistent directories and
mislabeled this profile's own skills/plugins/cron/memories as the
default profile's — the exact cross-profile confusion the hint and
`classify_cross_profile_target` exist to prevent.

The default-profile branch is unchanged.

Fixes #72894

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 14:11:57 -07:00
Teknium 30e9449403 fix(agent): keep legacy ambient home strings byte-identical when no agent home
The root-derivation for the profile-hint text now only engages when the
agent's own home is known; otherwise the line resolves exactly as before
(via this module's get_hermes_home), fixing
test_coding_prompt_preserves_legacy_workspace_order.
2026-08-14 13:56:58 -07:00
Teknium 8fd0f9c1d2 fix(agent): derive profile name from the hermes ROOT, not the ambient home
On a correctly bound profile session get_hermes_home() returns the profile
dir itself, so relative_to(home/'profiles') never matched and every profile
misreported as 'default' (with wrong paths in the profile hint text). Use
get_default_hermes_root() for both the name derivation and the default-data
root string. Adds regression tests for the bound-profile and root-home cases;
sabotage-verified the bound-profile test fails against the old resolution.
2026-08-14 13:56:58 -07:00
Teknium 2f25dec349 fix(agent): scope the skills index + active-profile line to the agent's OWN home
A bot profile's system prompt could list the DEFAULT profile's ~80
skills and print 'Active Hermes profile: default' — while the live
skills_list() correctly showed the bot's real (often empty) set. The
agent plans against that index, so a false inventory makes it claim
capabilities it doesn't have, waste context tokens, and lose trust.

Root cause (confirmed empirically): the skills-prompt builder and the
active-profile line resolve the home through get_hermes_home(), which
reads a HERMES_HOME ContextVar. ContextVars do NOT propagate into
threading.Thread, so an agent build running on a thread that didn't
bind the profile's home falls back to the launch (default) home and
builds default's index. A bare no-override thread builds default's
full 7621-char block; the same thread with the fix builds empty.

Fix: resolve the agent's OWN home from its dedicated _session_db.db_path
(ground truth, ContextVar-independent) and pass it explicitly:
- build_skills_system_prompt(skills_dir_override=...) scopes the index,
  the disk snapshot, and external-dir resolution to that home
- the active-profile line derives the profile name from the same home
Both fall back to ambient resolution when no db is present, so the CLI
and default-profile paths are unchanged.

Regression tests: an empty bot profile yields an empty skills block on
a bare thread even with ambient HERMES_HOME bound to a skills-rich
default; agent-home resolution from session_db.db_path. 3/3.
2026-08-14 13:56:58 -07:00
kshitij 367f0c21ed feat(agent): resolve sequential tool deadline via timeouts.tools.sequential_call (#85125 2a)
Follow-up on the #84795 salvage: the sequential deadline gets its own
resolver key. Unset, it inherits the concurrent batch deadline (same
value, same HERMES_CONCURRENT_TOOL_TIMEOUT_S bridge) so the two executor
paths cannot drift by default; set, it can be tuned or disabled
independently. Documented in cli-config.yaml.example; 5 contract tests.

Deliberately NOT on run_bounded_sync: the executors extend deadlines
dynamically during human approval waits (authorization-gate excluded
seconds) — the shared primitive is fixed-deadline. Noted in the docstring.
2026-08-15 02:21:39 +05:30
fangliquanflq 61645cde82 fix(agent): exempt clarify from sequential tool deadline
Clarify waits on a human for up to 3600s or unlimited. The generic sequential timeout was aborting that wait at 420s and leaving the prompt and worker active.
2026-08-15 02:21:39 +05:30
fangliquan 82a1b5a115 fix(agent): suppress late timeout observer events 2026-08-15 02:21:39 +05:30
fangliquanflq ededa8c4f1 fix(agent): bound sequential tool calls 2026-08-15 02:21:39 +05:30
teknium1 f79440e0f4 feat: /loop — recurring in-session wakeups (Claude Code parity)
Ports Claude Code's /loop (and its /proactive alias) across every Hermes
surface. /loop [interval] <prompt> re-runs a prompt or slash command on a
recurring cadence inside the live session; omitting the interval enables
self-paced mode (starts at the floor, backs off exponentially while the
agent's replies stop changing, snaps back on change — local digest
comparison, zero extra LLM cost).

Stop conditions: agent-emitted LOOP_COMPLETE marker, --times N,
--until <condition> (judged by the existing goal_judge aux task,
fail-open), /loop stop, and a loops.max_ticks backstop budget.

Core: hermes_cli/loops.py (LoopState + LoopManager + shared
dispatch_loop_command), persisted per session in SessionDB state_meta
(loop:<sid>) so /resume picks it up; migrates across compression
boundaries like /goal. New SessionDB.list_meta_prefix() powers the
gateway's cross-session scan.

Surfaces:
- CLI: /loop handler + idle-fire and post-turn-complete hooks in
  process_loop (mirrors the /goal hook shape; Ctrl+C pauses the loop)
- Gateway: /loop handler with route capture, mid-run control-verb guard,
  post-turn tick completion, and a supervised loop_wakeup_watcher that
  injects due wakeups into idle chats via the synthetic-message path
- TUI/dashboard/desktop: command.dispatch handler + per-session
  notification-poller wakeup driver + post-turn completion in the turn
  dispatcher; /loop added to the desktop slash palette
- /goal mixing: an active non-parked goal owns the idle boundary — loop
  ticks defer until it finishes, pauses, or parks; real user input always
  wins over both

Config: loops.{min_interval_seconds,max_ticks,self_paced_floor_seconds,
self_paced_ceiling_seconds}. Docs page + sidebar entry. 77 new tests.
Slack's 50-slash cap: /version moves to /hermes version to free the
native slot for /loop.
2026-08-14 13:40:19 -07:00
kshitij d6a5cb9725 Merge pull request #85147 from kshitijk4poor/feat/unified-deadline-layer
feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
2026-08-15 01:10:40 +05:30
joaomarcos 11c5aae104 fix(compaction): gate checkpoint replay/prune on current request eligibility
A captured native-compaction checkpoint lives in the persisted
codex_reasoning_items sidecar, but the wire restructure that follows it
(prune_pre_checkpoint_items) ran unconditionally: the native gate only
decided whether context_management went into the request, and no signal
from it ever reached _chat_messages_to_responses_input.

So a single checkpoint kept deleting every pre-checkpoint item from all
later requests — after a mid-session swap out of the gpt-5.6 family,
after compression.enabled: false, after the rejection kill switch, and
after a session resume that reloads the sidecar from state.db. The model
receiving the opaque blob was no longer the one able to decode it, and
nothing was logged.

Thread a single native_compaction_eligible boolean, derived from the same
value that gates the context_management field, into the converter. When
ineligible: do not replay type: "compaction" items and do not prune. Safe
because native compaction never truncates Hermes' local history, so the
fallback still carries the full conversation.

All Responses call sites are covered: build_kwargs and convert_messages
derive the flag via _native_compaction_active, the auxiliary/compression
client is explicitly ineligible, and the converter defaults to False
(pre-feature wire) so future call sites are safe by construction.

Fixes #85914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 10:39:25 -07:00
lkz-de 83d373aae6 fix(cli): guard expensive startup model overrides
Run the expensive-model warning for explicit startup `-m` / `--provider`
overrides before the chat loop starts, and fail closed for non-interactive
invocations that select an expensive or known-confusing model.

Also classify Nous paid-model 404s that say credits are required as billing
exhaustion so they fail fast with billing guidance.

Tested:
- scripts/run_tests.sh tests/hermes_cli/test_cli_startup_model_cost_guard.py tests/hermes_cli/test_model_cost_guard.py tests/agent/test_error_classifier.py -- --tb=short -q
2026-08-14 01:31:31 -07:00
fangliquanflq 8cf9e8a61b fix(compression): ignore background process notifications 2026-08-14 01:08:13 -07:00
Teknium fbec3c78fd fix(compression): clear preflight block on provider-confirmed rearm
A provider-confirmed rearm (#85846) resets the shared attempt budget, but
an earlier insufficient-progress verdict left _preflight_compression_blocked
armed, keeping the pre-API gate dark for the rest of the turn — a later
pressure spike could still grow unchecked until the provider overflow
handler fired. Clear the blocker and the stale pressure reading inside the
provider-confirmed rearm branch: the prompt is proven back below the
threshold, so the old verdict describes a request shape that no longer
exists.

Builds on @h-mascot's #84995, whose commit is preserved on this branch;
his rearm condition was superseded by #85846's latch-verified variant, but
the blocker-clear half was correct and is kept.
2026-08-13 23:43:54 -07:00
Henry Mascot 2b9da1f252 fix(compression): re-arm same-turn attempt budget 2026-08-13 23:43:54 -07:00
Gille 380e4da3d1 fix(compression): rearm budget from verified usage 2026-08-13 23:08:19 -07:00
ai-ag2026 72c828ca2c fix(agent): refund per-turn compression budget after real progress
A marathon tool turn burned all compression_attempts on *successful*
pre-API compactions; the gate then went permanently dark and the context
grew unchecked until the provider rejected the request terminally
("Context length exceeded: max compression attempts (3) reached", session
f087963205f9, 2026-08-01). The budget now refunds at loop top when the
assembled request is back under threshold * 0.8 AND the compressor's own
should_compress() agrees the pressure is gone.

Anti-thrash intent of the cap is preserved (#11529): no-progress passes
never reach the refund margin, divergent-signal cases (should_compress
still True) keep the budget burnt, and the insufficient-progress blocker
is untouched.

Validation: new behavioral suite (7) + all 130 compression/context tests
via scripts/run_tests_hermetic.py.

Rückbau: Commit revertieren; kein Zustand, keine Migration.

(cherry picked from commit 041b489d566bfb1d6816d53d9acd5ac50b8d5af0)
2026-08-13 23:08:19 -07:00
joaomarcos e331b834fe fix(usage): preserve OpenAI-wire cache writes at canonical accounting boundary (#85706)
Salvaged from PR #85702 by @JoaoMarcos44, composed onto the mapping-safe
_usage_get reads (PR #74591 by @RelaxJonh) and the flat cached_tokens /
Anthropic-name fallbacks (PRs #66105, #52571):

- cache-write precedence in the chat_completions branch:
  details.cache_write_tokens > details.cache_creation_input_tokens >
  usage.cache_creation_input_tokens > usage.cache_write_tokens
- codex_responses branch reads details.cache_write_tokens (GPT-5.6+
  documented name) with cache_creation_tokens fallback (from PR #70522)
- _usage_count(): clamp malformed negative counters to 0
- all reads in every branch are mapping-safe via _usage_get
2026-08-13 22:09:37 -07:00
RelaxJonh 11de734331 fix(usage): support dict-shaped usage objects in normalize_usage (#74314)
When the Responses API returns usage as a plain dict (e.g. from a
middleware or proxy that deserialises JSON to dict instead of a typed
SDK object), normalize_usage() used getattr() exclusively, which
silently returned 0 for every field on a dict.

Add _usage_get() helper that reads via .get() for dicts and getattr()
for attribute-style objects. All accessor sites in normalize_usage()
now use this helper, so token counts and cost are correct regardless
of the usage object's type.

Regression tests: two new tests feed the same payload as both a dict
and a SimpleNamespace through the codex_responses and
chat_completions branches, asserting identical output and non-zero
values.
2026-08-13 22:09:37 -07:00
mehmetkr-31 5b82e69693 fix(agent): read Kimi/Moonshot top-level cached_tokens in normalize_usage (#65722)
Kimi/Moonshot's native API (api.moonshot.cn / .ai) reports context-cache hits
as a top-level ``usage.cached_tokens``. The chat-completions branch of
normalize_usage() walks a fallback chain of
prompt_tokens_details.cached_tokens -> cache_read_input_tokens ->
prompt_cache_hit_tokens; none of those names match, so direct Kimi sessions
normalized to cache_read_tokens=0. The hits were invisible in accounting and
the cached prefix was billed at the full input rate.

Appended as the last link in that chain, so it only fills a genuine zero and
cannot override a provider that reports the nested OpenAI shape or DeepSeek's
prompt_cache_hit_tokens.

Rebuilt on current main rather than rebased — the branch was ~3400 commits
behind. The DeepSeek half of the original branch is dropped: 03c0b00f4
(#65678) landed prompt_cache_hit_tokens on main, so this is Kimi-only as the
review asked. The scripts/release.py addition to the frozen LEGACY_AUTHOR_MAP
is dropped too; contributors/emails/mehmet.kar@std.yildiz.edu.tr already
exists on main.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:09:37 -07:00
Chengxi Zhou 69bc3159e8 fix(agent): fallback to Anthropic-style token fields in normalize_usage
Local OpenAI-compatible servers like mlx_vlm.server emit
input_tokens/output_tokens in chat completion responses instead of
prompt_tokens/completion_tokens. The OpenAI Python client preserves
these as extra attributes, but normalize_usage() only looked at
OpenAI-style field names, causing input_tokens to always be 0.

This made the context progress bar stay at 0% forever and prevented
auto-compression from triggering.

Fix: add Anthropic-style fallback (input_tokens/output_tokens) in the
default else-branch of normalize_usage, with OpenAI-style names taking
priority.

Fixes #14686
2026-08-13 22:09:37 -07:00