Commit Graph

2147 Commits

Author SHA1 Message Date
CoolStar aaca343110 fix(agent): preserve Bedrock redacted reasoning replay 2026-09-01 08:30:26 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
joaomarcos 9d3f1de994 fix(anthropic): alias session_search/memory OAuth billing-classifier triggers
Anthropic subscription OAuth (claude_code credential) misroutes Hermes
sessions carrying the session_search or memory toolset into the
extra-usage lane, surfacing as HTTP 400 "You're out of extra usage" on
a valid subscription. Live-verified via the
anthropic-ratelimit-unified-representative-claim response header
(deterministic lane oracle, no dependency on the laggy usage counter):
tool schemas are innocent, the trigger is three specific system-prompt
sentences (session_search recall + two skill_manage sentences),
required jointly — breaking any one clears the classifier.

Two independent layers:
1. OAuth wire alias (anthropic_adapter.py, transports/anthropic.py):
   session_search -> chat_history_lookup, memory -> context_notes in
   tool name, description, and (session_search only) system-prompt
   prose, with wire-collision guarding and a normalize_response
   reverse-map that keeps GH-25255 registered-tool precedence. Also
   routes named tool_choice through the same normalizer, closing a gap
   where a forced tool_choice would leak the raw trigger string and
   stop matching tools[].
2. Prompt-preserving reword (prompt_builder.py): rewords the two
   triggering SKILLS_GUIDANCE sentences while keeping the same
   meaning, still naming skill_manage, and leaving the Skill Safety
   Rule section untouched. Applies to all auth paths since it's a
   prompt-copy change, not a wire-level transform.

Two layers rather than one because the three-sentence AND-condition
means a classifier tightening could start firing on either remaining
leg alone.

Fixes #65365
2026-09-01 16:27:20 +05:30
Teknium 530aa7b10f test(cache): fix stale Studio construction-order paragraph in the bridge witness docstring
The module intro still described the pre-b449d7c1 order (row first, agent
second); _reply_scope and numbered item 2 already state the correct one.
2026-09-01 02:14:35 -07:00
joaomarcos 41ec9d3591 fix(cache,tests): durable generation docstring + studio-bridge order (PR #98811, refs #96811)
agent/prompt_cache_scope.py: durable per-(source,session_key) monotonic generation; list _RESET_END_REASONS; note never garbage-collected. tests/agent/test_studio_bridge_affinity.py: build AIAgent before db.create_session; docstring corrected.
2026-09-01 02:14:35 -07:00
joaomarcos 5abba03a99 test(cache): pin the affinity contract on Studio's real bridge construction
The reproduction reported on #96811 is Hermes Studio's group chat, and it is
the one half of the issue that cannot be closed from inside this repository.
This states why in executable form, and pins the contract the host adoption
depends on so a later refactor cannot quietly break it.

Root cause, traced to the host. Studio reaches Hermes as a LIBRARY, not
through the gateway. Its bridge mints groupRuntimeSessionId(room, profile,
name) -- a gc_run_ prefix truncated to 96 characters plus a fresh UUID4 hex --
for every reply, writes the session row itself, and constructs AIAgent(...)
with platform / session_id / session_db and no routing identity of any kind:
no gateway_session_key, no chat_id, no user_id, no parent_session_id. So the
declared scope is unreachable, the row is a lineage root, and
resolve_prompt_cache_scope() correctly falls through to the physical id, which
moves on every reply. No stable carrier crosses the boundary. Hermes must not
recover one from the id's syntax -- that is #79017's collision class, and the
two negative controls merged via #97704 exist to keep it from trying.

What does cross the boundary is a value Studio already has:
groupBridgeSessionId(room, profile, name, sessionSeed, runtimeConfig) is
stable for one conversation in one room, already carries the room, profile,
agent name and the room-owned sessionSeed, and is already hashed and
length-bounded. Passing it as gateway_session_key is the entire adoption, and
it is a Studio-side change; this PR keeps Refs #96811 for exactly that reason.

What this suite adds is the half that IS reachable here. The existing suites
simulate Studio's id SHAPE against synthetic agent doubles; none of them runs
the host's construction path, so nothing today would fail if that path stopped
honouring a declaration. These tests build a real AIAgent in the bridge's own
order -- row written first, agent second, new physical id per reply -- and
state what the adoption buys:

- three replies with distinct physical ids hold ONE affinity scope, and the
  scope carries neither the room nor the member name;
- two members of one room, and two rooms, never share a bucket;
- a new sessionSeed rotates the scope, which is Studio's own conversation
  boundary and needs no reset observed by Hermes;
- equal keys under different row sources never collapse, because the row's
  source -- not agent.platform -- is the identity the peer queries match on;
- a tool child and a background-review fork on the same declared key keep
  their own scope, so #79161 survives this construction path too;
- and an undeclared bridge is byte-identical: per-reply scope, and a
  compression rotation still walks its lineage.

Against origin/main the two declared-conversation assertions fail --
"assert 3 == 1", and the raw gc_run_ id leaking as the routing key -- which is
the reproduction; the remaining eight pass there and here, because they
describe behaviour that must not change.

Reported by @cervantesh, whose re-check of Studio main@86d0c95375 located the
missing consumer on the real path and asked for exactly this witness.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos e5bce4df4b fix(cache): make the conversation generation survive a backwards clock
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.

latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.

Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos d7995bffaf fix(cache): qualify the declared key with the conversation generation
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.

No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.

That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.

It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.

The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
kshitijk4poor 29d4c0ebfd refactor(compression): extract preflight seed predicate onto ContextCompressor
Address review follow-ups on the seed fix:

- Move the 'seed only from the 0 state' guard from an inline block in
  build_turn_context() into
  ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
  the predicate with the rest of the speculative-seed lifecycle
  (snapshot_preflight_display_tokens /
  rollback_interrupted_preflight_display_tokens). Callers now use the
  method via a getattr guard so test doubles and external context
  engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
  described the old >=0 guard ('treats any negative value as no real
  usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
  method on the compressor fixture, eliminating mirror-drift risk (the
  helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
  low until their next report) in the method docstring.
2026-09-01 14:33:23 +05:30
Turgut Kural 80b836fa9e fix(compression): preflight display-seed must not overwrite real provider usage
The preflight rough-estimate seed used 'last >= 0' semantics, so any
provider that reports real prompt_tokens got its reading replaced by
the schema/reasoning-inflated rough estimate whenever the estimate was
larger. last_prompt_tokens feeds both the CLI context meter (cli.py)
and the post-response compression gate (conversation_loop.py), so one
seed made the bar jump to an inflated number AND could push the
real-usage gate over threshold on estimator noise.

Observed in a production CLI session on a 1M-token window with a
reasoning-heavy history: the status bar showed ~492K real provider
prompt tokens; the next turn's preflight estimated ~685K (rough
estimates can inflate 1.4-2.5x on reasoning-heavy sessions, #81481)
and seeded it into last_prompt_tokens; compression then fired at ~69%
of the window while real usage was ~49%.

Policy change: a real provider reading (>0) always wins. The seed now
only fills the 0 state ('no reading yet'), keeping the status bar live
for usage-less providers (#34282's motivation); -1 remains protected
as the post-compression sentinel (#36718).

Updates TestPreflightSentinelGuard to encode the new policy and adds a
regression test with the measured 492K/685K numbers.
2026-09-01 14:33:23 +05:30
Brooklyn Nicholson 5c6dbe22c3 fix(compression): take reasoning_content when the summarizer leaves content empty
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30
Teknium 8dbf07e950 fix: satisfy the no-locked-pure-readers gate and close the DB before rmtree
_record_db_file_identity's PRAGMA fallback is a pure read — route it
through _read_ctx() instead of the writer lock (Pattern C gate).
test_codex_turn_persists_each_message_exactly_once leaked a live
SessionDB into shutil.rmtree, racing the WAL sidecars ('Directory not
empty' on CI); close the handle first and rmtree with ignore_errors.
2026-08-31 14:02:55 -07:00
theo 8207862212 fix(compression): stop timeout paths from blocking retries 2026-08-31 13:00:33 -07:00
theo 0a8b25e0ab fix(auxiliary): forward service tier on Codex Responses 2026-08-31 13:00:33 -07:00
Teknium 13c3958df5 fix(auxiliary): credit-limited 402s clamp to the affordable budget instead of failing
Second leg of the masoria 20-minute 'Summarizing' stall (Aug 31 2026
bundle): after the Codex timeout, compression fell back to OpenRouter,
which defaulted the omitted output cap to the model's full 65,536-token
window and rejected with '402 ... can only afford 7117' — on an account
whose balance easily covered a summary. Three fallbacks, three 402s,
zero summaries.

- _create_with_progress: when a 402 names an affordable budget, retry
  ONCE with that cap (minus 64-token margin, 512-token floor); plain
  exhaustion 402s and within-budget 402s re-raise unchanged. Single
  funnel covers primary, fallback, and retry call sites.
- extends the #41055 OpenRouter max_tokens preservation onto the current
  _build_call_kwargs gate (explicit caps survive; None still omitted).

Live A/B against a mock OpenRouter enforcing a 7117-token budget with
the real SDK + real adapter: main fails with the exact bundle 402;
fixed branch retries clamped and returns the summary.
2026-08-31 13:00:24 -07:00
liuhao1024 b26a1eae8f fix(auxiliary): preserve max_tokens for OpenRouter to prevent free-tier 402
_build_call_kwargs strips max_tokens for non-Anthropic providers to
avoid wire-format issues (Copilot, ZAI, GPT-5). However, OpenRouter
free/limited-credit tiers need max_tokens because the model's full
output window exceeds the credit budget, causing HTTP 402.

Without max_tokens, the 402 triggers fallback to a text-only model
which then fails with 'unknown variant image_url, expected text'.

Include max_tokens when provider is 'openrouter' or base_url contains
'openrouter.ai'.

Fixes #41035
2026-08-31 13:00:24 -07:00
Teknium 093d1ebd0b test(auxiliary): watchdog timeout asserts owner-thread FD release
Adapt the #99660 blocked-before-first-event watchdog test to the FD-ownership
contract: the stranger-thread Timer marks the timeout but never close()s;
the owning thread surfaces the TimeoutError and releases the FDs on unwind.
2026-08-31 13:00:18 -07:00
dsad 8a5b49d86a fix(auxiliary): never release Codex client FDs from the timeout Timer
_CodexCompletionsAdapter.create arms a daemon threading.Timer that calls
client.close() when the aux Responses stream exceeds its timeout. On a
stalled stream -- the failure the timeout exists for -- the Timer is the
only thing that fires, so the close runs on a thread that does not own
the in-flight httpx connection.

That is the FD-ownership violation the repo already fixed twice on the
main transport (#29507, #67142, #70773): close() releases the raw TLS fd
while the owner's OpenSSL BIO still caches that integer, the kernel
recycles it into the next open() in the process -- a SessionDB or
kanban.db handle -- and the owner's unwinding TLS flush writes an
application-data record into that database file.

agent/auxiliary_client.py had no thread-ownership machinery at all: the
guarded twins (_retire_shared_openai_client, _abort_request_openai_client)
live in run_agent.py and are unreachable from this adapter, which holds no
AIAgent reference.

Dispatch on ownership the way chat_completion_helpers already does: from a
stranger thread only force_close_tcp_sockets() (shutdown(SHUT_RDWR), which
is FD-safe from any thread), and let the owning thread release the FDs when
it unwinds. The owner-thread caller (_check_cancelled) keeps closing
directly. Cache eviction (#23432) is unchanged.
2026-08-31 13:00:18 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Teknium 58f5b1e277 fix(model_metadata): parse Google's 'supports up to N' context-limit phrasing
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.

Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.

Reported by @Artemonim in #57275 (residual claim 5).
2026-08-31 12:20:02 -07:00
Darafei Praliaskouski 4252aecc2e fix(agent): cap compaction threshold floor at 85% of the context window
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only
degraded to the 85% trigger when it met or exceeded the effective
window exactly (#14690). Near-minimum windows slipped through: at
context_length=65536 the threshold passed through at 64,000 — 97.7%
of the window, ~1.5K tokens of output room — so pre-API compaction
effectively could not fire.

Providers that silently truncate over-window prompts instead of
rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never
deliver the reactive context-overflow backstop either. Observed live
on a 65,536-token local model: the session rode into the window
ceiling and each length-continuation retry re-sent a window-filling
prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room)
until the turn died with "Response remained truncated after 4
continuation attempts" — every retry paying a full multi-minute
prefill.

Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the
effective input budget whenever the floor is the binding term. An
explicit threshold_percent above 85% is user intent and stays
uncapped; windows where the floor lands at/below the cap are
unchanged.
2026-08-31 12:19:29 -07:00
kshitijk4poor fe0cfdf99c fix(redact): keep dotted config-key scans linear past the keyword pre-gate
The _CFG_SECRET_WORD_RE pre-gate only skips secret-FREE text. A compaction
payload containing one real secret assignment plus a long opaque dotted run
still reaches _CFG_DOTTED_RE's backtrackable '*' prefix, which re.sub retries
from every byte of the run — quadratic while holding the GIL (same class as
the _ENV_ASSIGN_LOWER_RE fix in this branch, #99255).

Anchor each attempt to the start of a key run with a negative lookbehind.
Match set is unchanged: any match starting mid-run implies a leftmost match
at the run start, verified 20/20 identical over a dotted-config corpus.
30k-char adversarial run: 102s -> 0.015s.
2026-09-01 00:25:45 +05:30
Teknium f50b5bb0fa fix(compression): dead Codex summary streams fail over in 60s instead of stacking 5-minute waits
The Codex auxiliary Responses adapter enforced a single absolute
deadline (300s floor for compression). A dead stream held the entire
budget before fallback ran, and repeated compression attempts stacked
those waits into 20+ minute 'Summarizing thread' stalls (masoria debug
bundle, Aug 31 2026). Meanwhile a healthy-but-slow reasoning summary
was killed at the same absolute deadline even while producing tokens.

Replace the absolute kill with progress-aware deadlines:
- 60s no-progress window for the first substantive payload AND between
  payloads; keepalive/lifecycle frames do not re-arm (mirrors the
  commit-fence gating, #96707)
- a live stream re-arms per token and is bounded only by
  _aux_stream_total_ceiling() (max(600s, 4x configured timeout)), the
  same backstop the streamed chat.completions path already uses
- the compression critical-path retry gate now distinguishes failure
  cost: a cheap first-token no-progress failure retries the same
  provider once; mid-stream stalls and ceiling hits still skip straight
  to provider fallback (#54465 semantics preserved)

Live A/B (real OpenAI SDK against a local SSE server, real adapter):
dead keepalive-only stream: main waits the full budget; fixed fails
over at the window. Slow-but-alive stream (tokens past the configured
timeout): main kills it mid-generation; fixed completes.
2026-08-31 11:52:51 -07:00
anhtahaylove 99d037eeb9 fix(compression): count streamed reasoning details as progress 2026-08-31 11:18:53 -07:00
anhtahaylove 4d95d8eca8 test(context): cover preflight timeout provider boundary 2026-08-31 11:18:53 -07:00
anhtahaylove de49e1ef05 fix(context): fail closed when preflight compression stalls 2026-08-31 11:18:53 -07:00
fangliquanflq ba0f5839d4 fix(redact): keep lowercase assignment scans linear 2026-08-31 11:18:41 -07:00
teknium1 d3a45a9ce4 test(agent): update enqueue-after-close contract to the #94736 self-heal
The old contract (write after close() raises AttributeError and drops
the token delta) is superseded: the persistence boundary now reopens
the connection, so the delta lands. Assert the new, stronger contract.
2026-08-31 10:51:19 -07:00
HexLab98 97f2a571b5 test(terminal): cover hung-wait bound, parent-tid interrupt, and cron inactivity watchdog
Pin that execute() returns at the wall-clock deadline when the inner wait
never returns, that /stop on the tool-worker tid still kills the subprocess,
that the cron inactivity helper fires while the caller thread is blocked,
and that ContextVars plus the activity callback reach the deadline worker.
2026-08-31 10:42:39 -07:00
Teknium d8f8a07ee3 fix(compression): truncated summaries no longer become compaction checkpoints (port of earendil-works/pi#7048)
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.

Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
  main-model fallback (a larger output budget may finish the summary), and
  on terminal failure ABORTS compression preserving the session unchanged
  (new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
  exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
  recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
  existing retry/backoff loop.

_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.

Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.

Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
2026-08-31 10:39:55 -07:00
fangliquanflq c6ee4e0809 fix(prompt): skip bundled AGENTS.md for desktop launch cwd 2026-08-31 10:10:25 -07:00
Teknium b7ebe6456f fix(xai): request-local alias provenance + collision-safe wire aliasing
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:

- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
  alias map THIS request emitted; the transport stashes it
  (_last_wire_aliases) and normalize_response reverses ONLY those aliases.
  A real user/plugin/MCP tool named hermes_tool_search is never silently
  dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
  bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
  that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
  a prior request can't leak into the next response's dispatch.

Refs #95003
2026-08-31 10:09:04 -07:00
liuhao1024 5e2f8b9865 fix(xai): alias the reserved tool_search bridge name on chat completions
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
2026-08-31 10:09:04 -07:00
David Zhang de123be524 fix(xai): alias the reserved tool_search bridge on the wire (#95003)
xAI reserves the function name `tool_search` for Grok's native
server-side Tool Search and rejects the client declaration outright:

    HTTP 400 {"code":"invalid-argument","error":"The function name
    tool_search is reserved for the tool_search tool"}

Hermes's progressive-disclosure bridge registers exactly that literal
(`TOOL_SEARCH_NAME` in tools/tool_search.py) and assembly is not
provider gated, so with the default `tools.tool_search.enabled: auto`
every grok turn fails the moment the catalog crosses the threshold —
mid-session, which reads to the user as a session reset.

Same treatment as the two collisions already handled on this
transport (xAI `web_search` #48108, OpenCode reserved names #85589):
alias to `hermes_tool_search` on the wire in build_kwargs, map back in
normalize_response so Hermes dispatch and the bridge contract are
untouched. `tool_describe` / `tool_call` are not reserved by xAI and
are left alone.

Folds the per-provider rename helpers into one `_alias_reserved_tools`
owner parameterized by the reserved-name tuple, and extends the
existing `_RESERVED_ALIAS_TO_NAME` reverse map so the dispatch-side
un-aliasing needs no new branch.

Scope note: this covers the Responses transport, which is where every
api.x.ai route lands by default (`_fallback_api_mode` maps api.x.ai →
codex_responses, and the xai provider profile declares it). An xAI
model forced onto `api_mode: chat_completions` would still hit the
400; that path has no provider-specific tool rewriting today and would
need the symmetric hook in agent/transports/chat_completions.py. Happy
to add it here if you'd rather have both in one change.

Tests: new TestXaiReservedToolSearchAlias covering the wire alias,
non-xAI backends keeping the canonical name, composition with the
native web_search swap, and the normalize_response round trip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012vLaAmnsdii3Gm9jMDs5gw
2026-08-31 10:09:04 -07:00
zengzheqing 37ae0e1d39 fix(curator): remove terminal from the consolidation fork (issue #96962)
The curator LLM fork was steered by its own prompt to re-home skill
support files with terminal `mkdir -p ... && mv ...`. A terminal move
writes the same bytes with NO ledger entry, so the archive that follows
snapshots an already-stripped package (files: 1) and `hermes curator
rollback` restores a hollow skill — SKILL.md back, references/ gone.

Remove the capability rather than guard it: the fork's enabled_toolsets
drops "terminal", so terminal and process disappear together and there
is no shell to parse, no process stdin to feed, no remote-backend
divergence — a heuristic command guard over a Turing-complete input
space can guarantee none of that. Every mutation the pass needs has a
ledgered skill_manage action (write_file / remove_file / delete), and
the prompt now steers exactly those. Reading works through skill_view.

Tests pin both halves: the call-site kwarg (["skills"] only), the
resolved surface (no execution/write tools), and the prompt steering
(no mkdir -p / mv shapes).
2026-08-31 10:08:13 -07:00
Teknium c0667439ec fix(compression): rotation heals stale automatic ended_at stamps instead of wedging (#88197)
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).

Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.

TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
  rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
  session_reset (deliberate boundary) => still aborts BEFORE the #47202
  flush, preserving #88411's no-growth contract where an abort remains
  correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
2026-08-31 09:58:11 -07:00
Teknium 8b73720fa7 fix(memory): tolerate bare-signature v2 providers when forwarding checkpoint requirement
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
2026-08-31 09:57:22 -07:00
Soju06 5db9058cdb fix(memory): forward checkpoint requirement to v2 providers
MemoryManager.on_pre_compress() detects checkpoint API v2 providers,
selects the normalized evidence list for them, and re-raises their
failures under require_checkpoint — but it never tells the provider
that a checkpoint is required: the call passes only the messages.

A v2 provider therefore runs in its default best-effort mode, swallows
durable-write failures, and returns normally; the host then treats the
checkpoint as succeeded and lossy compression proceeds. With
compression.checkpoint_required: true this silently defeats the
guarantee the option exists to provide.

Forward require_checkpoint only to providers advertising the requested
checkpoint API version. Legacy providers keep the strict one-argument
on_pre_compress(self, messages) contract, so bundled v1 providers
(honcho, mem0, supermemory, ...) are unaffected.

Regression tests cover required and best-effort signaling, legacy
signature compatibility, and required-mode failure propagation.
2026-08-31 09:57:22 -07:00
kshitijk4poor 3aee290899 refactor(compression): name the split-failure cooldown; drop duplicate tests
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
  comment pinning WHY it is the timeout ladder's first rung (transient
  lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
  condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
  test_compression_rotation_state.py / test_context_compressor.py; kept the
  force-bypass test (only site pinning that behavior for split failures) and
  the E2E test (now asserting the named constant).
2026-08-31 14:09:42 +05:30
kshitijk4poor 81ab11b821 test(compression): E2E-pin the split-failure cooldown through a real SessionDB
The salvaged unit tests drive _record_compression_failure_cooldown directly;
this drives the real _compress_context split-failure path (archive boom on a
real SessionDB) and asserts the cooldown recording fires with the
session_split_failed error class.
2026-08-31 14:09:42 +05:30
VVV 087cc49a26 fix(compression): refresh lease in-transaction before publish; arm cooldown on split failure
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):

1. publish_compression_child gains require_lease_refresh: the lease is
   extended inside the same transaction as the expiry check (same conn, no
   TOCTOU), giving a worker whose refresher thread died from transient DB
   failures one final chance to keep its completed work.

2. A failed compression split now records a 60s failure cooldown, so the
   next turn cannot immediately re-trigger the same compression.

Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
2026-08-31 14:09:42 +05:30
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00
Teknium ed3562bbbc test(compression): drop moot digest-loop tests; match stamped backoff errors
PR #98628 removed _build_chunk_digests, so the two lean chunk-digest
cancellation tests reintroduced by the #97512 cherry-pick target a
deleted mechanism — removed. The #96775 stall-interrupt assertions now
match the stall_interrupted marker inside the strategy/kind-stamped
durable error instead of assuming it is the prefix.
2026-08-30 19:46:21 -07:00
Teknium ad925a08da fix(compression): scope worker-teardown grace to the total-ceiling path
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
2026-08-30 19:46:21 -07:00