Commit Graph

2233 Commits

Author SHA1 Message Date
Teknium dc7e1b7ab9 fix(webhook): load URL-resolved profile's skills under multiplex
A `/p/<profile>/webhooks/<route>` request resolved the profile from the URL
but ran the route script, prompt render and `skills:` lookup with no
profile scope — the runner only enters `_profile_runtime_scope` later,
around `handle_message` — so routed webhooks loaded the launch (default)
profile's skills and logged "Skill not found" for the routed profile's own.

- gateway/platforms/webhook.py: add `_profile_scope(profile)` (nullcontext
  when no prefix was resolved; `_profile_runtime_scope(get_profile_dir(p))`
  otherwise, same helper the runner uses) and wrap the script / render /
  skill-injection block in it. Bare routes are unchanged.
- agent/skill_commands.py: `scan_skill_commands` scanned the import-time
  `SKILLS_DIR` (frozen to the launch home), so even a correctly scoped call
  listed default's skills; the #88023 home-keyed cache alone could not fix
  that. Use the call-time `_skills_dir()` there and at the two other
  SKILLS_DIR-relative sites in the module.
- agent/skill_utils.py: `normalize_skill_lookup_name` used the same frozen
  root, so a routed profile's absolute skill_dir was rejected by
  `skill_view` ("must be a relative path within the skills directory").
  Resolve against `_skills_dir()` — the root `skill_view` itself enforces.

Fixes #67277

Co-authored-by: Juani Lezcano <tky.juani@gmail.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium 98c27aa25c feat(webhooks): stamp the emitting profile on outbound webhook payloads
Under a multiplexed gateway every profile's outbound webhooks share one
delivery worker, so receivers could not tell which profile fired an
event. Add a top-level `profile` field to the payload, resolved at fire
time from the bound Hermes home via get_active_profile_name() ("default"
outside profiles). Documents the field in the wire-format section.

Reported by @vszgdcn8cj-ctrl.

Fixes #92674
2026-09-02 07:00:13 -07:00
chelsealong 600b9c7e7e fix(gateway): scope force-reload hook re-registration to its own home
re_register_config_hooks() cleared the entire process-global idempotence
set on every force-reload, so a profile-local plugin force-reload dropped
another live profile's ledger key without touching its still-registered
callback — the next registration call for that profile then appended a
duplicate. Scope the clear to the reloading profile's own home, and give
outbound webhooks the same force-reload restoration shell hooks already
had, since unload() wipes both from the shared _hooks dict.
2026-09-02 07:00:13 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Aldo 6ff9426d22 fix(anthropic): track the current fast-mode model matrix (Opus 4.8 / Opus 5)
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):

- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
  only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
  requests silently run at standard speed and bill standard rates
  (usage.speed: 'standard'). Today's allowlist therefore shows 4.6
  users a fast toggle that does nothing, while denying it to the two
  models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
  select fast inference via the model field and are explicitly
  excluded from the param gate.

Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.

## How to test

scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q

113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
2026-09-02 05:33:13 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
Teknium bc71b8bc95 fix(compression): anchor on the LAST intent row — newer user turn outranks older steer (#100053 follow-up)
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.

Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
2026-09-02 04:12:12 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
kshitijk4poor 2adb1a4ea6 refactor(agent): clone the review snapshot once at the spawn chokepoint
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.

Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
2026-09-02 13:28:41 +05:30
kshitijk4poor 26f0de23cf fix(agent): clone the /refine snapshot too, not just the automatic review
Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).

Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
2026-09-02 13:28:41 +05:30
Teknium 37f3ba110a fix(auth): auto-heal single-use OAuth grants already forked across profiles (#100339)
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.

`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.

Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.

Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
2026-09-02 00:58:29 -07:00
Teknium 3038493ee6 fix(auth): never fork single-use OAuth grants across profiles (#100339)
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:

1. `hermes profile create --clone-all` and the dashboard/TUI
   `mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
   verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
   drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
   `providers.<id>` device-code blocks, and the PKCE singleton file; API
   keys are still copied. The clone reads the root grant through the
   existing credential-pool root fallback.

2. A named profile with no local rows BORROWS the root grant via
   `read_credential_pool()`'s fallback, but every persist
   (`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
   rows into the profile's own auth.json — materializing a fork on the first
   rotation. `persist_pool_entries()` now routes borrowed single-use rows
   back to the root store (update-only, under the root lock; never falls
   back to a local copy). A borrowed `hermes_pkce` rotation commits its
   singleton to the root `.anthropic_oauth.json`, the borrower never prunes
   root-seeded rows it cannot see the backing file for, and
   `hermes -p <profile> auth add` persists only the profile's own rows.

Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.

Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).

Closes #100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 00:58:29 -07:00
Teknium 25d954c2cf fix(guardrails): hard stops catch replays, never legitimate iteration
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:

- Edit -> re-run is progress. A successful mutating call (write_file/patch,
  a green terminal/execute_code, browser actions, job/message/cron/memory/
  skill mutations) marks progress for every failing signature still being
  counted this turn; the next identical retry restarts its streak instead
  of accumulating toward exact_failure_block_after. A pure replay never
  mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
  (terminal, execute_code, process pollers, browser_navigate, web_extract)
  same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
  task loops with a live parent/client and do real edit -> re-run work.

Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
  unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
  this commit:        COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
2026-09-02 00:26:57 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
benbenwyb cd2d3089fb fix(agent): guard repeated skill reads
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.

Add regression coverage for repeated skill_view results under hard-stop guardrails.
2026-09-02 00:26:57 -07:00
João Vitor Cunha 384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
João Vitor Cunha ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
Teknium 30c9d40974 fix(compression): stop the Codex and Anthropic aux summary streams at the host deadline too (#99692)
PR #99779 gave the streamed chat.completions consumer the host's absolute
compression deadline. The two wires that consume their streams internally
still ran on their own, always-larger budgets after the host gave up:

- Codex Responses: clamp the re-armable watchdog's hard ceiling to the
  published host deadline, so a live (re-arming) stream is severed the
  instant the host stops waiting instead of at max(600s, 4x timeout).
- Anthropic Messages: the per-event hook now raises at the host deadline
  and on an explicit hard cancel; create_anthropic_message lets that
  TimeoutError abandon the stream (the with-block closes it) instead of
  swallowing it as a callback failure.

Sabotage-verified: without the Codex clamp the new deadline test hangs past
its 25s harness cutoff; without the Anthropic hook the three Anthropic tests
fail.
2026-09-01 23:56:06 -07:00
joaomarcos 904e5bb572 fix(compression): stop the summary stream at the host's own deadline
CompressionCommitFence.set_total_ceiling_seconds documents its deadline as
"shared by the host and worker", but only the host ever read it. The worker's
streamed summary bounds itself with _aux_stream_total_ceiling() instead —
max(600, 4 * aux_timeout) — which is >= the host's total ceiling for every
configured timeout AND starts counting later (after pool admission,
_serialize_for_summary, prompt build and TTFT). A stream that outlives its
abandoned host is therefore not an edge case; it is the guaranteed outcome of
every total-ceiling timeout.

8207862212 closed the first half: a cancelled fence now releases the
compression owner, freeing its pool slot and session lease. Its own comment
leaves the second half open — the isolated provider daemon that holds the
socket keeps streaming "until the auxiliary stream's longer absolute ceiling
expires". With the #99692 reporter's auxiliary.compression.timeout: 600 that
is 2400s of an orphaned ~500K-token summary the fence is already guaranteed to
refuse, and because the session never shrank, every following turn stacks a
fresh orphan on top of the last.

Publish the fence's deadline as an absolute monotonic instant
(CompressionCommitFence.deadline_monotonic) and give the auxiliary layer the
return leg it was missing: aux_stream_deadline() installs it thread-locally,
_ChatStreamAccumulator.feed() stops the stream once it passes, and
_run_protected_sync_provider_call propagates it onto the provider daemon
(thread-locals do not cross that boundary, so an owner-thread-only install
would be inert on exactly the path large-session compression takes).

Absolute, not relative: the deadline is unaffected by however long dispatch and
TTFT took before the accumulator was constructed. Checked as well as — not
instead of — the existing ceiling, so every caller without a host deadline is
byte-for-byte unchanged, and the "timed out" phrasing keeps _is_timeout_error
classification identical to a request timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 23:56:06 -07:00
Teknium c5c9aa8d44 fix(gateway): hygiene compaction keeps the profile secret scope under multiplexing
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').

- gateway/run.py: run both hygiene executor hops (detached-agent path and
  codex app-server path) inside copy_context().run, keeping the default
  executor so a fence-cancelled hung summary never occupies a gateway
  agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
  class failure — abort and preserve the session instead of dropping the
  middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
  ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
  stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
  run_in_executor); UnscopedSecretError classified as access failure.

Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
2026-09-01 22:28:52 -07:00
MattMaximo 6f625f7381 fix(tools): propagate caller contextvars in DaemonThreadPoolExecutor.submit
Some bundled CPython runtime builds strip stdlib ThreadPoolExecutor's
copy_context() propagation, so work submitted to the daemon pool runs in a
bare context. Under the multiplexed gateway this dropped the profile
secret scope in pool workers: the context-compression timeout fence
resolved auxiliary provider keys (SURPLUS_API_KEY) with
UnscopedSecretError, silently degrading LLM compression to lossy
deterministic summaries and driving re-read loops in affected sessions.

Restore stdlib semantics in submit() by snapshotting the caller's context
and running the callable inside it (a no-op re-application on runtimes
that already propagate). Mirrors the gateway's
_run_in_executor_with_context pattern.

Tests: daemon pool worker sees caller contextvars; scoped get_secret works
in a daemon-pool worker under multiplex while scoped misses still fail
closed (no env leak).
2026-09-01 22:28:52 -07:00
AgentLinker cac9db7caf fix(codex): retired stream requests must not synthesize a completed response
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.

Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.

Fix: publish a per-request retirement token so the worker can tell it has been
retired.

- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
  `agent._active_codex_stream_request_token` before handing off to the worker
  (codex_responses only) and clears it at all four kill sites plus the worker's
  own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
  which can raise — every other call site wraps it in try/except, and a leaked
  token would let a later worker mistake itself for the owning attempt. The
  request-local `_codex_request_retired` mirror also swallows the transport
  error our own force-close causes, so the worker's local error cannot replace
  the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
  `TimeoutError` from `interrupt_check` when it no longer owns the request —
  raising rather than breaking, because a break returns the partial `final`.
  The four stream callbacks also drop post-retirement frames so an abandoned
  attempt cannot stream tokens into the live turn's bubble (the gateway caches
  AIAgent instances per session).

`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.

Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
2026-09-01 22:14:06 -07:00
Teknium aac8d4b9e9 test: sweep two main-side tests onto the renamed gui_tour / process_manage names 2026-09-01 21:50:19 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium 419232d49b fix(codex): extend Happy-Eyeballs racing to Codex OAuth/auth clients; pin async native racing
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:

- hermes_cli/auth.py Codex OAuth clients (token refresh at
  auth.openai.com/oauth/token, device-code login, token exchange, usage
  probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
  every connect eats the full timeout per AAAA before IPv4 is tried, so
  auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
  had no explicit racing wired.

Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
  installs the existing _HappyEyeballsSyncBackend on a ready-built sync
  httpx.Client's direct transports (default transport + mounts), skipping
  proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
  host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
  five Codex OAuth/probe endpoints. Best-effort: falls back to default
  serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
  RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
  no custom backend needed. Documented in build_keepalive_http_client and
  pinned by tests (contract test on the anyio signature + a live
  regression test where a blackholed 100::1 IPv6 addr hangs and local
  IPv4 wins in ~250ms instead of the serial connect timeout).

network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.

Refs #13834; follows #94388 (9cce8725).
2026-09-01 12:08:11 -07:00
teknium1 37f5f1ff98 test(redact): corpus-level before/after coverage for value-aware gating (#96607) 2026-09-01 12:07:16 -07:00
fangliquanflq a8ddb231aa fix(redaction): gate ambiguous assignment values 2026-09-01 12:07:16 -07:00
xxxigm 0bee5ff408 fix(auth): look up keyed custom providers by durable pool slug
hermes auth add stores providers.<key> credentials under the config
slug, but runtime only tried custom:<display-name> and then sent the
no-key-required placeholder. Try the slug first, keep the legacy
namespace as fallback, and thread provider_key/key_env through named
custom resolution.
2026-09-01 22:42:27 +05:30
xxxigm 43470980bf test(auth): cover keyed providers.<key> credential-pool lookup
New-style providers store keys under the durable config slug, but
runtime still looks up custom:<display-name> and sends a placeholder.
2026-09-01 22:42:27 +05:30
CoolStar aaca343110 fix(agent): preserve Bedrock redacted reasoning replay 2026-09-01 08:30:26 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
Teknium 51609a35f6 fix(auth): purge silent OpenRouter paid-default adoption (#81952 class fix)
Three kills at the shared chokepoints:

1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
   while the active config.yaml is corrupt (AuthError code=corrupt_config).
   A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
   model.provider and tier-3/4 silently adopted the PAID openrouter provider
   against the user's real (unparseable) intent. New probe:
   hermes_cli.config.get_active_config_parse_failure(), recorded in the
   existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
   a fixed file clears the block immediately. Explicit provider requests
   are untouched.

2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
   (nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
   google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
   honored untouched (paid-lane warning retained).

3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
   process per provider) when a credential is newly ingested — ingestion
   itself stays allowed.

Fixes #81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
2026-09-01 07:00:38 -07:00
joaomarcos 9d3f1de994 fix(anthropic): alias session_search/memory OAuth billing-classifier triggers
Anthropic subscription OAuth (claude_code credential) misroutes Hermes
sessions carrying the session_search or memory toolset into the
extra-usage lane, surfacing as HTTP 400 "You're out of extra usage" on
a valid subscription. Live-verified via the
anthropic-ratelimit-unified-representative-claim response header
(deterministic lane oracle, no dependency on the laggy usage counter):
tool schemas are innocent, the trigger is three specific system-prompt
sentences (session_search recall + two skill_manage sentences),
required jointly — breaking any one clears the classifier.

Two independent layers:
1. OAuth wire alias (anthropic_adapter.py, transports/anthropic.py):
   session_search -> chat_history_lookup, memory -> context_notes in
   tool name, description, and (session_search only) system-prompt
   prose, with wire-collision guarding and a normalize_response
   reverse-map that keeps GH-25255 registered-tool precedence. Also
   routes named tool_choice through the same normalizer, closing a gap
   where a forced tool_choice would leak the raw trigger string and
   stop matching tools[].
2. Prompt-preserving reword (prompt_builder.py): rewords the two
   triggering SKILLS_GUIDANCE sentences while keeping the same
   meaning, still naming skill_manage, and leaving the Skill Safety
   Rule section untouched. Applies to all auth paths since it's a
   prompt-copy change, not a wire-level transform.

Two layers rather than one because the three-sentence AND-condition
means a classifier tightening could start firing on either remaining
leg alone.

Fixes #65365
2026-09-01 16:27:20 +05:30
Teknium 530aa7b10f test(cache): fix stale Studio construction-order paragraph in the bridge witness docstring
The module intro still described the pre-b449d7c1 order (row first, agent
second); _reply_scope and numbered item 2 already state the correct one.
2026-09-01 02:14:35 -07:00
joaomarcos 41ec9d3591 fix(cache,tests): durable generation docstring + studio-bridge order (PR #98811, refs #96811)
agent/prompt_cache_scope.py: durable per-(source,session_key) monotonic generation; list _RESET_END_REASONS; note never garbage-collected. tests/agent/test_studio_bridge_affinity.py: build AIAgent before db.create_session; docstring corrected.
2026-09-01 02:14:35 -07:00
joaomarcos 5abba03a99 test(cache): pin the affinity contract on Studio's real bridge construction
The reproduction reported on #96811 is Hermes Studio's group chat, and it is
the one half of the issue that cannot be closed from inside this repository.
This states why in executable form, and pins the contract the host adoption
depends on so a later refactor cannot quietly break it.

Root cause, traced to the host. Studio reaches Hermes as a LIBRARY, not
through the gateway. Its bridge mints groupRuntimeSessionId(room, profile,
name) -- a gc_run_ prefix truncated to 96 characters plus a fresh UUID4 hex --
for every reply, writes the session row itself, and constructs AIAgent(...)
with platform / session_id / session_db and no routing identity of any kind:
no gateway_session_key, no chat_id, no user_id, no parent_session_id. So the
declared scope is unreachable, the row is a lineage root, and
resolve_prompt_cache_scope() correctly falls through to the physical id, which
moves on every reply. No stable carrier crosses the boundary. Hermes must not
recover one from the id's syntax -- that is #79017's collision class, and the
two negative controls merged via #97704 exist to keep it from trying.

What does cross the boundary is a value Studio already has:
groupBridgeSessionId(room, profile, name, sessionSeed, runtimeConfig) is
stable for one conversation in one room, already carries the room, profile,
agent name and the room-owned sessionSeed, and is already hashed and
length-bounded. Passing it as gateway_session_key is the entire adoption, and
it is a Studio-side change; this PR keeps Refs #96811 for exactly that reason.

What this suite adds is the half that IS reachable here. The existing suites
simulate Studio's id SHAPE against synthetic agent doubles; none of them runs
the host's construction path, so nothing today would fail if that path stopped
honouring a declaration. These tests build a real AIAgent in the bridge's own
order -- row written first, agent second, new physical id per reply -- and
state what the adoption buys:

- three replies with distinct physical ids hold ONE affinity scope, and the
  scope carries neither the room nor the member name;
- two members of one room, and two rooms, never share a bucket;
- a new sessionSeed rotates the scope, which is Studio's own conversation
  boundary and needs no reset observed by Hermes;
- equal keys under different row sources never collapse, because the row's
  source -- not agent.platform -- is the identity the peer queries match on;
- a tool child and a background-review fork on the same declared key keep
  their own scope, so #79161 survives this construction path too;
- and an undeclared bridge is byte-identical: per-reply scope, and a
  compression rotation still walks its lineage.

Against origin/main the two declared-conversation assertions fail --
"assert 3 == 1", and the raw gc_run_ id leaking as the routing key -- which is
the reproduction; the remaining eight pass there and here, because they
describe behaviour that must not change.

Reported by @cervantesh, whose re-check of Studio main@86d0c95375 located the
missing consumer on the real path and asked for exactly this witness.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos d63e5d8a10 fix(cache): source-qualify the peer identity and gate the declared bind
Both blockers from @andrexibiza's review of 28a2d7f0ee.

1. The generation lookup was not in the same identity domain as recovery.
   latest_conversation_boundary() selected on session_key alone, while
   _declared_conversation_session() is qualified by (source, session_key).
   X-Hermes-Session-Key accepts any authenticated caller-supplied string, so an
   API conversation may legally carry the same key as a Telegram row in one
   database -- a /new over there rotated this conversation's gwk_ generation
   while recovery correctly refused to cross the same line, moving the affinity
   identity out from under a physical identity that had not moved.

   The boundary read now takes (session_key, source), and the carrier is
   'source|key|generation' rather than 'key|generation' -- keying on the string
   alone would also collapse two same-key conversations from different sources
   onto one routing key, since this value leaves the process verbatim as
   OpenRouter's sticky session_id and xAI's x-grok-conv-id. The source comes
   from the agent's own session row, falling back to the platform the row will
   be created with before it lands.

2. The declared key's stated lower precedence did not survive settlement. Both
   handlers let stored_session_id / an explicit body session_id win, then called
   _bind_declared_conversation() unconditionally. record_gateway_session_peer()
   does SET session_key = ? across compression ancestors, so a request carrying
   conversation A's chain plus header key B silently rebound A to B: A could no
   longer be recovered by its own key, and B recovered A's session.

   Recording is now gated on the declared key having actually selected or
   minted the session, on both paths. Behind that gate the bind itself refuses
   to overwrite a row already bound to a different key, so a future caller
   cannot reintroduce the same defect by opting in wrongly.

test_declaration_outranks_the_lineage_root asserted the pre-qualification
contract by comparing a DB-backed agent against a DB-less one; it now makes the
stronger statement it was written for -- one declared conversation reached
through two different physical ids on the same peer.

Refs #96811

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
joaomarcos e5bce4df4b fix(cache): make the conversation generation survive a backwards clock
Self-review of the generation marker. MAX(ended_at) alone is wall-clock: an
NTP correction between two resets writes a SMALLER boundary, MAX keeps
returning the older one, and the next conversation silently reuses the
previous generation -- two conversations on one routing key, which is the
defect this PR exists to remove.

latest_conversation_boundary now returns (count, latest_ended_at) and the
marker is 'count:ended_at'. The two halves fail under different conditions --
a backwards clock defeats the timestamp, retention pruning of an old ended row
decrements the count -- so a generation repeats only if both happen at once.
The pair is deliberately biased toward changing: a spurious change costs one
cold prompt-cache bucket, a repeat would merge two conversations.

Pinned by test_a_backwards_clock_does_not_reuse_a_generation, which rewrites
the second boundary to land before the first and asserts three conversations
still resolve to three distinct scopes.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos d7995bffaf fix(cache): qualify the declared key with the conversation generation
The declared key is a per-CHAT identifier and outlives the conversation it
names: reset_session() mints a fresh physical id on /new but keeps the key, and
the idle/daily/suspended policy resets do the same. Hashing the key alone
therefore mapped the conversation before a reset and the one after it onto one
gwk_ scope -- the lifecycle violation @cervantesh raised on #97158 and
@kshitijk4poor reproduced on #97709.

No counter is introduced. The generation that must rotate is already durable:
every one of those boundaries closes the outgoing row with an
_RESET_END_REASONS end_reason, so SessionDB.latest_conversation_boundary reads
the most recent one and declared_conversation_scope hashes 'key|generation'.

That makes the carrier stable across a host's per-response physical ids -- a
host that never resets writes no boundary, so every reply hashes the same value
-- while rotating on every conversation replacement, /new and the policy
auto-resets alike. ended_at only moves forward, so a retired generation can
never be reused: no ABA.

It also cannot drift from the rest of the codebase's notion of a conversation
boundary, because find_latest_gateway_session_for_peer fences on the same set.

The read is on the memoized resolution path, not per API call, and both lookups
fail closed: an unqualified key would span a /new, so a DB error degrades to the
physical-id scope. A SessionDB without the lookup keeps the previous behaviour.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
kshitijk4poor 29d4c0ebfd refactor(compression): extract preflight seed predicate onto ContextCompressor
Address review follow-ups on the seed fix:

- Move the 'seed only from the 0 state' guard from an inline block in
  build_turn_context() into
  ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
  the predicate with the rest of the speculative-seed lifecycle
  (snapshot_preflight_display_tokens /
  rollback_interrupted_preflight_display_tokens). Callers now use the
  method via a getattr guard so test doubles and external context
  engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
  described the old >=0 guard ('treats any negative value as no real
  usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
  method on the compressor fixture, eliminating mirror-drift risk (the
  helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
  low until their next report) in the method docstring.
2026-09-01 14:33:23 +05:30
Turgut Kural 80b836fa9e fix(compression): preflight display-seed must not overwrite real provider usage
The preflight rough-estimate seed used 'last >= 0' semantics, so any
provider that reports real prompt_tokens got its reading replaced by
the schema/reasoning-inflated rough estimate whenever the estimate was
larger. last_prompt_tokens feeds both the CLI context meter (cli.py)
and the post-response compression gate (conversation_loop.py), so one
seed made the bar jump to an inflated number AND could push the
real-usage gate over threshold on estimator noise.

Observed in a production CLI session on a 1M-token window with a
reasoning-heavy history: the status bar showed ~492K real provider
prompt tokens; the next turn's preflight estimated ~685K (rough
estimates can inflate 1.4-2.5x on reasoning-heavy sessions, #81481)
and seeded it into last_prompt_tokens; compression then fired at ~69%
of the window while real usage was ~49%.

Policy change: a real provider reading (>0) always wins. The seed now
only fills the 0 state ('no reading yet'), keeping the status bar live
for usage-less providers (#34282's motivation); -1 remains protected
as the post-compression sentinel (#36718).

Updates TestPreflightSentinelGuard to encode the new policy and adds a
regression test with the measured 492K/685K numbers.
2026-09-01 14:33:23 +05:30
Brooklyn Nicholson 5c6dbe22c3 fix(compression): take reasoning_content when the summarizer leaves content empty
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30