Commit Graph

3725 Commits

Author SHA1 Message Date
Uttkarsh Tiwari 36ae32524c fix(compression): preserve terminal lifecycle for lock skips 2026-08-27 12:42:38 +05:30
Uttkarsh Tiwari 7a21bfe68a fix(compression): suppress duplicate completion notices 2026-08-27 12:42:38 +05:30
lesseradmin 1341dfbd12 fix(compaction): exclude operational notifications from tail anchor and auto-focus (#92703)
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:

- _is_actionable_user_turn (tail anchor) only checked role/content, so a
  notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
  operational notices leaked into the compact focus hint.

Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.

Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.

Fixes #92703
2026-08-26 21:38:20 -07:00
Teknium f1d05ce7d8 fix(browser): real-profile snapshot is a first-class secret store + preserve channel identity
Addresses two P1 review blockers (kshitij / @kxee) on the real-profile feature:

Credential-store lifecycle for ~/.hermes/browser-profile/ (copied Cookies/
Login Data):
- exclude the singular 'browser-profile' dir from backup AND import
  (_EXCLUDED_DIRS drives both) — was silently archiving cookies/logins
- add a browser-profile/ directory-PREFIX read-deny to agent/file_safety.py,
  same class as auth.json / mcp-tokens
- secure the snapshot dir through the canonical hermes_cli.config._secure_dir
  (honors managed/NixOS group-share + HERMES_UID/GID), not a bespoke chmod

Channel identity (#95549 invariant — never normalize Beta/Dev/Canary to
stable, which would drive a different account's profile):
- detect recognized pre-release channels FIRST (Win ProgIds, macOS bundle ids,
  Linux .desktop) and return UNSUPPORTED_CHANNEL
- macOS bundle match is now EXACT (was startswith); Linux/Win channel-before-
  stable ordering; real_profile_data_dir/chromium_executable reject the sentinel
- _real_profile_cdp fails closed with a channel-specific message, never snapshots

Tests: channel-not-normalized (linux/darwin/windows), wrong-principal fail-closed,
backup exclusion, read-guard block/allow, snapshot dir secured. 187 browser +
222 backup/file_safety pass. Live re-verified: real Gmail inbox still loads.
2026-08-26 19:25:33 -07:00
Alexander Prendota 613164dadc fix(acp): key the ACP runtime exclusions on the scheme, not on one vendor
An ACP client talks to a CLI over subprocess stdio: it returns a plain
completion object rather than an iterable stream, and it does not implement
the Responses API surface. Both exclusions spelled out `acp://copilot`, so the
next ACP client silently inherited the wrong defaults — a Responses upgrade
its shim cannot serve, and a streaming call that tries to iterate a
`SimpleNamespace`.

Match on the `acp://` scheme instead. `acp+tcp://` was already handled this
way; copilot-acp's behaviour is unchanged, and the new tests pin that a
non-ACP URL still upgrades, so this is not a blanket opt-out.
2026-08-26 10:10:11 -07:00
Alexander Prendota 37fd61d13b fix(background-review): skip the fork when the provider can't emit tool calls
The review fork's entire job is to emit `memory` / `skill_manage` tool calls,
and by default it inherits the parent's live runtime. A provider that IS an
autonomous agent reaches Hermes through a client shim; if that shim cannot
carry Hermes tool calls back, the fork is a guaranteed no-op that still pays
for a full agent spawn — a whole CLI process, sometimes a JVM — on every
review cadence.

A client declares `SUPPORTS_HERMES_TOOL_CALLS = False` and the fork is
skipped with a warning naming the `auxiliary.background_review.{provider,model}`
override that routes the review to a normal model instead. Anything that says
nothing is assumed capable, so ordinary providers are untouched.

The check runs before the thread-scoped silence so the warning is not
swallowed, and only resolves the review runtime once the cheap capability
test has already failed, so the normal path does not resolve it twice.
2026-08-26 10:10:11 -07:00
Alexander Prendota 07200e9cd6 feat(agent): fold an agent-as-provider's own tool work back into the turn
Most providers are models: they ask Hermes to run a tool and Hermes runs it,
so the transcript and the loop's counters see every tool iteration. Some
providers are agents — an ACP CLI behind a client shim, or the codex
app-server, which already takes an analogous path in `agent/codex_runtime.py`.
They execute their own read/edit/execute tools inside their own session, and
by the time Hermes sees the response that work is done.

Those calls must never come back as pending `tool_calls` — Hermes would
re-run finished work. But summarising them into `reasoning` blinds two
subsystems:

- the self-improvement loop, which distils memories and skills by replaying
  `messages`; a one-line activity feed teaches it nothing;
- the skill-review nudge, whose `_iters_since_skill` counter only moves on
  Hermes tool iterations, of which there are none.

So a client may hand both back on the completion object —
`hermes_projected_messages` (completed assistant(tool_calls) + tool(result)
rows) and `hermes_provider_tool_iterations` — and
`splice_provider_projection` applies them. Rows go through `append_message`
like every other live-transcript append, so they carry a timestamp and
persist the same way the codex projection path's rows do.

The splice is append-only, sits before this turn's assistant message so the
order reads call -> result -> answer, and is a no-op for every client that
sets neither attribute, i.e. every ordinary OpenAI-compatible provider.
Garbage attribute values are tolerated rather than allowed to break the turn.
2026-08-26 10:10:11 -07:00
Alexander Prendota 083c5920e5 refactor(acp): share one OpenAI bridge between the ACP clients
ACP has no OpenAI `tools`/`tool_calls` channel: a prompt is text and a
response is text plus the agent's own tool notifications. Hermes' agentic
surface — memory, todo, skill_manage — is dispatched from OpenAI-shaped
tool_calls, so on an ACP provider it only works if the schemas travel into
the prompt as text and the calls are parsed back out of the response text.

copilot-acp already carried that bridge as private module-level helpers.
Lift it verbatim into `agent/acp_openai_bridge.py` so every ACP client
shares one implementation of the wire contract instead of re-deriving it —
`agent/claude_code_acp_client.py` (#81375) is currently a third copy of the
same four functions, and each copy is a place the `<tool_call>` contract can
drift.

copilot-acp is migrated onto it as the in-tree consumer and loses 176 lines
of duplication; its prompt shape is unchanged, which the new tests pin.

Two things the shared version adds over the copy:

- `render_tool_bridge_sections(..., allowlist=)`. A CLI with no tools of its
  own forwards Hermes' whole toolset (copilot, unchanged: no allowlist). A
  CLI that *is* an autonomous agent must forward only Hermes' agent-level
  tools — re-offering the overlapping read/edit/execute ones makes Hermes
  re-run work the agent already finished.
- `StreamChunks`, a list subclass that keeps response-level attributes.
  Hermes reads provider extras off the object returned by
  `chat.completions.create`; the old plain-list return silently dropped them
  whenever a caller asked for `stream=True`.
2026-08-26 10:10:11 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
kshitijk4poor dcfdc8deec fix(deadline): document the inline-mark contract; pin the ordering invariants (Phase 3a salvage round)
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.

New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
  invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
  async completion test)
2026-08-26 17:47:22 +05:30
Ayush Nangia 9ee2744097 fix(deadline): Phase 3a review round — mark ordering, loop offload, annotation
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
  describes the state at timeout; a recycling cleanup never poisons the
  healed replacement)
- the async flavor offloads the mark off the event loop
  (asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
2026-08-26 17:47:22 +05:30
Ayush Nangia 7a3aaf0143 feat(deadline): SuspectableBackend protocol — mark timed-out backends suspect
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
2026-08-26 17:47:22 +05:30
Teknium 3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium 65605d4a7a fix(computer_use): pitch background-FIRST (not background-only) in schema, prompt block, and skill 2026-08-26 04:50:13 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
TonyRainforest fa210e5a96 fix(compressor): abort compression on empty-content provider degradation to prevent context loss (#94448)
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.

- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py

Fixes #94448
2026-08-26 13:02:27 +05:30
kshitijk4poor 5885289c47 refactor(codex): route transports _pair_ids through shared fc_ canonicalization
The reuse reviewer found a third copy of the fc_->call_ synthesis in
agent/transports/codex.py _pair_ids (item_id[3:] spelling, which is why
the len('fc_') grep missed it). All three sites now share
_canonical_call_id_from_fc(), keeping the pairing invariant in one place.
2026-08-26 12:58:35 +05:30
kshitijk4poor e3bb3e7f8c refactor(codex): extract shared fc_->call_ canonicalization helper
/simplify-code reuse+quality reviewers both flagged the byte-identical
fc_->call_ synthesis blocks in the assistant and tool-result branches as
a correctness coupling — the two sites MUST stay in lockstep or pairing
breaks. Extract _canonical_call_id_from_fc() and route both through it.
Mutation check: pairing regression test fails when the tool-branch call
is stubbed out, green after restore.
2026-08-26 12:58:35 +05:30
kshitijk4poor 635232ec4e fix(codex): canonicalize fc_-only tool-result ids to match the call side
The sweeper review on #49224 flagged that the assistant branch synthesizes
call_<suffix> from an fc_-only id while the tool-result branch kept the raw
fc_... string — so an oversized pair hashed to two DIFFERENT clamped
surrogates and the function_call_output arrived unmatched (HTTP 400).

Canonicalize the tool-result side to the same call_<suffix> before
clamping. Also fixes the pre-existing short-fc_ pairing mismatch
(call_short123 vs fc_short123). Regression test covers both lengths.
2026-08-26 12:58:35 +05:30
kshitijk4poor 31485d50ea fix(codex): sanitize replayed function_call.name to Responses API pattern (#31666)
A degenerate tool name stored in conversation history (dots, spaces,
unicode from an earlier model degeneration) bricks every subsequent
Codex Responses turn with a non-retryable HTTP 400:
  Invalid input[N].name: string does not match pattern '^[a-zA-Z0-9_-]+'

The 400 replays forever until the user manually starts a new session.

Add _sanitize_replayed_fn_name() — replaces invalid chars with '_'
(runs collapsed), degrades all-invalid names to 'fn' instead of empty
(an empty name would trade one 400 for a preflight ValueError).  Applied
at both replay sites: the chat-message converter and the preflight
choke-point.  Live tool-definition names are left untouched — they must
match the dispatch registry exactly.  Pairing is by call_id, so
renaming a replayed function_call is safe.

call_id overflow (the sibling half of #49224) was already fixed on main
by #73492 (_clamp_responses_call_id); this commit covers the remaining
invalid-name defect.

Credit: @Morad37 (#31678 — identified the bug, the replay sites, and
the regex contract), @lubosxyz (#49224 — replace-not-strip semantics
and 'fn' fallback to avoid the empty-name trap).

Fixes #31666
2026-08-26 12:58:35 +05:30
alt-glitch 62b2d78025 fix(tool-search): bridge batch barrier, listing truncation, source indexing (salvage #92693, part 1)
Four fixes to the tool-search deferral layer, split from PR #92693 (the
availability-cache staleness fix ships separately):

1. The parallel batch planner now peels the tool_call bridge wrapper and
   decides admission on the underlying tool — supports_parallel_tool_calls
   works again when deferral is active. Unparseable wrappers stay
   sequential barriers; bridged calls get exactly the admission the same
   call gets direct. tool_search/tool_describe lookups batch concurrently.
2. _short_desc no longer truncates listing lines at 'e.g.', hostnames, or
   version strings — a sentence terminator must be followed by whitespace.
3. BM25 indexes the source label (e.g. 'linear' for mcp-linear), so
   service-name queries reach tools whose own name omits the service; the
   dead 'mcp' prefix token is stripped.
4. Substring-fallback docstring corrected (token misses, not zero-IDF).

Salvaged from #92693 by @alt-glitch with authorship preserved.
2026-08-25 15:24:15 -07:00
Teknium 9de5460c12 feat(cron): acked failure signatures stop re-pinging — durable incidents + ack CLI (salvage #94692) (#95017)
* feat(cron): durable failure incidents with signature dedup and ack

Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.

- cron/incidents.py: lazily-created incident table (detected -> alerted ->
  reviewed -> closed lifecycle; closed is per-signature terminal), sha256
  signature dedup over job_id + normalized error, redacted/truncated error
  storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
  suppress the per-run failure ping when the exact signature is acked (both
  the normal failure path and the processing-raised retry path). Best-effort:
  an incident-store error never breaks the cron run or delivery. Streak nudge,
  alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
  `hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
  classification, lazy-schema, scheduler gating, and CLI coverage.

Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.

* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome

Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
  lives in INCIDENT_STATES so future slices can add states without a
  table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
  delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
  delivery outcome (registered in cron_health monitoring) instead of
  the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
  remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.

---------

Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>
2026-08-25 14:02:40 -07:00
Teknium 4032a15ad0 refactor(prompt): remove the ~1.2K-token Nous Subscription block from the system prompt (#95005)
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-08-25 12:59:29 -07:00
fangliquanflq 7a67bd07a7 feat(docker): support shared container identities 2026-08-25 04:00:27 -07:00
Teknium 1ee524f77d fix(memory): bind the checkpoint gate to post-turn micro-compaction too
Independent review caught a compaction authority the gate missed:
post-turn micro-compaction (turn_finalizer -> _micro_compact) absorbs the
oldest exchanges into a rolling summary with no pre-compress checkpoint
hook in its path, and both compression.checkpoint_required and
compression.micro_compact could be enabled together — assistant evidence
could vanish into a summary the checkpoint filter later excludes, without
ever reaching the durable provider.

- agent_init: checkpoint_required forces micro-compaction off (warned),
  mirroring the native-compaction suppression
- turn_finalizer: defense-in-depth guard at the call site (attribute is
  plain mutable state a future path could flip on a live agent)
- behavioral regression test with a sabotage control (gate off proves the
  harness reaches the call site; gate armed proves zero calls)
- docs + config example mention the suppression; stale v1 test header fixed
2026-08-25 03:55:55 -07:00
Teknium 9e551d2931 refactor(memory): renumber checkpoint API — v1 is the implicit historical contract, v2 opts into fail-closed checkpoints
Per review: existing providers should not be retroactively re-versioned or
handed a changed payload. Version 1 is now the implicit historical
on_pre_compress() contract (best-effort, raw message list) that every
pre-existing provider is already on; the fail-closed checkpoint contract
becomes version 2. MemoryManager routes the raw transcript to v1 providers
unchanged and hands the host-normalized evidence list only to v2+ checkpoint
providers, so the plugin surface contract for shipped providers is
byte-identical with the gate off.
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 8cc379b528 fix(memory): bind the checkpoint gate to every compaction authority
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):

- codex app-server: in "native"/"off" auto-compaction mode (native is the
  default) Hermes preflight is skipped and the codex agent compacts its
  own thread inside run_turn() — the compress_context() rejection was
  unreachable. init_agent now refuses checkpoint_required together with
  api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
  testable guard), and run_codex_app_server_turn() fails closed as
  defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
  native_compaction_context_management() now returns None while the gate
  is armed, so context_management never goes on the wire and the
  checkpoint-aware Hermes compressor stays authoritative. The suppression
  is logged once per process, not silently applied.

Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 70d0b1fffb fix(memory): review follow-ups for the pre-compress checkpoint contract
Addresses the review on #93996:

- gateway: hygiene and manual /compress load the memory provider only when
  compression.checkpoint_required is enabled (skip_memory=not required).
  The historical fast path — no provider init, no best-effort hook — is
  back for everyone who did not opt in, so default behavior is truly
  unchanged.
- conversation_compression: assistant messages carrying both prose and
  tool_calls keep their prose in the checkpoint evidence (the tool_calls
  payload is stripped, the original message is not mutated); pure
  tool-call wrappers without prose are still dropped.
- tests: legacy-database regression proving the _compressed_summary column
  is added by the declarative _reconcile_columns() path on a plain reopen
  (no version-gated migration needed — append_message works right after),
  plus coverage for the prose-preserving filter.
- docs: providers must implement idempotent, content-keyed checkpoint
  writes — a fail-closed block means the next attempt re-runs
  on_pre_compress over largely the same transcript.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 1104ffe0b9 feat(memory): opt-in fail-closed pre-compress checkpoint contract (API v1)
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.

This adds an opt-in, provider-agnostic checkpoint contract:

- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
  by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
  historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
  on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
  failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
  (default false, documented in cli-config.yaml.example). When enabled,
  compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
  uncompressed transcript is preserved) unless a checkpoint-capable provider
  confirms the durable checkpoint. Providers receive normalized direct
  user/assistant evidence: tool rows, system messages, tool-call wrappers,
  and prior compaction summaries are filtered host-side into one stable
  contract. codex_app_server compaction is rejected under the gate because
  it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
  migration via _reconcile_columns) so summary provenance survives process
  restarts; only the resume model history carries the marker, keeping
  get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
  (skip_memory=False) so a required checkpoint also guards those rewrites.

The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.

Refs #93986
2026-08-25 03:55:55 -07:00
Teknium 76e306c458 refactor(tools): remove expired BFL FLUX 3 promo core tools (migration v39); FLUX 3 stays via video_gen/FAL for subscribers (#94599)
* refactor(tools): remove expired bfl_flux3_* promo tools; FLUX 3 rides the video_gen provider surface

* test: relay-cutover migration asserts >= v38, not the version literal
2026-08-25 02:45:10 -07:00
Mark McDonald eaf6545ab4 docs(gemini): update to use latest gemini models 2026-08-25 13:14:30 +05:30
Guilherme Artiles 610e2e02fb fix(curator): tell the background reviewer to read before it writes
`_background_review_read_before_write_guard` refuses a background-review
`skill_manage` write when the target file was not loaded via `skill_view` in
the same review turn (patch, edit, write_file over an existing file,
remove_file).

`CURATOR_REVIEW_PROMPT` never says so. It lists `skill_view` only under "read
the current landscape", so the reviewer goes straight to the write and every
mutation is refused. The failure is silent from the outside: the curator run
completes, writes nothing, and reads like a pass that simply found nothing to
consolidate. On our deployment that was 32 of 32 attempted writes rejected over
48h before anyone read the logs.

This adds the missing instruction to the toolset block, plus a test that fails
if a future guarded action is added to `skill_manager_tool` without being named
in the prompt — the guard and the prompt have to drift together or not at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 00:31:14 -07:00
Teknium a70d2ffce5 fix(background-review): teach review prompts the enforced read-before-write handshake
The skill_manage guard (added in #55906) refuses any patch/edit of an
existing SKILL.md, or overwrite/removal of an existing support file,
unless the exact target was loaded via skill_view during the review.
Neither _SKILL_REVIEW_PROMPT nor _COMBINED_REVIEW_PROMPT ever mentioned
this, so models routinely issued the write without the pre-read, got
refused, and burned review iterations (#62397).

Both prompts now carry a Read-before-write section scoped to the
guard's actual contract: existing targets only, exact-path pre-read for
support files, transcript quotes don't count, new skills/new support
files exempt, and a bounded one-view-one-retry recovery instead of a
loop. Direction follows #60331 by @kkwills13 with the scope corrections
requested in review (existing-target-only wording, no delete claim,
bounded retry, contract tests for both prompt variants).

Fixes #62397.
2026-08-25 00:18:35 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
Teknium 9cce872505 docs: note httpcore pin dependency in _enable_happy_eyeballs 2026-08-24 21:46:05 -07:00
Nathan Shan d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
Teknium 8172be0e8e fix(memory): log when a configured provider's tools are gated off by toolset config
Follow-up to the cherry-picked gate-parity fix: the silent 'return 0'
in inject_memory_provider_tools made #81014 undiagnosable — a
configured provider looked half-on with no hint which config key
suppressed its tools. Now an INFO line names the withheld providers
and the gating keys.
2026-08-24 21:45:30 -07:00
Chen Jin b45b028573 fix(agent): gate memory provider system_prompt_block on toolset config (#81014)
The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provider_tools_enabled()` via platform_toolsets or
disabled_toolsets. Result: the agent received instructions to call
`mnemosyne_remember`, `mnemosyne_recall`, etc., that did not exist in
its tool surface.

Centralize the gating into `memory_provider_tools_exposed(agent)`, use
it from both `inject_memory_provider_tools` and the system prompt
assembly path, and add regression tests covering:
* memory toolset enabled -> both tools and prompt block exposed,
* memory in disabled_toolsets -> neither exposed,
* memory not in enabled_toolsets and not built-in -> neither exposed,
* the built-in "memory" tool present as an opt-in -> both exposed,
* parity between `inject_memory_provider_tools` and
  `memory_provider_tools_exposed`.
2026-08-24 21:45:30 -07:00
Teknium 3733e4aff5 fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept
Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).
2026-08-24 20:25:10 -07:00
Teknium a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
kshitij d2a095df41 Merge pull request #94184 from kshitijk4poor/fix/85125-3b-mcp-recovery
fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b)
2026-08-25 03:35:38 +05:30
fangliquanflq 1420176393 fix(lsp): abort diagnostics waits after transport death 2026-08-25 01:48:43 +05:30
fangliquanflq 2f506c2023 fix(lsp): retire clients when the protocol reader exits 2026-08-25 01:48:43 +05:30
kshitijk4poor 2f33833de8 fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b)
Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:

- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
  poisoned state never does I/O; the NEXT caller pays once for a health
  probe that clears the suspicion or forces a reconnect. A single
  teardown-vs-keepalive race or auth-lock corruption can no longer park
  a connection permanently — park stays reserved for genuinely
  exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
  reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
  a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
  in-flight RPC races a child-watcher task, so a dead subprocess fails
  the call immediately with a retryable timeout instead of riding out
  the full 300s. Deliberate teardown/reconnect also fails in-flight
  calls now instead of leaving them attached to a dying transport.

Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.

Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.

Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
2026-08-25 01:34:55 +05:30
Vignesh Ramesh a0795acc83 fix(codex): identify Hermes requests 2026-08-24 11:25:04 -07:00
kshitij 6e534df114 Merge pull request #93773 from kshitijk4poor/fix/codex-sdk-transform-bypass-93650-v2
fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
2026-08-24 16:11:18 +05:30
kshitij 4341cf1df7 fix(approval): clamp approvals.timeout at the config-read chokepoint (#83220)
Widen the salvaged gate-site clamp to the bug class. The gate min() from
the contributor PR capped the gate bound at 360s, which would break the
#79719 contract (gate must extend while a legitimate >360s approval
prompt is answerable) and left the sibling overflow sites live: the CLI
prompt thread.join, the gateway poll deadline, and human_wait_ceiling
all consume the same config value.

Clamp once in _get_approval_timeout() via agent.deadline.MAX_SAFE_TIMEOUT_S
(1 year - semantically unbounded, platform-safe). The gate keeps its
approvals.timeout-tracking behavior above 360s; 7 regression tests pin
lock-acquire/thread-join safety and the gate-extension contract.

Bilateral E2E: with approvals.timeout=1e20 in a real config.yaml, main
crashes every consumer with OverflowError; this branch survives all 5
probes.
2026-08-24 16:11:03 +05:30
RelaxJonh 0dd0f6e648 fix(agent): clamp authorization gate lock timeout to prevent OverflowError on macOS
The _ConcurrentToolAuthorizationGate uses threading.Lock.acquire(timeout=...)
where timeout comes from human_wait_ceiling() (approvals.timeout + 60s).
When approvals.timeout is set very large (e.g. 999999999999 to effectively
disable timeouts), this overflows macOS timespec and raises:
  OverflowError: timestamp out of range for platform time_t

The gate only needs to serialize parallel dispatch — it should never wait
longer than the wedged-holder bound (_AUTHORIZATION_GATE_LOCK_TIMEOUT_S =
360s). Clamp the return value with min() so unbounded approvals.timeout
cannot overflow the lock acquire.

Fixes #83220
2026-08-24 16:11:03 +05:30
Teknium ec5e369fe6 Merge pull request #93784 from NousResearch/salv/81234-retry-carrier
fix: /retry and /undo no longer replay an older message after compaction (#81233, salvage #81234)
2026-08-24 03:31:29 -07:00