Commit Graph

28271 Commits

Author SHA1 Message Date
Brooklyn Nicholson 72cf8d1fac feat(desktop): export, import, rename and delete a board from the switcher
The board switcher could create and configure boards but not move,
rename, or remove one. Rename technically existed, buried as a field
inside "Settings…", which is why it read as missing; it now has its own
entry and the settings dialog is left owning scope alone.

Delete archives rather than erases — the board's directory moves to
boards/_archived/ and the toast names the path — and never appears for
`default`, which the backend refuses to remove.

The three dialogs had grown three copies of the same shell, the same
"invalidate the list and close" mutation tail, and the same name field,
so those are shared now instead of parallel-implemented.
2026-08-28 22:38:01 -05:00
Brooklyn Nicholson c57f8ad4e1 feat(desktop): PluginOs gains native save/open file pickers
Plugins had no sanctioned way to ask for a file path — the OS door
carried notify, openExternal, revealPath and writeClipboard, so anything
needing a dialog had to reach around the SDK for window.hermesDesktop.

pickSavePath and pickOpenPath wrap the existing selectSavePath /
selectPaths IPC with the door's usual contract: resolve null when the
bridge is missing or the user cancels, never throw at the plugin.
2026-08-28 22:38:01 -05:00
Brooklyn Nicholson 5e550838f7 feat(kanban): board export/import REST endpoints
POST /boards/{slug}/export and POST /boards/import, so the desktop and
dashboard can drive board transfer. Both exchange filesystem paths
rather than bytes, the same contract profile export/import uses and for
the same reason: the client runs its native save/open dialog on the
machine that hosts the backend, so a path is all either side needs, and
a board carrying a few hundred megabytes of attachments never has to
cross the renderer heap.

Also covers the pre-existing rename (PATCH) and delete endpoints, which
had no tests — including that delete archives to a restorable directory
and refuses to touch `default`.
2026-08-28 22:38:01 -05:00
Brooklyn Nicholson 3150e444b2 feat(kanban): export and import a whole board as a portable archive
`hermes kanban boards export|import` moves a board between machines:
tasks, comments, links, history, and attachments in one .tar.gz.

Two things make this more than a tar of the board directory. The
database is live — kanban runs in WAL mode, so a filesystem copy loses
whatever still sits in the -wal sidecar and tears if the dispatcher
commits mid-copy; export goes through SQLite's online-backup API
instead. And rows carry machine-local state: claims, worker PIDs,
absolute workspace and attachment paths, session ids, and the gateway
chat ids subscribed to task events. Shipping those verbatim is how an
imported board arrives holding a claim owned by a process on someone
else's laptop, or starts pushing task events into a stranger's Telegram
thread. Everything machine-local is stripped on export and re-stripped
on import, since an archive is untrusted input.

Imports always land as a NEW board, auto-suffixing the slug on
collision, so an import can never merge into or overwrite a board that
is already there. Tasks whose workspace was a directory or git worktree
on the source machine are parked in triage rather than left for the
dispatcher to claim and burn into the failure breaker.
2026-08-28 22:38:01 -05:00
Brooklyn Nicholson 110ecd238e refactor(profiles): extract the safe tar.gz primitives into archive_safe
Profile export/import owns the only hardened tar handling in the tree:
GNU-format writing (PAX fractional mtimes make macOS Archive Utility
throw "Error 94"), plus an extractor that rejects absolute paths, `..`
components, and non-regular members.

Kanban board transfer needs exactly that, and a second copy is how the
weaker of two extractors eventually ships. Move the four helpers to
hermes_cli/archive_safe and point profiles at them; no behavior change
beyond dropping a provably-unreachable fallback in the root-listing
helper, whose condition is a strict subset of the comprehension above it.
2026-08-28 22:38:01 -05:00
Gille 9a1eef7a29 fix(tools): narrow MCP OAuth lock scope 2026-08-28 20:16:26 -07:00
rob-maron f7c79efbac add tencent/hy4-preview to model pickers 2026-08-28 19:53:06 -07:00
Teknium 62e8126c69 fix(skill_manage): batch failure results carry the failing op's teaching payload (file_preview, hints)
Live A/B eval (old flat vs operations[] on qwen3.8-27b / gpt-5.6-terra /
claude-sonnet-5) caught a real regression: on a fuzzy-match miss the flat
path returns file_preview so the model can self-correct, but the batch
wrapper rebuilt the error dict and dropped every field except error/
failed_index. Sonnet, recovering blind, probed the file by writing and
reverting placeholder patches for 8+ turns (50k tokens vs 15k on the flat
arm). Batch failures now merge through all non-error fields from the
failing op's result.
2026-08-28 19:38:21 -07:00
hermes-seaeye[bot] ac6c8028e0 fmt(js): npm run fix on merge (#97517)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-28 23:47:58 +00:00
Brooklyn Nicholson 2a36a71578 fix(desktop): stop the project tree strobing while it re-probes an unreadable root
An unreadable root self-heals on a 3s timer, so the probe runs for as long as
the pane is open. Every forced reload cleared `rootError`, emptied `data` and
dropped `resolvedCwd` before reading, so each tick rendered "unreadable" →
blank → "unreadable" and flickered the header's project name with it. A local
ENOENT resolves well inside the 180ms skeleton delay, so the gap paints as a
bare blank frame rather than a loading state.

Re-reading the same root now probes underneath what is on screen; only a
different root, or the same path from a different backend, clears first.
2026-08-28 18:42:51 -05:00
Brooklyn Nicholson 0401e08884 fix(desktop): don't claim a stranger's cwd as the selected session's workspace
An unnamed `session.info` was treated as describing whatever the pane had
selected. The gateway stamps `stored_session_id: session_key or ""`, so every
not-yet-persisted session emits one, and `broadcast_session_info` / the
approvals loop re-emit for every live session at once. An unscoped event
applies exactly when no session is active, so with nothing selected each of
those repointed `$currentCwd` and claimed it for the null selection — the file
tree, coding rail and statusbar painted a folder no selected conversation
owned, until the next `releaseWorkspaceCwdOwner` dropped the claim and they
un-painted it.

Require the event to be bound to the pane's own runtime before an absent id
reads as the selection. The case the allowance exists for — a lazy session that
is the pane's runtime but is not persisted yet — still adopts and owns its cwd.
2026-08-28 18:42:51 -05:00
Teknium 3340bbbdad feat(a2a): client tools config-gated — disabled unless enabled (−561 tok/call on unconfigured installs) (#97421)
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)

* ci: retrigger after runner startup_failure on rerun attempt
2026-08-28 15:08:08 -07:00
Brooklyn Nicholson e60983a697 fix(desktop): let the HUD drag onto another monitor
Composer drag added renderer CSS-pixel deltas onto a window AppKit
clamps to the current display, so the bar could not follow the cursor
onto a second monitor (and drifted on mixed-DPI Windows). Track the OS
cursor in main and lift that clamp.

Co-authored-by: Biotrioo <biotrioo@protonmail.com>
2026-08-28 15:33:19 -05:00
brooklyn! a73b14c438 Merge pull request #96726 from NousResearch/bb/bot-mode-design-system 2026-08-28 15:06:08 -05:00
Mariano Nicolini 705a10850d refactor(nous): trim comments and drop unused code 2026-08-28 17:00:26 -03:00
Teknium dadbfd8990 refactor(patch): V4A mode gated to OpenAI-family mains — base schema is replace-only with real required (365 -> 195 for everyone else, -149; handler accepts both shapes from any model) (#97403) 2026-08-28 12:59:41 -07:00
Mohamed HAMMANE 6da0ae1cf5 fix(install): preserve project config for locked uv sync (#82446) 2026-08-28 12:36:15 -07:00
Teknium 25fcc8ad14 test(js): update node-engine-alignment fixtures for the 24.11 floor
The JS alignment suite pinned 24.0.0 as a supported Node; with the
engines arm raised to ^24.11.0 (babel 8 requires >=24.11), 24.0.0 and
24.10.x are now correctly rejected and 24.11+/24.18+ accepted.
2026-08-28 12:20:40 -07:00
Teknium 15eb5caf7a fix(install): Node 24.0–24.10 no longer passes the gates only to die at npm EBADENGINE
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.

- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
  package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
  (install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
  locked dependency's engines.node, and the installer gates must encode
  the same floors as the manifest — so the next babel-style floor bump
  turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
2026-08-28 12:20:40 -07:00
Jefferson Nunn c9fa2bba45 fix(install): tier-0 locked sync no longer trips over UV_NO_CONFIG
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:

  error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.

Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
2026-08-28 12:20:40 -07:00
Teknium 93de1d3430 vision_analyze diet + image routing: explicit aux vision backend becomes the de-facto route (reverses #29135) (#97339)
* refactor(vision_analyze): schema diet — routing mechanics removed (automatic; native path's own result teaches), region flow kept (~271 -> 181 tok/call, -33%)

* feat(image-routing): explicit auxiliary.vision backend is the de-facto image route — reverses #29135 (maintainer decision); native stays default when unset, image_input_mode:native stays absolute
2026-08-28 12:15:24 -07:00
Teknium 72874b0675 feat(skill_manage): operations[] is the call — each op names its skill; atomic with cross-skill rollback (#97295)
* feat(skill_manage): operations[] batch — several ops on one skill, atomic with rollback (memory-tool pattern); staged as ONE pending write under the approval gate

* refactor(skill_manage): operations[] IS the interface — single op = list of one (maintainer-directed); flat fields unadvertised handler compat; delete = sole-op routing

* guard(skill_manage): reject intra-batch same-file clobbers — double write/remove per path, full rewrite after an earlier SKILL.md edit; patch chains stay legal

* refactor(skill_manage): name-per-op — the call IS the operations array; cross-skill batches with all-touched-skills rollback

* guard(skill_manage): unify the intra-batch conflict guard — any destructive op on an already-touched file is rejected, with path normalization

Aggressive live testing found three holes in the two-part guard:
patch-then-write and patch-then-remove on the same supporting file
silently discarded the patch, and './references/x.md' //-style path
spellings slipped past the duplicate-write check. One rule now covers
the class: a destructive op (write_file/remove_file/full rewrite) on a
(skill, normalized-path) any earlier op touched is rejected pre-effect;
additive patches stay legal, so patch chains and write-then-patch still
work. Tests cover all three holes plus the pre-effect assertion.
2026-08-28 12:15:18 -07:00
Teknium 7d1c9aeab7 fix: log swallowed reclaim failures + pin ContextVar dispatch invariant (review follow-up for #91217) 2026-08-28 11:45:19 -07:00
69k4xmdfm2-blip fc5fdb8c2a fix(gateway): handoff is broken on multi-profile installs (wrong DB, wrong key, wrong bot)
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.

1. The watcher polls only the ROOT store.
   `_handoff_watcher` resolves `self._session_db` with no profile scope, which
   always yields the root `state.db`. But `/handoff` run under
   `hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
   store. Nothing ever reads it, so the CLI times out after 60s while the
   gateway is alive and connected. The watcher now iterates
   `[(None, None), *secondary_profiles]` and polls each inside
   `_profile_runtime_scope`.

2. The destination session key is built without the profile namespace.
   `_process_handoff` called `build_session_key()` with no `profile=`,
   producing `agent:main:...` while that profile's own adapter routes organic
   inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
   reads.

3. Delivery uses the PRIMARY profile's adapter and config.
   `self.adapters` holds only the default profile's adapters (secondaries live
   in `self._profile_adapters[name]`) and `self.config` only the default's home
   channel. A secondary profile's handoff was therefore sent by the wrong bot,
   to the wrong chat, while persisting the right session key and reporting
   `handoff_state='completed'` — a false positive that looks correct in the
   database and is wrong on the wire.

Two robustness fixes in the same path:

4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
   turn plus delivery; awaiting it inline meant one slow handoff stopped the
   watcher from even polling the other profiles. With the CLI's 60s deadline, a
   valid handoff could time out purely because another profile's was ahead of
   it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
   never claimed twice, and a bounded drain on shutdown.

5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
   one in-process dispatch, so a row still in that state at startup belongs to
   a gateway that died mid-dispatch. It can never reach a terminal state, and
   `request_handoff` refuses new requests unless the state is
   NULL/completed/failed — that session could never hand off again, silently.
   `reclaim_stale_running_handoffs()` now fails those rows once per store at
   watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
   may already have switched the session key and dispatched, so a blind retry
   risks double delivery.

Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.

Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
2026-08-28 11:45:19 -07:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Mariano Nicolini a51df3864e refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:01 -03:00
hermes-seaeye[bot] 9048530318 fmt(js): npm run fix on merge (#97358)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-28 18:38:38 +00:00
Brooklyn Nicholson a792d0794f fix(desktop): fill session-switch backfill in two frames instead of ten
FIRST_PAINT_BUDGET 20 + BACKFILL_STEP 60 prepended the rest of a 600-unit
page across ~10 visible commits. A 290-unit step keeps the interruptible
commits and removes the strobe.
2026-08-28 13:33:53 -05:00
Brooklyn Nicholson d229648511 fix(desktop): keep the session loader up while known history is empty
Brand-new drafts are empty on purpose. A routed session the list already
knows has messages must not drop the loader just because a runtime id is
bound — that is the blank frame during an unproven warm hold and a cold
switch.
2026-08-28 13:33:53 -05:00
Brooklyn Nicholson b6eb17d01c fix(desktop): hold unproven warm transcripts off the view
A compressed runtime cache is a legal tail, not display history. Publishing
it on session switch then replacing it with the persisted lineage is the
warm-path flicker. Gate that paint on persisted-display provenance and keep
the previous/empty view until REST authority lands.

Co-authored-by: xrbs00 <178640517+xrbs00@users.noreply.github.com>
2026-08-28 13:33:53 -05:00
yoniebans 00bbfc6900 docs(update): --yes help states the fork-upstream prompt is skipped, not accepted
The old text read as if --yes answers yes to every prompt. It accepts the config-migration and stash-restore prompts but skips the fork-upstream prompt without adding a remote (#97052 review); say so.
2026-08-28 13:33:56 -04:00
yoniebans be284cf5c0 fix(update): report when the official repo was not checked on the up-to-date path
Review follow-up on #97052 (helix4u): a fork with no upstream remote whose HEAD matches origin/main used to print plain "Already up to date!" under --yes even though official main was never consulted, so an unattended stale fork looked current. _sync_with_upstream_if_needed now returns whether the official upstream was actually checked, and the commit_count == 0 completion line says "Up to date with your fork (official repo not checked)." when it was not. Skip-as-decline semantics are unchanged: no prompt, no remote mutation, no decline marker. Caller-level regression test added for the fork + no-upstream + --yes + HEAD==origin/main path; helper tests now pin the return contract.
2026-08-28 13:33:56 -04:00
yoniebans b33fa127d2 fix(update): gate the fork-upstream prompt for --yes and non-tty runs
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes #60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
2026-08-28 13:33:56 -04:00
Brooklyn Nicholson 21cc469eb1 Merge remote-tracking branch 'origin/main' into bb/bot-mode-design-system
# Conflicts:
#	apps/desktop/src/plugins/hermes-bots/plugin.js
#	apps/desktop/src/plugins/hermes-bots/tests/group-chat.test.mjs
#	apps/desktop/src/sdk/profile-routing.test.ts
2026-08-28 12:05:11 -05:00
Brooklyn Nicholson 7584260842 docs(bots): name the room budget as the seam the limits PRs hook
GROUP_CHAT_MAX_ROUNDS and its four siblings carry over at the values
plugin.js shipped, so no rebase inherits a behavior change on top of a
rewrite. Making them configurable is live contributor work — #92213 for
per-room limits, #96842 for config plus a token budget — and both want the
same single seam, so say so where the constants are instead of adding a
config hook this PR has no consumer for.
2026-08-28 12:02:23 -05:00
Brooklyn Nicholson ebb20910fa fix(bots): scope a freshly created bot chat to the bots workspace
Creating a bot opened its chat with the workspace fields omitted, because
they were spread only when the caller passed a staleness probe — and the
create path is the one caller that has none. The composer reads that scope to
stand its branch rail down in a companion chat, so a just-created bot showed
the git rail until the next click reopened the same row scoped. Live-verified
on Linux against a real backend, and carried over from the old plugin.js
rather than introduced by the rewrite.

The probe answers whether to navigate. What the session IS never depended on
it: a freshly minted Bot Chat is a bot's chat no matter who asked for it. With
the gate gone all four openers in this file are the same call, so they become
one.
2026-08-28 12:02:23 -05:00
Brooklyn Nicholson a7db531a2a i18n(desktop): move the Bot Mode kickoff and the residual owned strings into the bundle
The intro a new bot is born with was the first line of its forever-chat and
shipped in English, so a non-English user met their bot in a foreign language
and the bot's reply followed the prompt's language. It now resolves through
the plugin bundle in all four locales. Attribution — the other half of #91827
— still needs the lazy or silent birth that issue proposes, since
prompt.submit IS the user-turn API; the intro itself stays, per AGENTS.md.

The rest is the class the review named rather than only the lines it cited:
every user-visible string the group room and the bot-scoped cron pane own now
lives in the bundle. Where core already ships the vocabulary in every locale —
Remove, weekday names, Daily/Hourly — the plugin reuses it instead of shipping
a second, worse translation. The frequency and weekday option lists stop being
module consts frozen at import, which pinned whichever locale loaded first.

Prompts addressed to a model, cron syntax, and the 'You' author marker stay
hardcoded on purpose, each for a stated reason, recorded in the bundle header.
2026-08-28 12:02:16 -05:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Teknium baa344dee7 refactor(process): schema diet — enum names the verbs, description keeps only non-obvious semantics; write-vs-submit trap teaching emphasized (306 -> 228 tok/call, -25%) (#97279) 2026-08-28 09:23:00 -07:00
Teknium 7b5e1911f8 refactor(todo): schema diet — item shape and merge semantics taught once, by the param schema (323 -> 232 tok/call, -28%) (#97257) 2026-08-28 08:58:12 -07:00
Teknium a9e72f1b58 refactor(read_file): schema diet + bundle anydoc 0.2.4 + typed NeedsOcrError/hosted-OCR wiring (#97195)
* refactor(read_file): capability-gate the anydoc format list; PDF coverage teaching lives in the response-time warning (426 -> 244/291 tok/call)

* feat(read_file): bundle firecrawl-anydoc 0.2.4 in core, typed NeedsOcrError handling, config-gated hosted OCR with local-OCR-first guidance

* refine(read_file): NEEDS-OCR warning hints at checking for an OCR skill without naming one; hosted_ocr knob unadvertised (maintainer-directed)

* simplify(read_file): drop the anydoc schema gate — bundled core dep makes absence a broken install, not a variant; formats stated unconditionally (263 tok/call)

* feat(read_file): PDF wording upgrades to 'scanned or text' when a trusted hosted-OCR route exists (direct key or explicit config; nous gateway excluded until Parse proxy works)

* simplify(read_file): FIRECRAWL_API_KEY is the ONLY hosted-OCR gate — nous gateway route removed (Parse proxy broken), config true no longer unlocks; false still disables
2026-08-28 08:46:11 -07:00
Mariano Nicolini 04647f15c8 fix(nous): only fall back to the reachable set when the overlap is empty
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.

Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:39:29 -03:00
Mariano Nicolini bafaac5e61 docs(nous): correct the subtract-only claim in the policy plan
The plan stated the policy set should only ever subtract from a list. That is
wrong when an allowlist names a model the curated manifest lacks, which empties
the picker instead of narrowing it — the behaviour fixed in 117e7fef88.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:26:30 -03:00
Mariano Nicolini 117e7fef88 fix(nous): surface allowed models the curated list does not carry
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.

When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.

Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:24:55 -03:00
Teknium 306db2776c test(codex): mid-turn compaction fixtures report realistic anchored usage
The usage anchor (#97206) now trusts provider-reported usage. These two
tests simulated a tool-heavy near-overflow turn while the shared fixture
reported a 12-token prompt — the anchored pressure check honestly
concluded no pressure. Give the scenario 18K anchored prompt tokens so
the tests pin the same compaction decision they always did.
2026-08-28 07:51:31 -07:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
Teknium c5b44e0756 chore: map contributor emails for TiberiuD and fedebyes 2026-08-28 07:51:23 -07:00
Teknium 5b31602c15 docs: reconcile positional pairing with shared _classify_tool_call_orphans (#97167) — classifier docstring reflects its remaining consumer; empty-id filter note updated 2026-08-28 07:51:23 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00