Commit Graph

32213 Commits

Author SHA1 Message Date
kshitijk4poor 193f05dec5 refactor(cron): one FIRE_CLAIM_TTL_SECONDS for claiming, one-shot re-arm, and stale-error recovery
The TTL was three copies of the literal 300 (claim_job_for_fire default,
rearm_oneshot, and the new stale-error guard). Hoisted so the three lanes
cannot drift; incident narrative in the guard comment cut to the WHY.
2026-09-07 00:39:44 +05:30
kshitijk4poor 37a0c1b95f chore: map r3x443 contributor email for #103740 salvage 2026-09-07 00:39:44 +05:30
r3x443 eb40bf060f fix(cron): skip stale-error re-arm while another process holds a live fire_claim
Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
2026-09-07 00:39:44 +05:30
kshitijk4poor 5f87ac74d7 refactor(gateway): verdict helpers live in run_shutdown; keep the early failure return; dedupe test setup
- _exit_with_failure_verdict / _resolve_gateway_exit_verdict move out of the
  run.py facade into run_shutdown.py, which already owns _restart_via_service
  and GATEWAY_SERVICE_RESTART_EXIT_CODE (no alias import needed).
- The running-shutdown tail keeps its early `return False` on a failure verdict,
  as on main, so a failure exit does not first drain cron/MCP; the helper still
  re-checks it for the startup-abort path.
- The three start_gateway tests share one _patch_aborted_startup helper.
2026-09-07 00:38:37 +05:30
kshitijk4poor 57238b8507 chore: map YuYigeng contributor email for #103208 salvage 2026-09-07 00:38:37 +05:30
YuYigeng e24047d8bd fix(gateway): unify shutdown exit verdicts
Adapt the shared exit-verdict resolver and service-restart fallback from JoaoMarcos44’s implementation in f77586ef7683c17d5890e63e38604a68d5ea11ce (#103248) for the earlier #103208 branch.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-07 00:38:37 +05:30
YuYigeng 51ddff0bb1 fix(gateway): restart after startup SIGTERM 2026-09-07 00:38:37 +05:30
Teknium 6d9c166455 Merge pull request #103529 from NousResearch/fix/terminal-auto-background
fix(terminal): an over-cap foreground timeout runs as a tracked background process instead of being refused (454 refusals in one run)
2026-09-06 12:06:04 -07:00
Teknium fdd140557b Merge pull request #103551 from NousResearch/fix/tool-output-caps
feat(file_tools): write_file flags a whole-file rewrite that mostly re-sends what is on disk (661 such writes, ≈$155, in one run)
2026-09-06 12:05:53 -07:00
Teknium cd658b6a46 Merge pull request #103492 from NousResearch/fix/approval-scanner-quoted-subshell
fix(approval): a grep inside "$(...)" no longer trips the hardline malformed block, and a quoted substitution body keeps its command boundaries (546 false blocks; review-found bypass closed)
2026-09-06 12:03:29 -07:00
Teknium 634109b7aa Merge pull request #103534 from NousResearch/fix/goal-judge-wait-on-subagents
fix(goal): the judge can WAIT on delegated subagents (4 of 5 nudges re-poked a waiting orchestrator; live 3/3 CONTINUE → 3/3 WAIT)
2026-09-06 12:03:15 -07:00
Teknium afb4e080f6 Merge pull request #103496 from NousResearch/fix/goal-judge-own-processes-only
fix(goal): the judge sees only its own session's background processes; pid/session waits expire after 30 min (parked 3h22m on a grandchild's poller)
2026-09-06 12:03:05 -07:00
Teknium d97f318717 Merge pull request #103526 from NousResearch/fix/nous-auth-stampede
fix(nous): adopt a same-account fresh key before expiry and start the keepalive in every process (620 hourly 401s → 0; review-found account takeover closed)
2026-09-06 12:02:58 -07:00
Teknium e5f8420be8 Merge pull request #103513 from NousResearch/fix/subagent-context-cap
fix(delegation): compression_threshold_tokens is opt-in (default off, children keep the 500K ratio trigger); validate the value
2026-09-06 12:02:49 -07:00
Teknium 3d5831fa59 Merge pull request #103486 from NousResearch/fix/nested-delegate-deadline-and-summary-budget
fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum
2026-09-06 12:02:43 -07:00
kshitijk4poor 2d77992750 fix(file-ops): reuse _kill_process_group_posix in the native rg runner
A bare os.killpg/signal.SIGKILL trips the Windows-footgun lane (the module is
imported on Windows even though the native lane never runs there). The local
environment already has the POSIX group killer with the TERM→KILL escalation
and setsid-escapee sweep; use it.
2026-09-07 00:28:46 +05:30
kshitijk4poor 9158bd8e0d refactor(file-ops): one _run_rg_bounded owns the native-vs-shell transport choice
Three call sites each repeated `if native: _run_rg_native(...) else: _exec(... | head -n N)`.
The choice now lives in _run_rg_bounded; callers pass the words, the bound, and the
one thing the native lane cannot express (a cd prefix → native_ok=False). The grep/find
pipeline keeps its explicit shell form because of the column cap.

Test file: one module-scoped LocalEnvironment instead of thirteen (~0.8 s each).
2026-09-07 00:28:46 +05:30
kshitijk4poor 4756a8115e refactor(file-ops): drop the _native_rg_enabled alias, build the files-lane rg command once
_native_rg_enabled was a pass-through to _native_read_enabled (a wrapper with
no behaviour); call the gate directly and note in its docstring that it
covers search too. The rg --files invocation was assembled twice (argv list
for the native lane, string for the shell lane) with room to drift; build the
command string once and hand it to either transport.
2026-09-07 00:28:46 +05:30
kshitijk4poor ef5a534c9a fix(file-ops): native rg runner honours deadline and /stop while rg is silent
The first version checked the deadline only after a line arrived, so an rg
that produced nothing for 60 s (huge tree, no hits yet) pinned the caller
past the timeout and ignored the interrupt flag that the shell path honours
via _wait_for_process. Drain on a daemon thread; the waiter owns deadline
(124) and interrupt (130) and kills the process group, so no rg or child
survives the return. Probe: silent 10 s process, timeout=2 → 2.0 s / 124;
interrupt at 0.5 s → 0.5 s / 130; zero stray processes afterwards.
2026-09-07 00:28:46 +05:30
kshitijk4poor 83063aeb87 perf(file-ops): run rg natively for search_files on local POSIX hosts
read_file already bypasses the backend shell on a local POSIX environment
(_read_file_native); search_files still paid two bash spawns per call — the
`test -e` existence probe and `set -o pipefail; rg ... | head -n N` — plus one
per zero-match probe. Measured on macOS against the repo's tools/ tree:
content search 85 ms → 15 ms, no-match search (three probes) 179 ms → 44 ms,
file-name search 73 ms → 12 ms; raw `rg` argv is ~15 ms, so the remainder was
transport.

Same gate and kill switch as reads (_native_read_enabled: LocalEnvironment,
not win32, HERMES_NATIVE_FILE_READ=0 disables). The argv builders and
_parse_search_output are unchanged and shared: _run_rg_native shlex-splits
the already-quoted words, streams stdout and stops after fetch_limit lines
like `head` would, and reports exit 0/1/2 (124 with partial output on
timeout) so the parser sees the shell contract. grep/find fallbacks, remote
backends, Windows and the multi-root cd form keep the shell path.

Two shell-observer tests in test_search_zero_match_and_multipath pin the
shell lane explicitly; they assert on command text, not behaviour.
2026-09-07 00:28:46 +05:30
kshitijk4poor 83467c28f7 test: drop the duplicate #90322 file — TestSearchHints already json.loads the truncated output 2026-09-07 00:20:11 +05:30
kshitijk4poor 8ed990e70f test: trim #90322 regression to one invariant, drop the hint-splitting workaround
The credentials test no longer has to strip trailing text before json.loads;
the new file keeps the single behaviour contract (truncated output round-trips
through json.loads and carries the next offset) — the existing TestSearchHints
cases already cover offset arithmetic.
2026-09-07 00:20:11 +05:30
liuhao1024 12ff3da3e4 test: update cross-tree search hint assertion to the structured _hint field
Follow-up to the pure-JSON search_files fix: the credential-filter test
in tests/agent/ also asserted the old appended '[Hint: ...]' text — the
local run covered tests/tools/ only, so CI slice 7 caught it. The _hint
field is present and correct in the CI failure output itself; only the
assertion syntax was stale.
2026-09-07 00:20:11 +05:30
liuhao1024 478e66475b fix(tools): keep truncated search_files output pure JSON
When results were truncated, search_files appended the pagination hint
as plain text after the serialized payload ("{...}\n\n[Hint: ...]"),
so the tool result was no longer parseable JSON — downstream tool-message
handling on providers strict about tool-content formatting could reject
or mishandle it, contributing to 400 upstream errors in sessions with
truncated search output (#90322).

Move the hint into the payload as a structured _hint field, matching the
existing _omitted/_warning side-channel convention in the same function.
The model-facing guidance (explicit next offset) is unchanged.

Fixes #90322
2026-09-07 00:20:11 +05:30
Gille 485aaf69b7 fix(local-runtime): find nvidia-smi in WSL driver path 2026-09-06 23:34:20 +05:30
jango ed406f929d fix(browser): make the real-profile attach daemon reapable (#100855, salvage #103284)
The `hermes-real-profile` agent-browser daemon (the attach lane for consented
real-profile browsing) ran with plain `_build_browser_env()`: no
`AGENT_BROWSER_SOCKET_DIR`, so it lived in agent-browser's default dir and no
reap path could see it. A wedged daemon + headless Chrome survived 47h across
two gateway restarts (#100855), and on macOS the genuine Chrome binary it held
made "Chrome won't open" for the user.

Give the attach lane the same contract every other lane already has:
`_prepare_session_socket_dir()` (per-session socket dir + `owner_pid` claim)
and `_agent_browser_command_env()`, and add the named dir to the orphan
reaper's scan. The existing `_reap_socket_dir` then applies its owner-liveness
and start-time-fingerprint rules unchanged; the daemon is listed as tracked so
the untracked-idle escape hatch never fires under a live user (per-task `rp_*`
sessions drive it over `--cdp`, so its own dir shows no activity). The daemon-side idle
timeout is NOT inherited: Chrome is launched by Hermes, not the daemon, so a
self-exiting daemon would leave Chrome holding the copy dir while the next
attach re-runs the snapshot overlay over it.

When a reaped daemon's Chrome (Hermes-launched, own process group) still
holds the copy dir, `_real_profile_cdp` re-attaches to it instead of running
the snapshot overlay over a live profile. DevToolsActivePort outlives a
crashed Chrome and its port can be recycled, so the file's browser id must
match `/json/version` before it is trusted; an attach failure on a live
Chrome fails closed rather than overlaying.

Tests: attach/get/close commands carry the reaper-visible socket dir and
owner_pid and no idle timeout; a dead-owner real-profile daemon is reaped by
`_reap_orphaned_browser_sessions`; a surviving Chrome is re-attached, never
overlaid (all red on main).
2026-09-06 23:31:05 +05:30
kshitijk4poor 57f05e2142 refactor(bot-mode): re-authorize message_agent on staged snapshots without copying the agent
The salvaged fix ran the injector against a copy.copy(agent) whose
tools/valid_tool_names pointed at the staged pair, wrapped in a blanket
try/except. A shallow copy of a live AIAgent (locks, DB handle, in-flight
attribute writes from the late-binding thread) is a workaround for the gate
mutating in place, and the except turned any failure into a silently
published snapshot WITHOUT message_agent.

Extract the gate as tools.bot_mode_dm.message_agent_authorized(agent) (the
same predicate ensure_message_agent_tool already used), and have the
snapshot builder append message_agent_tool_schema() to the staged list
directly when it passes. No copy, no swallow; same tests, same live
behaviour (compaction / between-turns / resume keep the tool; ordinary
sessions scrubbed).
2026-09-06 23:11:12 +05:30
jango b2b026dc22 fix(bot-mode): retain message_agent across tool rebuilds (#102864) 2026-09-06 23:11:12 +05:30
Teknium dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
Teknium 0390ace817 Merge pull request #103756 from NousResearch/feat/evals-postmortem-harness
feat(evals): post-mortem harness — forensics lanes, live A/Bs and review probes for the #102117 run fixes, reproducible on any state.db
2026-09-06 10:31:30 -07:00
kshitijk4poor 5f406d88ea test(cron): the run_job FakeFuture double answers done() like the real Future
run_job's finally now asks the worker Future whether it is still running before deciding who tears the session down; the heartbeat test's stand-in future lacked done().
2026-09-06 22:56:51 +05:30
kshitijk4poor b114641c88 refactor(state): fold the speculative-open retry into acquire's loop; keep real inode-swap tests
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
  the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
  type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
  those files carry no windows_only marker so they never run on Windows, and the monkeypatched
  predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
  db.close() already releases a registry-shared handle, so the described NameError leak only
  existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
  teardown; replacement not published before the last close settles) plus the raising-close
  settlement; comments say the WHY once.
2026-09-06 22:56:51 +05:30
kshitijk4poor 40488a4e54 refactor(cron): detached-worker teardown as a sibling module, no guards around a real Future
The deferral helpers from the previous commit were appended to the cron/scheduler.py facade and
carried ~70 lines of fallbacks (getattr/callable checks, result() waits, a teardown_registered
flag) for hypothetical Future doubles; the only producer is _cron_pool.submit, always a real
concurrent.futures.Future whose done()/add_done_callback() cannot raise. Also drops the
_finalize_cron_session_db passthrough. Behaviour unchanged; the test now asserts the
finalize+teardown contract on the real Future instead of a patched wrapper.
2026-09-06 22:56:51 +05:30
joaomarcos aad74f26f9 fix(state): coordinate SessionDB teardown with active writers
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.

Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.

Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.

The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.

Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.

Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.

Fixes #102827

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
2026-09-06 22:56:51 +05:30
Teknium 583fb3e200 evals(postmortem): cache concurrency probe, the instrument behind #104284/#104421 and api#227
N concurrent real AIAgent sessions on a growing tool loop; per-call cache_read /
cache_creation / response id / upstream / prefix shas; consecutive pairs classified
ideal / stuck / collapse. argparse (--provider nous|openrouter|anthropic, --wire,
--pin, --settle, --ttl, --model); wire defaults to what Hermes would pick. Summary
JSON carries bad pairs with both response ids. Smoke-run against Nous from the repo
path (picked chat via nous_api_mode; 6/6 ideal).
2026-09-06 10:23:38 -07:00
Teknium 1690a7a11f Merge remote-tracking branch 'origin/main' into feat/evals-postmortem-harness 2026-09-06 10:20:50 -07:00
kshitijk4poor 311b980bb6 fix(prefix-cache): drop the workspace pin at session boundaries; bind session cwd for /context
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).

The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.

Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
2026-09-06 22:45:31 +05:30
joaomarcos 0ed7acb051 perf(prefix-cache): pin session-start workspace snapshot across compaction rebuilds
- pin session-start workspace snapshot (_frozen_workspace_snapshot) on first build so dynamic git probes don't churn Tier 2 during compaction rebuilds in active coding sessions
- replay pinned snapshot across rebuilds when cwd matches; re-probe only on cwd switch
- honor coding_context invariant that workspace is a session-start snapshot, preventing prefix-cache divergence at offset ~4,662
- add invariant tests covering workspace snapshot pinning across git mutations and cwd transitions
- addresses upstream prompt divergence identified in #103326
2026-09-06 22:45:31 +05:30
Teknium 5e645791ac Merge pull request #104421 from NousResearch/feat/nous-anthropic-wire-auto
feat(nous): anthropic_wire=auto picks the session's wire from the first response's upstream (gated until GMI native is measured)
2026-09-06 10:04:56 -07:00
Teknium c5594ec4b3 fix(delegation): a crash mid-unit keeps the children that already finished
Since #104299 a background delegate_task call is split into completion units
(one per `group`, one per ungrouped task). A multi-child unit still joined on
all its children before anything was written durably, so an owner crash
between the first and last child lost the finished work and replayed the whole
unit as "outcome unknown" — the restart-granularity gap that #104233 (Xipong's
#76228/#76229 direction) solved with a second row per child.

Each finished child of a detached unit is now recorded on the unit's OWN row
(`record_unit_child` → result_json {results, partial}) as its future lands;
the real result overwrites it at finalize. `recover_abandoned_delegations`
replays recorded children with their real summaries and marks only the
unfinished ones unknown, naming the count. No new rows, no new consumer shape.
`task_indexes` is persisted so recovery knows a split unit's members.
2026-09-06 09:21:58 -07:00
Teknium 9a5a78187a fix(tui): post-turn control publish only when automation state exists (keeps plain-session event stream unchanged) 2026-09-06 09:21:45 -07:00
Teknium e27b8c5b92 test(desktop): trim session-control suites to behaviour contracts
session-control.test.tsx 1277→381 lines (32→8 cases): legacy-vs-structured
precedence, action dispatch, send continuation (idle + busy-queue), failure
banner, heartbeat countdown. store/session-control.test.ts 648→472 lines
(26→16): ordering/stale-token invariants, method-not-found downgrade, and a
gateway-switch wipe test replacing the removed rebind seam tests; the
setTimeout-spy 'no timer' test dropped.
2026-09-06 09:21:45 -07:00
Teknium efc30169c3 fix(desktop): queue goal-resume continuation when busy; wipe session controls on gateway switch
- session-control-goal.tsx: a 'send' dispatch against a busy session now
  parks the kickoff on the composer queue (the backend already resumed the
  goal) instead of reporting continuationFailed. Shared helper
  queueKickoffIfSessionBusy() extracted from slash.ts so both paths agree.
- gateway-switch.ts: wipeSessionListsForGatewaySwitch clears
  $sessionControlBySession (runtime-id keyed; new backend re-mints ids).
  Dead resetSessionControlAfterGatewayRebind removed; clearSessionControl
  now called from the session delete path beside clearQueuedPrompts.
- session-control.tsx: read/hydration failures use controlUnavailable copy,
  not actionFailed.
- i18n: heartbeatDueWaitingForIdle added to ja/ru/zh-hant; new
  continuationQueued/continuationBusy/controlUnavailable in all locales.
2026-09-06 09:21:45 -07:00
Teknium b0f87d7a7a fix(tui): publish session.control.update on every control mutation path; drop the Desktop-only heartbeat driver
- /goal and /loop are command.dispatch built-ins (never reach the slash worker), so the
  worker-branch publish never fired for the most common typed commands and the structured
  card stayed stale until the next turn. The publish helper now lives beside _snapshot_control
  in methods_session_control.py and is called from both slash paths plus the end of the turn
  tail — after the goal judge / loop tick evaluation, which mutate state AFTER message.complete.
- session.control skips the duplicate emit for dispatcher-backed actions (still exactly one
  update per mutation).
- Removed tui_gateway/desktop_heartbeat_driver.py, the entry/ws hooks and the hermes_cli lock
  rework: #104224 drives /heartbeat from the existing per-session poller for every TUI client.
- test_session_control.py: dropped presentation-shaped cases, added the publication invariants
  (both slash paths, unknown dispatch publishes nothing, real _run_prompt_submit turn publishes
  post-judge). 3/3 new tests fail on the contributor backend.
2026-09-06 09:21:45 -07:00
Jerry Gooch 45b7dabfb9 fix(desktop): run and surface Heartbeats reliably
(cherry picked from commit 77af78f1d8b7cb736d5d86e4fc7722449528a269)
2026-09-06 09:21:45 -07:00
Jerry Gooch 4f008c36bb fix(desktop): preserve session control action state
(cherry picked from commit 0efb410ed49ed6977352b313a639ff3f9d9375c1)
2026-09-06 09:21:45 -07:00
Jerry Gooch dffd8d62c2 feat(desktop): add session automation controls
(cherry picked from commit 8a61c2bcba8e4c7b3adcc172d4e5457721f76990)
2026-09-06 09:21:45 -07:00
Jerry Gooch bfddf556bf feat(desktop): hydrate structured session controls
(cherry picked from commit fec4bab6a191c474456956a860d86957d435552b)
2026-09-06 09:21:45 -07:00
Jerry Gooch 8cb2bcc8c1 feat(desktop): expose structured session controls
(cherry picked from commit 0dae0d9e7a36b57bc0489fbf0d6089de4f025d16)
2026-09-06 09:21:45 -07:00
Teknium 45646f4d09 feat(nous): anthropic_wire=auto decides a session's wire from its first response
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.

`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.

GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.

Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.

Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
2026-09-06 09:11:04 -07:00