The TTL was three copies of the literal 300 (claim_job_for_fire default,
rearm_oneshot, and the new stale-error guard). Hoisted so the three lanes
cannot drift; incident narrative in the guard comment cut to the WHY.
Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
- _exit_with_failure_verdict / _resolve_gateway_exit_verdict move out of the
run.py facade into run_shutdown.py, which already owns _restart_via_service
and GATEWAY_SERVICE_RESTART_EXIT_CODE (no alias import needed).
- The running-shutdown tail keeps its early `return False` on a failure verdict,
as on main, so a failure exit does not first drain cron/MCP; the helper still
re-checks it for the startup-abort path.
- The three start_gateway tests share one _patch_aborted_startup helper.
Adapt the shared exit-verdict resolver and service-restart fallback from JoaoMarcos44’s implementation in f77586ef7683c17d5890e63e38604a68d5ea11ce (#103248) for the earlier #103208 branch.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
fix(approval): a grep inside "$(...)" no longer trips the hardline malformed block, and a quoted substitution body keeps its command boundaries (546 false blocks; review-found bypass closed)
fix(nous): adopt a same-account fresh key before expiry and start the keepalive in every process (620 hourly 401s → 0; review-found account takeover closed)
fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum
A bare os.killpg/signal.SIGKILL trips the Windows-footgun lane (the module is
imported on Windows even though the native lane never runs there). The local
environment already has the POSIX group killer with the TERM→KILL escalation
and setsid-escapee sweep; use it.
Three call sites each repeated `if native: _run_rg_native(...) else: _exec(... | head -n N)`.
The choice now lives in _run_rg_bounded; callers pass the words, the bound, and the
one thing the native lane cannot express (a cd prefix → native_ok=False). The grep/find
pipeline keeps its explicit shell form because of the column cap.
Test file: one module-scoped LocalEnvironment instead of thirteen (~0.8 s each).
_native_rg_enabled was a pass-through to _native_read_enabled (a wrapper with
no behaviour); call the gate directly and note in its docstring that it
covers search too. The rg --files invocation was assembled twice (argv list
for the native lane, string for the shell lane) with room to drift; build the
command string once and hand it to either transport.
The first version checked the deadline only after a line arrived, so an rg
that produced nothing for 60 s (huge tree, no hits yet) pinned the caller
past the timeout and ignored the interrupt flag that the shell path honours
via _wait_for_process. Drain on a daemon thread; the waiter owns deadline
(124) and interrupt (130) and kills the process group, so no rg or child
survives the return. Probe: silent 10 s process, timeout=2 → 2.0 s / 124;
interrupt at 0.5 s → 0.5 s / 130; zero stray processes afterwards.
read_file already bypasses the backend shell on a local POSIX environment
(_read_file_native); search_files still paid two bash spawns per call — the
`test -e` existence probe and `set -o pipefail; rg ... | head -n N` — plus one
per zero-match probe. Measured on macOS against the repo's tools/ tree:
content search 85 ms → 15 ms, no-match search (three probes) 179 ms → 44 ms,
file-name search 73 ms → 12 ms; raw `rg` argv is ~15 ms, so the remainder was
transport.
Same gate and kill switch as reads (_native_read_enabled: LocalEnvironment,
not win32, HERMES_NATIVE_FILE_READ=0 disables). The argv builders and
_parse_search_output are unchanged and shared: _run_rg_native shlex-splits
the already-quoted words, streams stdout and stops after fetch_limit lines
like `head` would, and reports exit 0/1/2 (124 with partial output on
timeout) so the parser sees the shell contract. grep/find fallbacks, remote
backends, Windows and the multi-root cd form keep the shell path.
Two shell-observer tests in test_search_zero_match_and_multipath pin the
shell lane explicitly; they assert on command text, not behaviour.
The credentials test no longer has to strip trailing text before json.loads;
the new file keeps the single behaviour contract (truncated output round-trips
through json.loads and carries the next offset) — the existing TestSearchHints
cases already cover offset arithmetic.
Follow-up to the pure-JSON search_files fix: the credential-filter test
in tests/agent/ also asserted the old appended '[Hint: ...]' text — the
local run covered tests/tools/ only, so CI slice 7 caught it. The _hint
field is present and correct in the CI failure output itself; only the
assertion syntax was stale.
When results were truncated, search_files appended the pagination hint
as plain text after the serialized payload ("{...}\n\n[Hint: ...]"),
so the tool result was no longer parseable JSON — downstream tool-message
handling on providers strict about tool-content formatting could reject
or mishandle it, contributing to 400 upstream errors in sessions with
truncated search output (#90322).
Move the hint into the payload as a structured _hint field, matching the
existing _omitted/_warning side-channel convention in the same function.
The model-facing guidance (explicit next offset) is unchanged.
Fixes#90322
The `hermes-real-profile` agent-browser daemon (the attach lane for consented
real-profile browsing) ran with plain `_build_browser_env()`: no
`AGENT_BROWSER_SOCKET_DIR`, so it lived in agent-browser's default dir and no
reap path could see it. A wedged daemon + headless Chrome survived 47h across
two gateway restarts (#100855), and on macOS the genuine Chrome binary it held
made "Chrome won't open" for the user.
Give the attach lane the same contract every other lane already has:
`_prepare_session_socket_dir()` (per-session socket dir + `owner_pid` claim)
and `_agent_browser_command_env()`, and add the named dir to the orphan
reaper's scan. The existing `_reap_socket_dir` then applies its owner-liveness
and start-time-fingerprint rules unchanged; the daemon is listed as tracked so
the untracked-idle escape hatch never fires under a live user (per-task `rp_*`
sessions drive it over `--cdp`, so its own dir shows no activity). The daemon-side idle
timeout is NOT inherited: Chrome is launched by Hermes, not the daemon, so a
self-exiting daemon would leave Chrome holding the copy dir while the next
attach re-runs the snapshot overlay over it.
When a reaped daemon's Chrome (Hermes-launched, own process group) still
holds the copy dir, `_real_profile_cdp` re-attaches to it instead of running
the snapshot overlay over a live profile. DevToolsActivePort outlives a
crashed Chrome and its port can be recycled, so the file's browser id must
match `/json/version` before it is trusted; an attach failure on a live
Chrome fails closed rather than overlaying.
Tests: attach/get/close commands carry the reaper-visible socket dir and
owner_pid and no idle timeout; a dead-owner real-profile daemon is reaped by
`_reap_orphaned_browser_sessions`; a surviving Chrome is re-attached, never
overlaid (all red on main).
The salvaged fix ran the injector against a copy.copy(agent) whose
tools/valid_tool_names pointed at the staged pair, wrapped in a blanket
try/except. A shallow copy of a live AIAgent (locks, DB handle, in-flight
attribute writes from the late-binding thread) is a workaround for the gate
mutating in place, and the except turned any failure into a silently
published snapshot WITHOUT message_agent.
Extract the gate as tools.bot_mode_dm.message_agent_authorized(agent) (the
same predicate ensure_message_agent_tool already used), and have the
snapshot builder append message_agent_tool_schema() to the staged list
directly when it passes. No copy, no swallow; same tests, same live
behaviour (compaction / between-turns / resume keep the tool; ordinary
sessions scrubbed).
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
run_job's finally now asks the worker Future whether it is still running before deciding who tears the session down; the heartbeat test's stand-in future lacked done().
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
those files carry no windows_only marker so they never run on Windows, and the monkeypatched
predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
db.close() already releases a registry-shared handle, so the described NameError leak only
existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
teardown; replacement not published before the last close settles) plus the raising-close
settlement; comments say the WHY once.
The deferral helpers from the previous commit were appended to the cron/scheduler.py facade and
carried ~70 lines of fallbacks (getattr/callable checks, result() waits, a teardown_registered
flag) for hypothetical Future doubles; the only producer is _cron_pool.submit, always a real
concurrent.futures.Future whose done()/add_done_callback() cannot raise. Also drops the
_finalize_cron_session_db passthrough. Behaviour unchanged; the test now asserts the
finalize+teardown contract on the real Future instead of a patched wrapper.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
N concurrent real AIAgent sessions on a growing tool loop; per-call cache_read /
cache_creation / response id / upstream / prefix shas; consecutive pairs classified
ideal / stuck / collapse. argparse (--provider nous|openrouter|anthropic, --wire,
--pin, --settle, --ttl, --model); wire defaults to what Hermes would pick. Summary
JSON carries bad pairs with both response ids. Smoke-run against Nous from the repo
path (picked chat via nous_api_mode; 6/6 ideal).
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).
The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.
Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
- pin session-start workspace snapshot (_frozen_workspace_snapshot) on first build so dynamic git probes don't churn Tier 2 during compaction rebuilds in active coding sessions
- replay pinned snapshot across rebuilds when cwd matches; re-probe only on cwd switch
- honor coding_context invariant that workspace is a session-start snapshot, preventing prefix-cache divergence at offset ~4,662
- add invariant tests covering workspace snapshot pinning across git mutations and cwd transitions
- addresses upstream prompt divergence identified in #103326
Since #104299 a background delegate_task call is split into completion units
(one per `group`, one per ungrouped task). A multi-child unit still joined on
all its children before anything was written durably, so an owner crash
between the first and last child lost the finished work and replayed the whole
unit as "outcome unknown" — the restart-granularity gap that #104233 (Xipong's
#76228/#76229 direction) solved with a second row per child.
Each finished child of a detached unit is now recorded on the unit's OWN row
(`record_unit_child` → result_json {results, partial}) as its future lands;
the real result overwrites it at finalize. `recover_abandoned_delegations`
replays recorded children with their real summaries and marks only the
unfinished ones unknown, naming the count. No new rows, no new consumer shape.
`task_indexes` is persisted so recovery knows a split unit's members.
- session-control-goal.tsx: a 'send' dispatch against a busy session now
parks the kickoff on the composer queue (the backend already resumed the
goal) instead of reporting continuationFailed. Shared helper
queueKickoffIfSessionBusy() extracted from slash.ts so both paths agree.
- gateway-switch.ts: wipeSessionListsForGatewaySwitch clears
$sessionControlBySession (runtime-id keyed; new backend re-mints ids).
Dead resetSessionControlAfterGatewayRebind removed; clearSessionControl
now called from the session delete path beside clearQueuedPrompts.
- session-control.tsx: read/hydration failures use controlUnavailable copy,
not actionFailed.
- i18n: heartbeatDueWaitingForIdle added to ja/ru/zh-hant; new
continuationQueued/continuationBusy/controlUnavailable in all locales.
- /goal and /loop are command.dispatch built-ins (never reach the slash worker), so the
worker-branch publish never fired for the most common typed commands and the structured
card stayed stale until the next turn. The publish helper now lives beside _snapshot_control
in methods_session_control.py and is called from both slash paths plus the end of the turn
tail — after the goal judge / loop tick evaluation, which mutate state AFTER message.complete.
- session.control skips the duplicate emit for dispatcher-backed actions (still exactly one
update per mutation).
- Removed tui_gateway/desktop_heartbeat_driver.py, the entry/ws hooks and the hermes_cli lock
rework: #104224 drives /heartbeat from the existing per-session poller for every TUI client.
- test_session_control.py: dropped presentation-shaped cases, added the publication invariants
(both slash paths, unknown dispatch publishes nothing, real _run_prompt_submit turn publishes
post-judge). 3/3 new tests fail on the contributor backend.
Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.
`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.
GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.
Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.
Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.