_call_llm_impl applied auxiliary_max_tokens_param whenever
fast_compression_cap was non-None — but _compression_fast_lane_controls
passes an explicit caller max_tokens straight through, so a compression
call that set its own cap had the param force-injected onto providers
where _build_call_kwargs deliberately omits it (ZAI vision hard-400s on
max_tokens; GPT-5/Copilot require max_completion_tokens). Pre-fast-lane
main omitted the param for that exact call shape (verified via
subprocess pinned to the pre-PR base).
Gate the forced param on 'max_tokens is None' so it applies only to caps
the certified lane itself produced — the same guard the fallback path
already uses.
Regression test pins the pre-PR wire shape. Mutation-checked.
_run_protected_sync_provider_call propagates the forward-progress hook to
its daemon worker but not the new _aux_dispatch/_aux_provider_response
timing hooks (both threading.local). When compression takes the protected
path — the common case, since the summary call runs under
aux_interrupt_protection with a hard-cancel source — provider_dispatch_ms
and time_to_first_progress_ms were silently absent from telemetry.
Also collapse the two byte-identical save/restore context managers
(aux_progress_hook, _aux_timing_hook) onto one _aux_thread_local_hook
implementation so the propagation semantics can never drift between the
progress and timing slots.
Regression test drives _run_protected_sync_provider_call with both timing
hooks installed and asserts the worker-thread notifies reach them.
Mutation-checked (reverting the propagation fails the new test).
Post-review hardening on the attempt-ownership commit:
- Write-time re-validation: the entry staleness check in
_restore_compressor_attempt_state runs before the durable-cooldown DB
I/O, so a fallback could claim the compressor in that window and the
stale setattr loop would still clobber its state. The in-memory writes
now re-validate AND execute under _COMPRESSOR_ATTEMPT_LOCK — the same
lock claims are taken under. The DB rollback stays outside the lock
(safe: the dangerous direction requires a prior claim, which the entry
check rejects). Both the quality reviewer and the lead's independent
pre-verification converged on this window.
New deterministic test: TestMidRestoreClaimRace (claim injected
between entry check and write via instrumented deepcopy).
- Documented gen-0 semantics on _claim_compressor_attempt: per-compressor
all-or-nothing, never mixed with gen>0 on one instance (reviewer
finding 2, verified unreachable — comment hardens against future
confusion).
Dropped after verification (reviewer finding 3): resetting
_SUMMARY_ROUTE_CONSUMED on pin_summary_route exit — the echo lives in
the worker thread's COPIED context (propagate_context_to_thread) and
dies with it; a probe confirmed the next attempt's context is clean.
Resetting it would break digests running after the with-block.
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:
1. Late-primary snapshot restore: the detached primary's unwind called
_restore_compressor_attempt_state with the PRIMARY's pre-attempt
snapshot. Landing after the fallback's commit it rolled
_previous_summary/cooldown/provenance/telemetry back to pre-primary
values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
cleared the callback the fallback had just installed, so the
fallback's F4 cancellation consult read None.
Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.
Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
summary onto the pinned healthy route: take_pinned_summary_route()
echoes the consumed route into a context-local
_SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
single-use contract for the SUMMARY call is unchanged — the
main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
accepted limitation on _retry_compression_on_fallback_chain
(built-ins idempotent; resuming mid-pipeline would couple the retry
to host callback ordering).
Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).
The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
resolve_compression_fast_lane and _compression_config_claims_fast_lane
each hand-parsed the same four config fields (provider, model,
reasoning_effort, max_output_tokens) with copy-pasted normalization and
int-coercion. Extract _fast_lane_config_fields() as the single source of
truth for both.
This also fixes a real inconsistency the duplication hid: certification
checked the literal string 'none' while _get_task_extra_body routes
reasoning_effort through parse_reasoning_effort, which treats 'false',
'disabled', and YAML boolean false as disabled too. A user writing
reasoning_effort: false got reasoning disabled but silently lost the
fast-lane cap. Certification now delegates to parse_reasoning_effort so
the two predicates can never disagree.
Regression test: every disabled-spelling certifies; empty/real efforts
do not. Mutation-checked (reverting to the literal check fails the new
test).
_compression_fast_lane_controls unconditionally did body = dict(extra_body)
before the early-return guard. Move the copy below the guard so non-
compression auxiliary calls (vision, title_generation, etc.) return the
original extra_body reference without a shallow copy.
The contributor's PR description notes this commit was deliberately
omitted as obsolete. Removing the planning doc that was accidentally
included via cherry-pick.
Three-reviewer pass (reuse/quality/efficiency) on the guard file:
- Drop dead `agent._interrupt_requested = False` setup: the
`_record_streamed_assistant_text` chain only consults
`_stream_writer_superseded()` (stream-writer TLS token), never
`_interrupt_requested` — verified by reading both call sites.
- Deduplicate the two trace-callback blocks into one
`_count_writer_statements` helper using the house idiom
(`statements.append` — 9 existing uses in tests/test_hermes_state.py)
instead of a mutable-dict counter closure; failure messages now dump
the captured SQL for direct diagnosis.
Dropped after verification: reviewer suggestion to flip xfail to
strict=True — its premise ("the fix PRs already remove the markers")
is wrong: #92166/#95380 predate this file and cannot remove markers
they don't contain, so strict=True would redden main's CI the moment
either merges. strict=False + follow-up marker removal is the
deliberate no-red-window ratchet.
Efficiency reviewer: no material findings (GC delta 0.06 on the ratio,
36.9MB peak, sqlite trace API stable since 3.14, dir convention OK).
Re-verified post-fold: 3x main runs (2 passed, 2 xfailed), flip checks
on both fix branches still pass with --runxfail.
Pattern B (O(N²) rebuild-per-delta in hot paths) has no lintable
signature, unlike Pattern A's ASYNC ruff gate: `s += frag` is quadratic
in a hot loop and harmless elsewhere. The only durable prevention is
behavioral — pin the scaling SHAPE of each known hot path and fail CI
when it regresses.
New tests/perf_guards/test_pattern_b_scaling.py, three guards:
1. Streamed-text accumulation (fix in flight: #92166)
Self-normalizing ratio: time at 16k deltas over 4k deltas.
Linear ≈ 4x, quadratic ≈ 16x; main measures 9.6x → strict bound 7.0.
xfail on today's main, PASSES on the #92166 branch (verified).
2. list_sessions_rich statement count (fix in flight: #95380)
Deterministic — counts writer-connection SQL statements via sqlite
trace callback, zero timing. Bounded-constant guard xfails on main
(measured N+1: 26 stmts / 12 sessions), PASSES on #95380 (verified).
A second, weaker budget guard (≤4·N+8) passes TODAY and catches a
regression from N+1 to N·M immediately.
3. Tool-call fragment assembly (#92242 shape)
Ratio guard on the buffered-parts accumulator model; sized so the
small case takes ≥3ms (sub-ms bases jitter on CI runners).
Flake hardening: min-of-K timing, ratio thresholds with ≥2x separation
from both measured-good and measured-bad, operation counts preferred
over timing. 10/10 identical outcomes across repeated local runs.
The xfail markers are the ratchet contract: each names its fix PR and
must be removed when that PR merges, flipping the guard to enforcing.
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
/simplify-code reuse finding: _write_marker reimplemented the
mkstemp->write->os.replace pattern that utils.atomic_write_text already
provides as the repo's shared atomic-text-write helper (and the shared
version adds fsync + cross-device/busy-file fallbacks).
Follow-up to the #95605 salvage, closing the review findings:
- _copy_alias no longer swallows OSError silently: it warns (a leftover
alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
(update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
it now asserts the whole layout (anchor + aliases) is complete, so a
partially-materialized alias set can never read 'active' in doctor —
the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.
5 new regression tests.
Review fixes from kokhlo's live-hardware review:
- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
PYTHONHOME=<venv> boots a staged copy that would otherwise die with
"No module named 'encodings'", papering over the exact prefix
failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
images) still skip; EACCES after our own chmod now refuses the
install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
so the managed-runtime layout (cpython-3.11-macos-* symlinked to
cpython-3.11.15-macos-*) no longer reports stale on a fresh install.
Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
The gate is the never-brick guarantee, so each direction gets its own
test: nonzero exit (dyld/encodings crash), build-time prefix leak,
timeout, and the deliberate OSError skip (a binary that cannot execute
here means the symlinked venv was equally dead — installing cannot make
things worse).
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).
Re-land:
- Keep the signed real-file copy of bin/python (identifier-pinned
via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
symlinks. Copies boot on every build we could reproduce and keep
the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
plus the venv prefix, abort and leave the live venv untouched
on failure.
Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.
Closes#95596.
In Bot Mode every roster click resolved the bot's canonical "Bot Chat" by
name and opened it as a tab. Nothing records a tab close (the plugin keeps no
closed set; core's tile bucket only forgets), so a Bot Chat the user had
closed came back beside every newer thread on every bot switch — close it,
start a new thread, visit another bot, come back: two tabs again, forever.
A row click is now "go to this bot": when the bot's workspace already holds
tabs, the one the user last had active is fronted and no chat is resolved or
opened. The canonical chat is opened only when the bot has nothing open, or
on the explicit asks — a new "Open Bot Chat" row-menu item and the Bots home
"Open chat" button (`openRosterBot(bot, { canonical: true })`).
- session-states: `focusWorkspaceOwnerSessionTile(ownerKey)` fronts the
owner's remembered-active tile (else its most recent) and reports it.
- sdk: `host.focusOpenWorkspaceSession(ownerKey)` exposes it to plugins;
feature-detected in the plugin so older shells keep the canonical open.
- hermes-bots: `focusExistingBotTab` short-circuits `openRosterBot`; the
claim it records carries only the fronted tab, and the session.reclaimed
re-resume now skips such claims so it cannot resurrect the closed chat.
Tests: vitest for the core helper, a node test for the click path (open tabs
win, nothing open → canonical, explicit canonical, older shell, throwing
host), and an e2e that seeds two bots with real "Bot Chat" rows, closes one,
starts a thread, switches bots and back, and asserts the Bot Chat stays
closed until asked for explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Tours had no switch at all, and the tips switch only covered the app's
own rotation — so "I don't want these" was answerable for half of one
feature and none of the other. Both are now a row in Settings →
Appearance, on by default, and off means off for Hermes as well: an
agent tip is dropped at the bridge and a tour request is refused in
words, so the agent hears that the walkthrough didn't run instead of
narrating a spotlight nobody can see.
Renames $tipRotationEnabled to $tipsEnabled to match what it now
governs. The storage key keeps the old name on purpose — renaming it
would read as unset for anyone who had already turned tips off, and
silently turning them back on is the one outcome worth avoiding.
The switches stop at the app's edge for now: the tools are still in the
model's schema and the calls simply don't land. Withdrawing them
entirely needs the setting to reach the backend, which is the next PR.
Made opt-in when it fired a tip 45 seconds into a launch and another
every six minutes, which is a cadence that owes you a choice. The pacing
has since become a settling delay of five to ten minutes per launch and
a six-hour cooldown persisted across them — roughly a tip a day, weeks
to walk the catalog. At that weight the switch has nothing left to
protect anyone from, and a discovery feature nobody meets is one nobody
has. The switch stays for whoever still wants it off.
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
The overlay table is only actions, pickers, and RPCs. Completions group
by backend kind, and argument mode comes from the warmed catalog so
/review and plugin commands stay typeable.
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
A first tip 45 seconds in and one every six minutes after walks the
whole catalog in an hour, which is the cadence of a notification rather
than a nicety. Games get this right by being almost absent: a tip while
you settle in, then nothing for the rest of the day.
Two clocks now have to agree. A per-launch settling delay of five to ten
minutes means opening the app is never met with a bubble, and a six-hour
cooldown persisted across launches means quitting and reopening isn't a
way to farm them — the old schedule lived in the effect and re-armed on
every mount. Flipping the switch on skips the settling delay and offers
immediately, since that clock guards a launch you came into with a
purpose, not a deliberate opt-in.
An agent tip starts the cooldown too: whoever just pointed at something,
the user has had their one interruption for a while.
The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.
Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
An ambient rotation that points at parts of the app you may not have found yet
— one accent bubble, an arrow, and an outline around the subject. No scrim and
no spotlight: a tip is a pointer beside your work, not a modal in front of it,
so it takes no focus, owns no Esc, and blocks nothing.
It only speaks when the app is genuinely quiet — nothing streaming, no dialog,
menu or tour up, window focused, a few seconds since the last keystroke — and
walks the catalog in order rather than shuffling, so tips arrive as a tour of
neighbouring parts of the app instead of unrelated ones. A tip with nothing on
screen to point at is skipped, not waited for.
Closing one with its ✕ retires it for good. That is the whole reason the ✕ is a
heavier gesture than letting the bubble time out, and Settings → Appearance is
the only way back — alongside the switch that turns the feature off entirely.
Tours and tips both address elements by selector, and everything they most want
to point at — the composer, the model pill, the nav rows, the profile rail, the
right-pane toggle — was reachable only by icon, position, or a translated
aria-label. None of those survive a re-render, a theme, or a locale change.
Two handles are needed per surface, not one, because where an arrow points and
what an outline wraps are different questions: a nav row's label carries the
`data-tour` handle so an arrow lands at the end of the word, and defers the
outline to the row via `data-tip-arrow-only`.
The collector also has to skip panes hidden by the keep-alive stack. An inactive
tab stays mounted under `visibility: hidden` to keep its scroll position, so its
rect is identical to the live tab's and no geometry test separates them — which
is how a tour could spotlight a background tab's composer.
The default popover is glass over the app's own chrome, which is right for
something the user opened and wrong for something the app said. The accent
variant fills the same box — same arrow, same placement engine — with a solid
brand colour so an unprompted surface reads as the app speaking.
Filling it with `primary` directly doesn't work across themes: a pale accent is
a perfectly valid primary (imported VS Code themes love a pastel), and the
honest `primaryForeground` for one is near-black, so the loud surface comes out
a pastel card whispering. `--dt-primary-solid` deepens the hue until a light
foreground clears AA — a no-op on an accent that is already deep, and darkening
only, so the hue survives.
forceResume already requested a main-route resume, but Bot Chat is a
tile. Reopening reused the warm cached transcript and skipped REST, so
cron bot-chat deliveries that landed while the panel was closed stayed
invisible until app restart. Refresh the tile transcript on explicit
open and merge the persisted tail into the cache.
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.