Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:
1. Late-primary snapshot restore: the detached primary's unwind called
_restore_compressor_attempt_state with the PRIMARY's pre-attempt
snapshot. Landing after the fallback's commit it rolled
_previous_summary/cooldown/provenance/telemetry back to pre-primary
values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
cleared the callback the fallback had just installed, so the
fallback's F4 cancellation consult read None.
Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.
Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
summary onto the pinned healthy route: take_pinned_summary_route()
echoes the consumed route into a context-local
_SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
single-use contract for the SUMMARY call is unchanged — the
main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
accepted limitation on _retry_compression_on_fallback_chain
(built-ins idempotent; resuming mid-pipeline would couple the retry
to host callback ordering).
Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).
The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
resolve_compression_fast_lane and _compression_config_claims_fast_lane
each hand-parsed the same four config fields (provider, model,
reasoning_effort, max_output_tokens) with copy-pasted normalization and
int-coercion. Extract _fast_lane_config_fields() as the single source of
truth for both.
This also fixes a real inconsistency the duplication hid: certification
checked the literal string 'none' while _get_task_extra_body routes
reasoning_effort through parse_reasoning_effort, which treats 'false',
'disabled', and YAML boolean false as disabled too. A user writing
reasoning_effort: false got reasoning disabled but silently lost the
fast-lane cap. Certification now delegates to parse_reasoning_effort so
the two predicates can never disagree.
Regression test: every disabled-spelling certifies; empty/real efforts
do not. Mutation-checked (reverting to the literal check fails the new
test).
_compression_fast_lane_controls unconditionally did body = dict(extra_body)
before the early-return guard. Move the copy below the guard so non-
compression auxiliary calls (vision, title_generation, etc.) return the
original extra_body reference without a shallow copy.
The contributor's PR description notes this commit was deliberately
omitted as obsolete. Removing the planning doc that was accidentally
included via cherry-pick.
PR review (andrexibiza, post-a69a9c351d) caught the one operator-facing
surface the identity change missed: the example config still promised
the profile-scoped ID is NOT sent and described the 30-day rotating
HMAC. Rewritten to state the stable install_id is transmitted as-is,
matching the sender, wizard, and docs A.2. Swept the repo for further
stale references: none remain (the HMAC text in relay-shared-metrics.md
A.2/A.3 is the intentional decision record).
Three-reviewer pass (reuse/quality/efficiency) on the guard file:
- Drop dead `agent._interrupt_requested = False` setup: the
`_record_streamed_assistant_text` chain only consults
`_stream_writer_superseded()` (stream-writer TLS token), never
`_interrupt_requested` — verified by reading both call sites.
- Deduplicate the two trace-callback blocks into one
`_count_writer_statements` helper using the house idiom
(`statements.append` — 9 existing uses in tests/test_hermes_state.py)
instead of a mutable-dict counter closure; failure messages now dump
the captured SQL for direct diagnosis.
Dropped after verification: reviewer suggestion to flip xfail to
strict=True — its premise ("the fix PRs already remove the markers")
is wrong: #92166/#95380 predate this file and cannot remove markers
they don't contain, so strict=True would redden main's CI the moment
either merges. strict=False + follow-up marker removal is the
deliberate no-red-window ratchet.
Efficiency reviewer: no material findings (GC delta 0.06 on the ratio,
36.9MB peak, sqlite trace API stable since 3.14, dir convention OK).
Re-verified post-fold: 3x main runs (2 passed, 2 xfailed), flip checks
on both fix branches still pass with --runxfail.
Pattern B (O(N²) rebuild-per-delta in hot paths) has no lintable
signature, unlike Pattern A's ASYNC ruff gate: `s += frag` is quadratic
in a hot loop and harmless elsewhere. The only durable prevention is
behavioral — pin the scaling SHAPE of each known hot path and fail CI
when it regresses.
New tests/perf_guards/test_pattern_b_scaling.py, three guards:
1. Streamed-text accumulation (fix in flight: #92166)
Self-normalizing ratio: time at 16k deltas over 4k deltas.
Linear ≈ 4x, quadratic ≈ 16x; main measures 9.6x → strict bound 7.0.
xfail on today's main, PASSES on the #92166 branch (verified).
2. list_sessions_rich statement count (fix in flight: #95380)
Deterministic — counts writer-connection SQL statements via sqlite
trace callback, zero timing. Bounded-constant guard xfails on main
(measured N+1: 26 stmts / 12 sessions), PASSES on #95380 (verified).
A second, weaker budget guard (≤4·N+8) passes TODAY and catches a
regression from N+1 to N·M immediately.
3. Tool-call fragment assembly (#92242 shape)
Ratio guard on the buffered-parts accumulator model; sized so the
small case takes ≥3ms (sub-ms bases jitter on CI runners).
Flake hardening: min-of-K timing, ratio thresholds with ≥2x separation
from both measured-good and measured-bad, operation counts preferred
over timing. 10/10 identical outcomes across repeated local runs.
The xfail markers are the ratchet contract: each names its fix PR and
must be removed when that PR merges, flipping the guard to enforcing.
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.
Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
role (unreadable/non-object/id-less payloads still reject rather than
block the queue) and now records the raw install_id in
sent_install_id; _body rewrites the payload's install_id from that
frozen column, keeping byte-identical resends anchored to one
recorded value.
Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.
Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.
Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.
258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
The English bundle on the parent left the other locales to fall through.
Ship them the same way kanban does, using core's nouns, so a language
switch no longer paints the Bots rail in English.
Deleting the Bots home deleted the dead end that "Select a Bot or group
first." was apologizing for: with nothing selected the center is now an
ordinary session and + works there. The refusal still stands for the cases
that remain — a group room, and a selected row too orphaned to route.
Ownership of an empty transcript is per session and is not known until each
plugin has loaded its own data — which is why the slot mounts contributors
rather than resolving a claim up front. Taking only contributions[0] then let
a plugin that DECLINES a session suppress the one that owns it: silently,
permanently, and only for the users whose installed plugins happen to
register in that order.
Every registration is mounted and answers for itself. Declining is free; two
claiming one session render both, a visible conflict instead of a silent drop.
The composer subscribed to $sessionStates while isBotChatSession reads the
scope set — and, to resolve a live id to the stored one it is filed under,
$sessionTiles too. It stayed roughly right only because $sessionStates
republishes per streaming token, so the stale answer was overwritten within a
frame. A quiet session had no such luck.
Adds useStoresSelector beside useStoreSelector for the general case: a scalar
whose inputs span several stores, recomputed when any notifies, still
re-rendering only when the scalar changes.
The migration, its commit marker, and the rollback around it are all still
live in data.ts, but the suite covering them went out with the vm/regex
harness and was never ported — leaving only the happy-path commit ordering
asserted incidentally by bot-delete.
Nine cases: the sole-local topology gate (and the two refusals — multi-source,
and a batch where any one profile has no local route), the marker rule that
makes v2 authoritative (committed, markerless, crash-window), and the three
failed-commit paths (roll back to the last committed generation, clear both
keys when there is none, survive a failed cleanup).
migratedLocalRoutes is module-private, so where the old suite asserted the
map's size this one asserts what the map is for: no route adopted, nothing
written. Mutation-checked — dropping the marker rule fails three of them.
The adopt branch runs a full extra round-trip past the point every sibling
open in createCanonicalChat is probed at — registry miss, mint, title
conflict, re-consult — so it is the likeliest of them to land after the user
has clicked another bot, and it was the only one that navigated unguarded.
Adopting the winner still settles identity, which is always correct to
return; only the workspace steal is gated.
cronjob(action='run', prompt=...) context was silently dropped when the
manual run forwarded to the gateway (#96010 follow-up): POST
/api/jobs/{id}/run took no body. The forward now sends {prompt} in the
request body; the api_server validates it (length cap + strict injection
scan, same as stored prompts) and trigger_job stamps it as a transient
manual_run_prompt alongside manual_run_at. run_one_job consumes the stamp
for that single fire and mark_job_run clears it, so it never persists
into the job definition or later scheduled fires.
Ports 7c91079 onto the split modules — it landed on main in plugin.js, which
this branch deletes.
Every roster click resolved the bot's canonical Bot Chat by name and opened
it as a tab. Nothing records a tab close (this plugin keeps no closed set;
core's tile bucket only forgets), so a Bot Chat the user had closed came back
beside every newer thread on every bot switch. A click is now "go to this
bot": when the bot's workspace already holds tabs, the one the user last had
active is fronted and no chat is resolved or opened. The forever-chat opens
only when the bot has nothing open, or on the explicit ask — a new "Open Bot
Chat" row-menu item.
The claim such a click records carries only the fronted tab, so the reclaim
listener skips it and cannot resurrect the closed chat. focusExistingBotTab
is feature-detected, so an older shell keeps opening the canonical chat.
Dropped from the port: the Bots home "Open chat" button, a surface this
branch removes. The .mjs test is replaced by a real one — its two
source-reading cases become a render of the menu item in bot-row.test.tsx
and, for the reclaim guard, the claim-shape invariant the guard reads.
Stubbing usePluginI18n there also means the row's localized labels render as
text in tests instead of empty.
Keeps plugin.js and its new .mjs test deleted. main's closed-chat fix
(7c91079) landed in both; it is a real behaviour change, so the commit that
follows ports it onto the split modules rather than dropping it with the
files. Its core half — focusWorkspaceOwnerSessionTile and the
host.focusOpenWorkspaceSession verb — merged cleanly and is used as-is.
/simplify-code reuse finding: _write_marker reimplemented the
mkstemp->write->os.replace pattern that utils.atomic_write_text already
provides as the repo's shared atomic-text-write helper (and the shared
version adds fsync + cross-device/busy-file fallbacks).
Follow-up to the #95605 salvage, closing the review findings:
- _copy_alias no longer swallows OSError silently: it warns (a leftover
alias symlink is the exact #95541 crash shape) and reports failure.
- Alias staging uses mkstemp (unique names) so concurrent ensures
(update + doctor --fix) can never promote a truncated interim copy.
- The anchor marker is written LAST and atomically (write-then-rename):
it now asserts the whole layout (anchor + aliases) is complete, so a
partially-materialized alias set can never read 'active' in doctor —
the next ensure retries the install instead.
- /.hermes-runtime/python/ store marker is derived from
managed_uv._RUNTIME_DIR_NAME instead of a hardcoded string.
5 new regression tests.
Review fixes from kokhlo's live-hardware review:
- The boot-gate probe now runs with PYTHONHOME / PYTHONPATH /
PYTHONSTARTUP / __PYVENV_LAUNCHER__ scrubbed: an inherited
PYTHONHOME=<venv> boots a staged copy that would otherwise die with
"No module named 'encodings'", papering over the exact prefix
failure the gate exists to catch.
- OSError is split by errno: ENOENT/ENOEXEC (fixtures, foreign-arch
images) still skip; EACCES after our own chmod now refuses the
install instead of silently accepting a broken copy.
- Marker writes and both marker comparisons go through os.path.realpath,
so the managed-runtime layout (cpython-3.11-macos-* symlinked to
cpython-3.11.15-macos-*) no longer reports stale on a fresh install.
Tests: +3 (env-scrub spy, EACCES refusal, symlinked-home state).
100 passed in the module + doctor neighborhoods.
The gate is the never-brick guarantee, so each direction gets its own
test: nonzero exit (dyld/encodings crash), build-time prefix leak,
timeout, and the deliberate OSError skip (a binary that cannot execute
here means the symlinked venv was equally dead — installing cannot make
things worse).
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).
Re-land:
- Keep the signed real-file copy of bin/python (identifier-pinned
via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
symlinks. Copies boot on every build we could reproduce and keep
the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
plus the venv prefix, abort and leave the live venv untouched
on failure.
Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.
Closes#95596.
In Bot Mode every roster click resolved the bot's canonical "Bot Chat" by
name and opened it as a tab. Nothing records a tab close (the plugin keeps no
closed set; core's tile bucket only forgets), so a Bot Chat the user had
closed came back beside every newer thread on every bot switch — close it,
start a new thread, visit another bot, come back: two tabs again, forever.
A row click is now "go to this bot": when the bot's workspace already holds
tabs, the one the user last had active is fronted and no chat is resolved or
opened. The canonical chat is opened only when the bot has nothing open, or
on the explicit asks — a new "Open Bot Chat" row-menu item and the Bots home
"Open chat" button (`openRosterBot(bot, { canonical: true })`).
- session-states: `focusWorkspaceOwnerSessionTile(ownerKey)` fronts the
owner's remembered-active tile (else its most recent) and reports it.
- sdk: `host.focusOpenWorkspaceSession(ownerKey)` exposes it to plugins;
feature-detected in the plugin so older shells keep the canonical open.
- hermes-bots: `focusExistingBotTab` short-circuits `openRosterBot`; the
claim it records carries only the fronted tab, and the session.reclaimed
re-resume now skips such claims so it cannot resurrect the closed chat.
Tests: vitest for the core helper, a node test for the click path (open tabs
win, nothing open → canonical, explicit canonical, older shell, throwing
host), and an e2e that seeds two bots with real "Bot Chat" rows, closes one,
starts a thread, switches bots and back, and asserts the Bot Chat stays
closed until asked for explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Tours had no switch at all, and the tips switch only covered the app's
own rotation — so "I don't want these" was answerable for half of one
feature and none of the other. Both are now a row in Settings →
Appearance, on by default, and off means off for Hermes as well: an
agent tip is dropped at the bridge and a tour request is refused in
words, so the agent hears that the walkthrough didn't run instead of
narrating a spotlight nobody can see.
Renames $tipRotationEnabled to $tipsEnabled to match what it now
governs. The storage key keeps the old name on purpose — renaming it
would read as unset for anyone who had already turned tips off, and
silently turning them back on is the one outcome worth avoiding.
The switches stop at the app's edge for now: the tools are still in the
model's schema and the calls simply don't land. Withdrawing them
entirely needs the setting to reach the backend, which is the next PR.
Made opt-in when it fired a tip 45 seconds into a launch and another
every six minutes, which is a cadence that owes you a choice. The pacing
has since become a settling delay of five to ten minutes per launch and
a six-hour cooldown persisted across them — roughly a tip a day, weeks
to walk the catalog. At that weight the switch has nothing left to
protect anyone from, and a discovery feature nobody meets is one nobody
has. The switch stays for whoever still wants it off.
Inference is now any args_hint without subcommands → text. Mixed is the
only remaining hint-token path. desktop= and the few argument_mode
overrides live on the registry entry; the side tables are gone. Catalog
aliases get their own dict copy. Composer tests seed the catalog so
/goal stays mixed without an overlay row.
The overlay table is only actions, pickers, and RPCs. Completions group
by backend kind, and argument mode comes from the warmed catalog so
/review and plugin commands stay typeable.
New commands and plugins declare argument_mode and desktop availability
on CommandDef / register_command. commands.catalog ships that map so
desktop does not need a second command list.
main's Bots-home flash fix (6f8be61) landed in plugin.js, which this branch
deletes. Both files stay deleted: the guard it adds (botOpenInFlight gating
botsHomeMayOpen) protects a surface this branch removes, and the two sites
where it renamed the generation bump to cancelBotOpen already bump here via
bumpBotOpenGeneration in shared.ts.
The forward dialed a hardcoded 127.0.0.1. The api_server adapter binds
extra.host -> API_SERVER_HOST -> 127.0.0.1, so mirror that chain when
dialing. Wildcard binds (0.0.0.0/::) listen on loopback, so keep dialing
loopback for those; bracket bare IPv6 literals.
The CLI has no 'trigger' subcommand ('trigger' is only an alias of the
cronjob TOOL's run action). Point operators at the real remediation:
start the gateway; its ticker owns relay-fronted delivery and fires the
job on schedule.
A manual 'hermes cron run' on a relay-fronted target has no live relay adapter
and no standalone sender, so it now forwards to the running gateway's
POST /api/jobs/{id}/run (marks due for the gateway ticker, which delivers via
the live relay adapter). Gateway unreachable -> the accurate 'start the gateway
or use cron trigger' error. Native topologies are untouched.
A manual in-process 'hermes cron run' has no live relay adapter, but the
delivery loop fell through to the native standalone path and hit the native
configured/enabled gate, misdiagnosing relay-fronted platforms ('not
configured/enabled') whose credential lives in the connector. Now, when
resolve_delivery_transport finds no transport AND the platform is in
relay_fronted_platforms(), emit the accurate 'start the gateway or use cron
trigger' remediation and skip the native gate. Native topologies unchanged.
A first tip 45 seconds in and one every six minutes after walks the
whole catalog in an hour, which is the cadence of a notification rather
than a nicety. Games get this right by being almost absent: a tip while
you settle in, then nothing for the rest of the day.
Two clocks now have to agree. A per-launch settling delay of five to ten
minutes means opening the app is never met with a bubble, and a six-hour
cooldown persisted across launches means quitting and reopening isn't a
way to farm them — the old schedule lived in the effect and re-armed on
every mount. Flipping the switch on skips the settling delay and offers
immediately, since that clock guards a launch you came into with a
purpose, not a deliberate opt-in.
An agent tip starts the cooldown too: whoever just pointed at something,
the user has had their one interruption for a while.
The two halves of tips were behind one switch, which meant the app
volunteering commentary at idle shipped on by default. Split them along
the line that matters: the rotation talks unprompted, so it now waits to
be asked for, while an agent tip stays ungated like the tour it mirrors
— Hermes raises one mid-conversation, in answer to something the user
said.
Drops the tool's config gate along with the config key it read. The
renderer mirrored that key with config.set, which has no branch for it
and answered "unknown config key" into a swallowed catch, so the opt-out
never reached the backend in the first place.
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
An ambient rotation that points at parts of the app you may not have found yet
— one accent bubble, an arrow, and an outline around the subject. No scrim and
no spotlight: a tip is a pointer beside your work, not a modal in front of it,
so it takes no focus, owns no Esc, and blocks nothing.
It only speaks when the app is genuinely quiet — nothing streaming, no dialog,
menu or tour up, window focused, a few seconds since the last keystroke — and
walks the catalog in order rather than shuffling, so tips arrive as a tour of
neighbouring parts of the app instead of unrelated ones. A tip with nothing on
screen to point at is skipped, not waited for.
Closing one with its ✕ retires it for good. That is the whole reason the ✕ is a
heavier gesture than letting the bubble time out, and Settings → Appearance is
the only way back — alongside the switch that turns the feature off entirely.