Tours and tips both address elements by selector, and everything they most want
to point at — the composer, the model pill, the nav rows, the profile rail, the
right-pane toggle — was reachable only by icon, position, or a translated
aria-label. None of those survive a re-render, a theme, or a locale change.
Two handles are needed per surface, not one, because where an arrow points and
what an outline wraps are different questions: a nav row's label carries the
`data-tour` handle so an arrow lands at the end of the word, and defers the
outline to the row via `data-tip-arrow-only`.
The collector also has to skip panes hidden by the keep-alive stack. An inactive
tab stays mounted under `visibility: hidden` to keep its scroll position, so its
rect is identical to the live tab's and no geometry test separates them — which
is how a tour could spotlight a background tab's composer.
The default popover is glass over the app's own chrome, which is right for
something the user opened and wrong for something the app said. The accent
variant fills the same box — same arrow, same placement engine — with a solid
brand colour so an unprompted surface reads as the app speaking.
Filling it with `primary` directly doesn't work across themes: a pale accent is
a perfectly valid primary (imported VS Code themes love a pastel), and the
honest `primaryForeground` for one is near-black, so the loud surface comes out
a pastel card whispering. `--dt-primary-solid` deepens the hue until a light
foreground clears AA — a no-op on an accent that is already deep, and darkening
only, so the hue survives.
Swapping no-redeclare for the TS-aware rule is the right fix, but touching
an eslint config makes the autofix bot's patches wait on a team review and
stalls its auto-merge. The two overload pairs that needed it carry an inline
disable instead; the rule swap can land on its own.
groupChatSyncMemberKey keyed on a field the descriptors never carry, so
every member hashed to the empty string and the round engine saw one member
where there were several; it keys on botRosterKey now, which is durable
across machines. botMetaWriteAt grew an entry per write and never dropped
one — it prunes past 60s, which is longer than the echo it exists to
suppress. aliasRouteIndex replaced the whole map on rebuild, so a slower
rebuild finishing second clobbered a newer one; a generation token means
only the newest result lands. Regression tests for each.
Dead since the split: EYE_X/EYE_Y, generatedSessionTitle, enabledMcp,
groupChatSyncDeletedRevision — each down to a declaration with no reader.
The bot source-status labels were half-localized, so they moved onto one
helper here rather than in the i18n pass. no-redeclare is disabled inline
over the two overload pairs it misreads, rather than in the shared config.
A raw checkbox in the routine editor where the SDK already exports one. The
same resizable-panel style block inline in three dialogs, now a
ResizableFrame beside the other dialog parts. Four inline styles that were
Tailwind spelled longhand — model-picker's was literally flex flex-col
gap-2. The twenty that remain are computed grid columns and avatar sizes,
which have to stay inline.
Both roster filter menus, the avatar picker's tabs. Renamed group.newDesc to
group.manageDesc since it describes managing an existing group, not making
one.
The getPluginCtx()?.i18n.t() sites only guarded the context, not i18n on it
— a plugin context without the bundle threw mid-render and painted an empty
rail. Guarded both.
petFrameCache is keyed by spritesheet URL over a 4500-pet gallery and holds
decoded PNG data URLs, so scrolling pinned every pet you passed until the
window closed. Capped at 120 — five pages, so scrolling back stays instant
and a miss only re-pays the fetch and crop. relayAgentsCache already swept
stale ids, but the sweep sits behind an early return that a shrink to one
connection skips; capped at 32, safe because every live connection is
rewritten each cycle so eviction can only reach ids that stopped being
fetched.
relay.ts also carried eight loose module-level lets, the state outlier
across the plugin's modules. Nothing renders from them, so they stay module
scope — collapsed into one record rather than moved to a store.
Six verbatim copies of the leading-glyph box down to two owners: transcript
lines get SCAFFOLD_GLYPH_CLASS from scaffold-row, sidebar rows get
SIDEBAR_ROW_LEAD. The bot rail's GatewayKindGlyph was a fork of core's
ConnectionGlyph down to the icon set, so it wraps the real one now and the
duplicate kind-to-icon tables are gone.
forceResume already requested a main-route resume, but Bot Chat is a
tile. Reopening reused the warm cached transcript and skipped REST, so
cron bot-chat deliveries that landed while the panel was closed stayed
invisible until app restart. Refresh the tile transcript on explicit
open and merge the persisted tail into the cache.
Row geometry lived inside chrome.tsx as private consts, so anything that
wanted to line up with a session row copied the literals instead — the lead
cell alone appeared verbatim in six files. Its own comment warns that owning
height anywhere else makes rows float 1-2px off sessions, which is exactly
what a copy invites. Split the measurements into row-geometry.ts, keep the
public names re-exported from chrome.tsx, and put the lead cell and
ConnectionGlyph on the plugin SDK so a plugin composes the same box rather
than re-deriving it. ConnectionGlyph takes a className now so a caller can
tint it without forking the component.
LruCache is the same move for a different leak: katex-memo.ts had a private
one with a hardcoded ceiling, and the bot rail had two Maps that never
evicted. One shared class, exported, katex's twin deleted.
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.
The plugin shipped as bundled JavaScript, so eslint never saw it. Now that it
is .tsx under src/, the whole ruleset applies: sorted JSX props, curly braces,
statement padding, unused imports.
Three of the ref writes the atom-mirror rule flagged are the cases its own
comment carves out — a scroll-position tracker fed by a DOM listener, a
previous-value tracker for the hidden -> visible edge (lagging a render IS its
contract), and a timer handle cleared on unmount. Those get the documented
disable. The fourth was a genuine mirror: McpSetupButton copied its profile
prop into a ref every render so two callers could read it. They read the prop
directly now, and the ref holds only the profile the component creates on
demand for the New Bot flow.
no-redeclare counts a TypeScript overload signature as a redeclaration of its
implementation, which is what botSelectionKey and botMetaKey tripped. Swapped
for the TS-aware version alongside the no-undef swap already there; hermes-ink,
whose vendored yoga bindings merge a const and a type under one name, extends
its existing carve-out to the new rule name.
A bot's canonical chat is the relationship — /new inside one would fork it into
a scratch session, so the composer reroutes /new to /compact there. The guard
deciding "is this chat the canonical one" read host.activeSessionId, which is
not on the host: the real atoms are host.state.activeSessionId (runtime id) and
host.state.focusedStoredSessionId (stored id). The optional chain swallowed it,
the comparison ran against null every turn, and /new reset forever-chats for as
long as the guard shipped.
It reads the focused STORED id now, which is the id space canonical_session
reports in. The comparison itself moves into isCanonicalChatOnScreen so a test
can drive it — matching either the durable registry row or the
compression-lineage tip, since a compacted Bot Chat is on screen under its tip
id while the registry still names it by the root.
The old harness sliced source text out of plugin.js, re-evaluated the
fragments through vm.runInNewContext, and in a good number of files simply
regex-matched the source for a symbol name. Nothing it asserted survived the
split into modules, and its blind spot was load-bearing: the /new guard
regex-matched clean for as long as it shipped dead.
These are the same contracts driven against the real modules through the real
imports, colocated beside the code they cover. Files whose only content was a
source regex or a re-implementation of the function under test are dropped
rather than translated.
An untouched bot chat was blank: core's splash stands down for any
session that exists, and nothing took its place. Claim the new
`chat.empty` slot and title the chat with the bot's face above its name
in the splash's lettering, so an empty conversation still says whose it
is. Rendering "HERMES AGENT" there would have been the wrong identity
for the surface.
The transcript hands its slot the RUNTIME session id while a canonical
Bot Chat is keyed by its stored one — the two id spaces behind the
#93080 misroute — so identity resolves through the focus store the rest
of the plugin already trusts. Metadata is read with `botRosterMeta`
rather than by name: it is keyed by the route it came from, so a by-name
read misses the entry a bot's avatar and rename actually live under.
The name, not the stack, sits on the center line; the face hangs above
it by half its own block.
Core has one empty state, the intro splash, and it belongs to a fresh
draft with no session selected. A session that exists but has nothing in
it yet falls outside that, and whoever opened it is the only one who
knows what should stand in the gap — core has no business learning a
bot's face and name.
Add `chat.empty` as a contribution area, resolved by the transcript the
same way `transcript.directives` resolves an inline widget. A
contribution mounts for every empty session and renders nothing for the
ones it does not own, so it can subscribe to its own stores and appear
once they load.
Pull the splash's lettering out of `intro.tsx` into a `Wordmark`
primitive so a second caller cannot fork it, and move the typeface,
weight and tracking into `.wordmark` beside the `.fit-text` rules they
travel with — as utilities they were one class-merge or one missed scan
away from silently rendering the wordmark in body text. `Wordmark` takes
its width, because fit-text sizes to fill and a short name set at the
splash's full width comes out enormous.
Creating a bot from the desktop dialog builds the profile tree but no
config.yaml, so the profile resolves no provider and its first turn dies
with "No LLM provider configured" — created, but unable to run. Every bot
made that way was dead on arrival.
Seed the active profile's model block at creation. It is a copy, not a
link: profiles stay independent islands and editing either afterwards never
touches the other. "Fresh" means fresh skills and SOUL, not unreachable.
Bot Mode kept its own unread map, so a bot's dot and the session dot could
disagree about the same chat. Route it through core's store instead, keyed
by the canonical chat's stored id. A hidden session has no listed row to
read a profile from, so the writers take an explicit profile hint rather
than assuming the live gateway's — otherwise the marker persists into a
bucket that never held it.
The composer also hands a blank repoPath to the branch/worktree rail in a
bot chat. The row already hides itself without a repo and stops probing git
and GitHub with it, so one prop does what a second composer would have.
Contributions could scope themselves to a workspace, and Bot Mode scoped
every sessions pane out of the center. Close the last bot chat and the main
zone had nothing left to render, so it collapsed and the sidebar stretched
across where the app used to be. Bot chats are ordinary tabs in the main
strip, beside session tabs, so the filter goes: workspace mode stays as a
SIGNAL consumers adapt to, never a decision about whether a pane renders.
A structural floor backs it up — a subtree hosting the registered main pane
never collapses, whatever happens to its tabs. The shared "+" now falls
through to an ordinary session when the selected owner has no route of its
own, instead of refusing with a toast.
Also lets a pane contribution declare defaultCollapsed, so Bot Mode's
scheduled-jobs pane arrives as the right-edge rail tab rather than an open
zone. It applies when the pane enters the tree, so a user's expand persists
and is never overruled on a later boot.
Every primitive Bot Mode hand-rolled was one the app already had but never
exported: the Panel master-detail family, SessionStatusDot, ColorSwatches +
PROFILE_SWATCHES, RowButton, DisclosureCaret, formatAgo, translateNow, and
the unread store behind the dot. Export them, with docblocks that say which
hand-rolled shape each one replaces so the next plugin doesn't repeat it.
warmAgent/ensureAgent accept an undefined connectionId alongside null, since
a roster row's is optional and both mean "no explicit source".
Bot Mode arrived as a 16,933-line plugin.js that reimplemented most of the
app: its own scroll container, color palette, status dots, empty states,
buttons, time formatting and cron surface, none of which could follow the
theme. Split it into 41 focused modules and route every one of those through
the primitives core already ships, so a Bot row now renders the same
SessionStatusDot, swatches and age labels as the session row beside it.
User-facing strings move into a plugin locale bundle instead of sitting
inline, and the UI settles on "bot" as the noun (model-facing prompt text
still says "agent"). The codemod scaffolding that drove the jsx() -> TSX
conversion retires with the conversion.
hermes-bots/plugin.js was 16,193 lines of hand-written JSX compiler output
— it imported { jsx, jsxs } from 'react/jsx-runtime' and called them
directly. Because the file was .js, eslint (scoped to plugins/**/*.{ts,tsx})
and tsc (allowJs: false) both skipped it entirely, so none of the design
system, import-fence, or type rules that govern the rest of the app ever
reached the largest UI surface we ship on by default.
Converting it back to JSX is an inverse-compile, not a rewrite, so it is
done by script rather than by hand:
- scripts/codemod/dejsx.mjs rewrites jsx()/jsxs() calls into JSX elements.
735 conversions, none skipped. Comments between children become
{/* … */} containers, since a bare // in children position is text.
- scripts/codemod/verify.mjs proves the result. It recompiles the .tsx
through esbuild — an implementation independent of the codemod — and
compares it to the original after normalizing away esbuild's own
rewrites (quote style, void 0, numeric format, string/template folding,
export hoisting) and alpha-renaming every binding per scope. Output:
IDENTICAL across 333,711 normalized characters.
Two rewrite classes are semantics-preserving but not byte-identical, so
they are named and counted rather than hidden: 12 spread-children folds
(JSX has no spread-children syntax; children={[...xs]} can only be written
{xs}, which React flattens identically and which is the idiom the whole
ecosystem writes) and 1 redundant key prop (passed both in props and as
the third argument; React's jsx runtime never copies key into props).
No behavior change. 16,193 lines become 16,050 of real TSX.
The gateway omits a policy-blocked model from `/v1/models` rather than
marking it, so after the preceding commits a restricted model is simply
absent from the pickers. That reads as "Hermes does not support this"
instead of "your organization disallows it".
Show one line when the org is governed, in the two flows where a user
picks a model. It enumerates nothing: model policy is an allowlist, so an
org admitting a handful of models blocks the whole rest of the catalog,
and graying hundreds of rows would be a worse UI than omitting them.
Driven by the `policy_present` claim, which is tri-state — the line shows
only when it is explicitly true, because an absent claim means an older
mint rather than an unrestricted org. The claim is stamped at mint time,
so the line can lag a policy change by up to the access token's lifetime.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The `/model` picker warms `provider_models_cache.json` in parallel before
its serial build loop, and Nous was collected into that prefetch because
the credential scan treats any auth.json providers entry as credentials
regardless of auth type.
Nothing reads the result. The picker's nous branch builds from the curated
list rather than `cached_provider_model_ids`, and Nous cannot reach the
api_key-only unified pathway that would call it. Because the prefetch
forces a refresh it also skips the cache read, so the entry is written and
never read — a live authenticated /v1/models round trip per picker open
for nothing.
Exclude it. Also add the plan this and the preceding commits implement.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.
The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.
Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.
Narrow all four against the authenticated catalog:
- `_login_nous`, which chooses the model the session starts on
- `_model_flow_nous`, the `hermes model` picker
- `list_authenticated_providers`, the `/model` picker
- `/api/model/recommended-default`, dashboard onboarding
The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.
The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.
For an org with no policy — the common case — the filter is a no-op and
every list is what it was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A Nous team admin can restrict which models and which serving providers
their org may use. The inference gateway applies that policy to
`GET /v1/models`, omitting blocked rows with no marker field, so the keys
of an authenticated catalog read are the reachable set.
Add the two pieces the pickers need:
`nous_policy_present()` reads the `policy_present` claim off the OAuth
access token, which costs no request. `/api/oauth/account` does not carry
the claim, so this reads the token rather than going through
`get_nous_portal_account_info`. The claim is tri-state — absent means an
older mint, which is not the same as "no policy" and must not be reported
as one.
`nous_policy_allowed_ids()` turns the authenticated pricing response into
that set, reusing the cache entry a caller asking for pricing already
populates rather than issuing a second round trip. It returns None —
"leave the list alone" — for an org with no policy, for an anonymous read
whose catalog is unfiltered, and for an empty read, each of which would
otherwise narrow a list on evidence that cannot support it.
`restrict_to_nous_policy()` applies the set while preserving the caller's
order, and keeps a `:free` sibling whose base model is reachable. The
gateway admits a row when any of its requestable ids passes and treats
anything unknown as a keep, on the grounds that over-listing costs a 403
from the authoritative gate while hiding a row the gate would serve is
unrecoverable from the client. This mirrors that.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.
That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.
Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.
`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.
Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback
Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.
Credit: @fedosis original PR #94996.
The timeout validation (isinstance(raw, (int, float)) and not
isinstance(raw, bool) and raw > 0 → float(raw)) was duplicated
between auxiliary_client.py:_fallback_entry_timeout and
conversation_compression.py:resolve_compression_fallback_route.
Extracted into _coerce_positive_timeout to prevent drift if the
config schema evolves.
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.
Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
Clicking a bot row moved the layout interaction tracker to the sidebar
group, so $focusedStoredSessionId fell back to the primary selection —
which is null in Bot Mode, because bot chats open as tiles and never set
$selectedStoredSessionId. The Bots plugin reads that null 'focused'
edge as 'the chat lost the center', releases its open claim, and the
Bots home re-asserts over the still-visible chat: the UI jumps to the
list instead of staying in the chat.
$focusedStoredSessionId now answers from the main zone's active tile in
Bot Mode before falling back to the selection, so a chrome-sidebar
click no longer fabricates a null edge; a genuinely closed chat (no
tile in main) still surfaces null and the home returns as before.
Regression tests cover the sidebar-click case (red before the fix),
plus guards for the closed-chat and sessions-mode derivations.
The write-durability contract test (test_repair_path_has_no_bare_connects)
enforces that every repair-path connection carries the macOS write
barriers; the restore rewrites the file header, so it belongs under the
same rule.
apply_durability_barriers() is the guest-connection entry point for
secondary state.db users (async delegation ledger) that must not run
journal-mode setup. database.synchronous (#90892) rides on the
journal-mode/pragma paths guests now skip, so apply it here directly —
otherwise guest connections silently run at the compile-time default.
Follow-up to the #93012 salvage.
Review feedback on #89681: the restore issued the switch pragma
directly via _set_journal_mode_no_wait, which bypassed the
vulnerable-SQLite WAL-reset gate — on the reporter's own runtime
(SQLite 3.50.4) it re-enabled WAL on the one file that had just
demonstrated it can corrupt, exactly what the gate exists to prevent
(a rebuilt file is a new database). It also skipped the WAL companions
(size limit, checkpoint barrier, synchronous=FULL) and used the
leave-WAL helper for an enter-WAL switch.
The restore now calls apply_wal_with_fallback, inheriting the gate,
the macOS-NFS silent-refusal handling and the companions from the
front door, and the pre-surgery mode is probed (best-effort; a
malformed file may refuse) and recorded in the report so the
before/after WARNING only fires when the comparison is honest.
Corruption can drop the WAL bit from the database header, and every
repair strategy rebuilds or rewrites the file in place — so a repaired
store comes back in the default journal mode (delete). The WAL-reset
gate at open time never sees the flip because it happens inside the
repair path, not at open (the open-time flip #89393 warns about is a
different door), leaving the operator no signal that a WAL store moved
to DELETE (#89674).
After a successful repair, re-apply the canonical database.journal_mode
through _set_journal_mode_no_wait (concurrent openers abort the flip
instead of sneaking it between transactions) and log a WARNING naming
the old and restored modes. Best-effort: a refused restore is logged,
never raised — the repair itself already succeeded. Probing the damaged
file for a pre-repair mode is deliberately not trusted: the corruption
itself is what drops the WAL bit, so the configured setting is the only
reliable target.