Covers the precise pinned-session resolver added in PR #88690: request
param shape, the preferred_session response field, hidden-row and
compression-lineage resolution, and older-gateway behavior.
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI
3.0 mode) as a rerunnable eval under evals/browser-use/, following the
toolperf_abeval / evals-compaction pattern.
- tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task
batteries (5 easy, 6 hard) exactly as run for the PR
- single_run.py: one cell = task x arm (base | pr | prns) x model x rep;
throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate
trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE
- orchestrate.py: resume-safe local-CDP battery driver
- orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use
provider plugin, browserbase via REST) with per-cell session lifecycle
- report.py: scorecard aggregation with vs-base token deltas
- README.md: design, run instructions, and the recovered Aug 8-10 2026
baseline scorecards (hard battery, backend matrix, easy round 1,
digest ablation)
The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness
and readouts were recovered verbatim from the benchmark session's tool-call
history in state.db, with hardcoded paths parameterized. Smoke-verified
live: report.py aggregation, and single-cell runs (pr + base arms) against
a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.
Follow-up to #88664. The remote @mention delivery path had three gaps that
made cross-machine DMs half-work:
1. No reply relay: deliverRemoteRosterMentions submitted the prompt and
toasted, but never polled — the handoff note promised a relay that never
came. Now a bounded poll (same shape as a group member turn: new
assistant message after the baseline, 180s cap) relays the recipient's
reply as a notification, or says it's still pending.
2. No sender attribution: the raw user text was submitted, so the
recipient's messaging protocol never recognized an agent-to-agent
message. Deliveries now carry the standard
"Message from 🤖 <sender> (@handle):" prefix.
3. Duplicate Bot Chats: every mention minted a fresh "Bot Chat" session.
ensureRemoteCanonicalChat now resolves the recipient's pinned canonical
chat from its profile ui_meta, falls back to resume-by-title, and only
creates when neither exists — mirroring ensureGroupChatSession.
Tests: remote-dm-delivery.test.mjs (pin-resume without create, attribution
prefix + reply relay via vm-run behavior tests, source contract for the
bounded poll). Plugin suite 187/187.
A resume racing a profile/connection swap can 404 on a backend that
does not own the session; the terminal-failure branch then dropped the
window to the blank new-chat route while the target session was alive.
goneSessionVerdict() now gates the draft fallback: a session created
this run, still listed on some profile, or looked up while a gateway
swap is in flight arms the bounded auto-retry latch instead of
discarding the route. Draft remains the calm-conditions path for
verifiably dead ids.
The BOTS sidebar previewed each profile's most recently active session
(last_session) but clicking the row opened the pinned canonical chat —
two different session identities, so the preview described one
conversation and the click landed in another.
- profiles.list gains an optional preferred_session_ids param
({profile: session_id}): an exact, existence-checked per-profile
lookup that resolves hidden rows and compression lineages to the
live tip (the same resolver session.resume uses) and returns a
preferred_session summary alongside the unchanged last_session.
- The hermes-bots plugin sends its canonical-chat pins with each
roster poll and previews preferred_session ?? last_session.
- openBotCanonicalChat verifies pins through the precise resolver
instead of a paginated, hidden-excluding session.list window that
misjudged real hidden pins as gone; transient lookup failures no
longer clear the pin or mint a replacement chat.
- Grandfathering: a bot with history but no pin adopts the previewed
session on first open instead of minting a new empty chat — the
behavior the design comment already promised.
Closes#88200
A failed intro used to clear the pin even though the chat was already made, so the next click opened a second one. A missing pin grabbed the newest session, even when a real Bot Chat was in the list.
Failed intros now keep the pin. A missing pin looks for a session titled Bot Chat. If that is not there, we try the stored pin instead of the newest row.
Ported from NousResearch/Hermes-Bot-Mode#59 after Bot Mode moved in-tree.
Community asks (Discord, Aug 17): unclear what happens across cloud vs
desktop, whether every bot replies in group chats, and how to persist
connections to multiple gateways.
Extends the Bot Mode user-guide page (landed on main today) with the new
cross-connection features:
- "Create on" picker: creating an agent on another registered machine, with
the remote-target caveats (clone source, staged capability checklists,
draft discard).
- Group chats: explicit "not every bot replies" explanation of the
round-robin/pass model, and rooms spanning machines with device badges.
- @mentions across machines via the Connections registry (no gateway switch).
- Bots-across-machines section: persistent SSH inventory, last-known rows,
and the stay-in-your-chat interaction model; cloud+desktop recipe.
- desktop.md Bot Mode section links to the full guide; multi-connection
page's Bot Mode reference points at the docs page instead of the old
standalone repo.
Builds on the salvaged #88598 multi-source roster work:
- New Agent "Create on" picker (multi-connection registries only): the
profiles.create/configure/describe/mcp.catalog calls route to the picked
connection's backend via host.requestProfile route descriptors — the
window's active gateway never switches. Remote-target drafts discard via
the remote CLI; appearance/title write into the remote profile's ui_meta
and asset store; the taken-name check is scoped to the target machine's
roster; the live SkillsView Capabilities tab (active-gateway-bound) falls
back to the staged checklists for remote targets.
- Group chats can seat bots from other registered connections: member turns
(session.create/resume, prompt.submit, reply polling) route to each
member's own source through the new requestForBot helper. Remote member
descriptors persist on the room record (bot-meta is active-gateway-scoped
by design); watermarks/sessions key by source-qualified member keys so
same-named agents on two machines never share state; @name-device handles
resolve in room mentions; room lines and turn prompts badge cross-machine
speakers with their device.
- pidIsOurDashboard: a dead remote PID (ps exits non-zero) now reads as
FOREIGN instead of throwing "Could not verify SSH backend process
ownership" — the misleading error from #88625's secondary report.
Tests: cross-connection-bots.test.mjs (route descriptors, requestForBot
routing, member keys, disambiguated mentions, device badges, source
contracts); plugin suite 180/180; electron vitest 220/220.
Stay on the default gateway. Bot Mode still lists every Connections
agent. Clicking a remote row no longer hops the window onto SSH —
@dixie / @bob-spark in this chat resolve against the roster and
Desktop delivers in the background via requestProfile.
Clicking Mac Mini / Spark (the device default row) passed the desktop
pool key as the remote Hermes profile. That profile does not exist, so
the chat never opened. Named profiles (bob, dixie) already sent a real
name and worked.
Also refuse to fall back to this-device's default chat pin when the
remote source did not actually become active.
A leftover sshConnections key made roster polling call
ensureRegistryBackend for every SSH source every ~5s. That spawned
remote dashboards, the mux died, and the renderer hit hermes:api
ECONNRESET / liveness-probe drops.
SSH inventory stays on the cached ls path only. Clicking a bot still
dials that one source.
A first inventory miss used to stick forever as a seeded default until
the user hit Test. Retry after 60s; a successful cache still never
re-probes, and Test still forces an immediate refresh.
Remembered remotes were restored whenever the union omitted their
connection id, so a deleted registry source could keep resurrecting
until remount. Only restore rows that still belong to a registered
source.
Clicking the local agent left connectionId null, so the roster treated
the registry primary (often an SSH box) as active and dropped its
profiles while inventing a "This device" shadow of default.
Inventory undialed SSH sources with a cached ls of ~/.hermes/profiles
instead of requiring the window to switch onto that machine. Hostile
HERMES_HOME values are rejected before the listing command runs.
Group chats could be created but never deleted — the only way out was
manually ungrouping every member, which still left the shared room log
in plugin storage forever.
The group-chat workspace header now has a trash button behind a
ConfirmDialog. Disband is soft: it clears every member's group
assignment (syncs cross-machine via ui_meta), drops the room log from
the atom and the persisted group-chats map, clears the needs-you badge,
and closes the room view. The members' per-group gateway sessions
("Group: <name>") are intentionally kept and remain reachable from each
bot's session browser. A room with a drive still in flight leaves a
runtime-only epoch-bumped tombstone so the round-robin loop bails at its
next member boundary; the tombstone is never persisted.
The skew banner told users to run the in-app update, but on machines where
the local desktop pack is the broken step (e.g. the get-windows win32
binding staging failure in #88251), rebuilding in place cannot succeed.
Reinstalling from the packaged installer sidesteps the local build
entirely, so offer it as the escape hatch:
- Settings > About skew banner gains a "Get the installer" button opening
https://hermes-agent.nousresearch.com/ via openExternal
- banner copy now mentions reinstalling when the update doesn't clear the
warning
- all 5 locales updated (new bundleOutOfSyncAction key)
Bot Mode ships built into the desktop app (default on) but only had a
short section in desktop.md. This adds a dedicated user-guide page
covering the Bots roster, creating and editing Bots, avatars, routines,
group chats, bot-to-bot messaging (agent.bot_mode_protocol), the
multi-connection roster, and CLI parity.
An inline ::preview widget could render and be clicked, but the click went
nowhere: the sandbox has no channel to the agent, so an interactive chart
was a dead end. Now the frame injects a second script beside the measurer
that gives the page one voice:
window.hermes.send('get-price eth')
<button data-hermes-send="get-price eth">ETH</button> (zero-script form)
The prompt rides postMessage up tagged with the mount token, then goes
through the composer's own send path (requestComposerSubmit -> prompt.submit)
flagged display_kind=hidden — the same row-typing auto-continue and internal
notifications already use. The agent wakes and takes a real turn; the
durable row persists (context, resume, DB audit); but NO bubble renders,
live or on reload. The user clicks ETH and the chart just changes — the
off-screen loop is click -> hidden turn -> agent rewrites the widget file ->
frame hot-swaps.
Trust boundary matches size reports and is tighter where it matters: mount
token required (frames can't forge each other's intents), string-only,
trimmed, capped at 500 chars, throttled to one intent per second per frame.
The gateway whitelists display_kind to "hidden" — the RPC can't mint
arbitrary row types — and the flag threads through both turn paths (inline
and compute-host isolation) so isolated sessions don't resurrect bubbles on
resume.
The desktop platform hint teaches the model to wire interactive widgets
with data-hermes-send and to answer clicks by updating the widget's file
rather than with prose; the SDK doc documents the contract.
Completes the project-local skills epic's remaining skill items (#48974,
#48975) on top of the discovery/trust work in #88566.
Quarantine (#48974): trust is a repo-level decision made once, but repo
skill content changes with every pull — the hub install path scans, a
checkout didn't. Every project SKILL.md dir now runs through the same
skills_guard scanner as hub installs (content-hash cached under
~/.hermes/cache/project_skill_scans/, never inside the repo). Verdict
'dangerous' quarantines the skill: excluded from the index, skills_list,
and slash commands via the single iteration chokepoint
iter_project_skill_files(), and skill_view refuses by name with an
explanatory error. Scanner failure fails closed. Verified against a real
injection fixture (6 findings: prompt_injection_ignore, deception_hide,
invisible_unicode, credential exfil patterns).
Non-interactive inheritance (#48975): find_project_root() now resolves
from TERMINAL_CWD (the per-surface workdir cron jobs and the terminal
tool already use) before falling back to process cwd. Cron/API/ACP
surfaces inherit a prior interactive trust decision by project identity:
job workdir inside a trusted repo => project skills load; untrusted or
no workdir => nothing loads; no surface ever prompts.
Tests: +10 cases in tests/agent/test_project_skills.py (real malicious
fixture, fail-closed, rescan-on-change, cache location, TERMINAL_CWD
inheritance matrix). Docs: quarantine + non-interactive sections in
skills.md.
hermes update moves the source tree, but the desktop UI (including bundled
plugins like Bot Mode) is compiled into the app binary at build time. A
terminal-side update — or an in-app update whose bundle-swap leg failed —
leaves a new runtime under an old renderer: About reports the new Hermes
version while the sidebar is missing that version's desktop features
(the 'no Bots tab after the Bot Mode update' reports).
Detect the skew by comparing the packaged install-stamp commit against the
tree (git rev-list --count <stamp>..HEAD -- apps/desktop) and surface it:
- Settings > About: amber warning banner pointing at the updater
- macOS native About panel: suffix on the version line
- hermes:version IPC gains bundleOutOfSync / bundleCommitsBehind
Fail-quiet by design (no stamp / fallback stamp / git failure = no warning)
so dev runs and non-git builds never see a false 'install is torn' alarm.
All 5 locales covered.
Exports createBudgetedLoop (+types) through the plugin SDK and migrates the
Bots plugin's face-animation clock onto it. The hand-rolled clock only
checked document.hidden; via the shared loop it now also pauses while the
window is minimized or unfocused, matching every other desktop render loop.
- sdk/index.ts: export createBudgetedLoop/BudgetedLoop/BudgetedLoopOptions
from @/lib/budgeted-loop with plugin-facing guidance
- hermes-bots/plugin.js: startFaceClock delegates scheduling to the SDK loop
when present (fps 15, idleWhen = no visible faces); paint body, IO-based
visibility tracking, and the 1Hz rescan are shared; the hand-rolled rAF
path remains as a feature-detected fallback for older desktops, per the
plugin's established SkillsView/McpTab pattern
- tests: new SDK-path case (fps/idleWhen wiring, visibility wake, re-entry
wake, dispose-on-stop) driven through an injected fake loop; verified to
fail with the SDK branch removed; existing fallback-path tests unchanged
CI's check:lint caught three real issues in the directive surface:
- TranscriptDirectiveLeaf called the contribution's render() inline in JSX
— the exact pattern no-restricted-syntax bans because the callback's hooks
land in the host and a plugin reload changes the host hook count
(React #310). The callback is now memoized and mounted via ContribRender.
- The inline frame mirrored its height state into a ref from the message
handler (the stale-read pattern no-restricted-syntax flags). Functional
setState reads current state directly; the shadow refs are gone.
- perfectionist import/export ordering in the frame and the SDK index.
The SDK doc still described the v1 frame (fixed height attribute, rail card
under the frame) — chrome that no longer exists. And plugin_storage's usage
example used `with plugin_db(...)`, which reads as auto-close but sqlite3's
context manager only scopes transactions; the example now closes explicitly.
The first inline frame was a full-width bordered box at a fixed height:
webpage-in-a-rectangle, not a widget. Now the frame disappears into the
message flow:
- Content-driven size. The injected measurer reports height (live) and
intrinsic width (adopted once, so %-width children can't feedback-loop the
frame toward zero). A sparkline shrink-wraps and sits flush left like an
inline image; a full-bleed page measures the whole viewport and stays
column-wide. The height attribute is now only a starting value.
- Theme bridge. A style prelude injects first with the app's resolved theme
tokens under stable names (--foreground, --muted-foreground, --accent,
--border, --card), the app font, zero body margin/padding, and a
transparent background — reference HTML written against those vars renders
native in any theme. Page styles override the prelude, so a page that
brings its own design keeps it.
- No chrome. Border, rounded box, and the rail-opener card under the frame
are gone; the fallback paths (non-HTML, remote gateway, unreadable file)
keep the classic card. The wheel gate went with the border — frames size
to content, so there is nothing to scroll inside, and widgets are fully
interactive.
- The desktop platform hint now teaches the default: an inline widget is
transparent, token-colored, flush left, no page chrome — only a standalone
page brings its own background. "Make me an inline sparkline" gets native
styling without the user spelling it out.
Some providers re-send the previous assistant text verbatim when a turn
continues past a tool call (a tool_calls row, then a stop row with identical
prose — both persisted). The turn merge folds both rows into one bubble, so
every paragraph in the reply rendered twice; inline ::preview frames made it
obvious. Repeated text parts now dedupe in the same pass as generated-image
echoes — the last occurrence wins.
The frame was a hardcoded 280px unless the model guessed a height attribute
— tall pages clipped (the flip-clock demo cut its last digit), short ones
floated in dead space. The opaque-origin sandbox means the parent can't
measure the document, but we own the srcdoc string: a tiny injected script
observes the document with ResizeObserver and posts its scrollHeight up via
postMessage, and the frame tracks it live within the 120-1200 clamp.
Reports are validated before they can move layout — per-mount random token
(two previews in one transcript, or a hostile page inventing messages,
can't move each other's frames), finite-number check, clamp. A 4px
tolerance stops vh-sized pages (which measure exactly what they're given)
from oscillating; an explicit height attribute still opts out of
auto-sizing entirely.
The first cut of the core ::preview consumer rendered the classic
preview-attachment card — a button into the right rail we already had, which
made the directive indistinguishable from an ordinary preview link. Now the
directive shows the thing itself: the workspace HTML file renders in a
sandboxed srcdoc iframe inline in the assistant message (opaque origin,
allow-scripts only — no reach into the app, its storage, or the bridge),
with an optional height attribute clamped to 120-1200px and the classic
card kept below as the rail escape hatch.
The frame waits for turn settle before reading the file (mid-stream it is
often mid-write), resolves relative paths against the session's own cwd,
and falls back to the plain card for non-HTML targets and remote gateways
(no local file door there).
Plugins that persist state have been writing into their own install tree
(<hermes home>/plugins/<name>/), which `hermes plugins update` git-pulls and
`hermes plugins remove` deletes — user data dies with the code that wrote it.
plugins/plugin_storage.py is the sanctioned home: plugin_data_dir(name) gives
one data root per plugin under <hermes home>/plugin-data/<name>/ (profile-
aware, created on first use, names validated against traversal), and
plugin_db(name) opens a WAL-mode SQLite database inside it. Secrets stay on
the existing secret-scope path — this is state, not credentials.
hermes-achievements, the in-tree offender, converts with a legacy-file
migration on first read.
The transcript becomes a contribution area (transcript.directives). A plugin
registers a named directive and the model addresses it by emitting
::name{key="value"} as its own paragraph; that leaf renders as the plugin's
component, wrapped in the contribution error boundary. Unclaimed or malformed
directives stay plain prose, so nothing changes for text that merely looks
like a directive (std::vector) or for users with the plugin disabled.
Core ships ::preview{file="..."} as the reference consumer (the existing
preview-attachment card), the desktop platform hint teaches the model the
syntax, and the SDK exports the area + types so runtime plugin.js files get
the surface through the normal plugins API.
Reconciles the salvaged Responses-API mandate with the bundled meta-ai
plugin that landed in #88565:
- meta-ai profile api_mode -> codex_responses (prompt caching engages
only on /v1/responses; 0% vs 93-99% measured). Custom endpoints with a
non-api.meta.ai base URL still fall through to chat_completions via
the host-driven mandate design.
- cli-config.yaml.example: point the example at MODEL_API_KEY (Meta's
documented env var) and note the bundled provider covers the default
endpoint
- tests/providers/test_meta_ai_profile.py updated for the new wire
Implement Claude Opus review findings for Meta API support:
- Document in agent/agent_init.py that provider="meta" without an api.meta.ai URL falls through to chat_completions by design (URL-driven wire selection).
- Comment on suppression guard in hermes_cli/runtime_provider.py noting api.meta.ai is handled by _detect_api_mode_for_url.
- Replace inline __import__ with top-of-module import in tests/hermes_cli/test_model_switch_openai_api_mode.py.
- Rename test_meta_retention_not_sent_when_overridden -> test_meta_retention_override_wins in tests/agent/transports/test_meta_codex_cache.py.
- Add test in tests/agent/test_meta_agent_init.py for provider="meta" fallback without api.meta.ai URL.
- Add test in tests/agent/test_auxiliary_client.py for prompt_cache_retention: "24h" under _CodexCompletionsAdapter.
Source: Claude Opus review findings for feat/meta-api-support.
Relocate host_mandated_api_mode check from top of api_mode cascade to
fallback else branch so URL-based provider-slug rewrites (e.g.
api.anthropic.com -> provider='anthropic') always run first. Previously
the mandate branch set api_mode for api.anthropic.com without rewriting
provider, leaving provider='' and causing credential_pool_matches_provider
to fail closed and discard anthropic-scoped pools (#63425 regression
introduced in 8f60e8263).
The mandate is now a true fallback for hosts without an elif branch
(api.meta.ai -> codex_responses for 93-99% prompt-cache hits vs 0% on
chat, plus future mandates) with lazy import + try/except preserved.
Add regression tests: provider=None + api.anthropic.com URL implies
provider='anthropic'/api_mode='anthropic_messages' and preserves an
anthropic credential pool; provider=None + api.meta.ai URL implies
codex_responses.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
Surface meta/muse-spark-1.2 in the Hermes model selector (CLI, desktop,
gateway) via the curated OpenRouter list and regenerated model-catalog.json.
The model is live on OpenRouter with tool calling; the picker intentionally
does not show the full OpenRouter catalogue.
Desktop shipped the same bug class four times in one week: a decorative
animation loop that never sleeps (bots face clock #88543, pixel egg #88406,
diffusion placeholder #88564, plus the #77651 hidden-renderer wave). Each fix
hand-rolled the same four behaviors. This extracts them into one helper in
src/lib/budgeted-loop.ts:
- fps budget (default 15) on top of rAF
- observability pause via the existing createRendererLoopPauseController
- idle dormancy: idleWhen() true after a draw parks the loop with zero
pending work until wake() — the piece every hand-rolled loop forgot
- teardown: dispose() cancels the frame, disposes the controller, and makes
wake() a no-op
DiffusionCanvas migrates onto it as the first consumer (net -40 lines at the
call site); its existing scheduling/budget/instance-cap tests pass unchanged
against the migrated implementation. Helper suite covers budget, pause,
park/wake, dormancy-survives-focus-churn, and dispose idempotency; sabotage
run (budget+dormancy stripped) fails 4/5.