Home Manager separates an installation from a daemon. This module put
both under `services.hermes-agent`, and `installPackage` added a program
to the PATH from a service module.
`programs.hermes-agent` now installs the command line application and
the desktop application. `services.hermes-agent` keeps the state, the
configuration and the daemons, and stays the authority: the new module
reads `hermesHome` and the backend address from it. A person can enable
one without the other, which is a machine with an application and no
gateway, or a headless gateway with no display.
The desktop application needs this split to work correctly. A launcher
that starts from the desktop menu reads no shell profile, thus the
HERMES_HOME that `home.sessionVariables` exports reaches an interactive
shell only. Home Manager writes `systemd.user.sessionVariables` to
environment.d, and this module puts no HERMES_HOME there, because that
file applies to each user unit. The application then opens ~/.hermes
while the services use `hermesHome`, and the person sees no sessions and
no keys. Thus the launcher carries the value itself, through a new
`extraEnv` argument on the desktop package.
The application also gets the Nix agent package, with
HERMES_DESKTOP_HERMES. The usual distribution of the Electron
application carries its own Hermes runtime and downloads more at the
first start. `hermesDesktop` is a passthru of the agent and pins
`finalAttrs.finalPackage`, so an override of `extraPythonPackages` or
`extraDependencyGroups` reaches both. One machine thus has one runtime.
`backend.sessionTokenFile` connects the application to the backend of
the service. Without it the module runs `hermes serve` and the
application starts a backend of its own, which gives two backends on one
HERMES_HOME. The backend reads the file into
HERMES_DASHBOARD_SESSION_TOKEN. The launcher reads the same file into
HERMES_DESKTOP_REMOTE_TOKEN, beside a HERMES_DESKTOP_REMOTE_URL that
names the address of the service.
Measurements against a live `hermes serve` on loopback show why that
shape is the correct one:
- `_resolve_session_token()` reads HERMES_DASHBOARD_SESSION_TOKEN, and
`_has_valid_session_token` accepts that value as a Bearer credential.
A request without it gets 401, and a request with the wrong value
gets 401.
- The /api/ws socket accepts a query parameter only. A header gets 403,
and `?token=` connects. Hermes Desktop builds exactly that URL, in
`apps/desktop/electron/connection-config.ts`. Thus a test of the HTTP
leg alone is a false positive.
- `resolveDesktopRemoteRoute` throws when the URL is set and the token
is not. Thus the two variables travel together or not at all.
The token enters no Nix store path. `makeWrapper --set` and a systemd
`Environment=` value both write a literal into the store, which all
users can read. Thus each side reads the file at start time. The
launcher does it through a new `extraRun` argument on the desktop
package, and the backend through the launcher script that
`backend.waitFor` already uses. launchd has no EnvironmentFile, so a
script is the one shape that works on Linux and on Darwin.
`backendArgv` gives the plain argv only when nothing must run before
the backend.
`services.hermes-agent.installPackage` is removed. It defaulted to true,
so a person who never named it still got the command line. A silent
removal thus gives them a machine with no `hermes` and no message. The
module refuses a configuration that sets it, and the text names the
exact replacement for the value they gave.
Checks:
- the launcher carries HERMES_HOME
- the launcher reports HERMES_MANAGED only when the services own the
configuration, because no activation writes a marker without them
- the launcher pins the agent package that `programs.enable` installs
- the launcher names the backend of the service, and gives a token
beside the URL
- the backend reads the session token
- each side reads the file at start time, and the token is no `--set`
value
- `programs.enable` alone starts no service
- `installPackage` is refused, with a message that names the
replacement, and its absence evaluates
Each check reads the wrapper of the real package, and not an option
value. Each one was tested with a mutation that breaks the behavior it
asserts.
Clicking a bot in the roster always reopened its pinned canonical Bot Chat.
Start a new conversation with bot A, click bot B, click back to A — the new
conversation was gone, replaced by the pinned transcript. A bot row is a
workspace entry point, so it has to land on the live conversation.
Two independent causes, both fixed here:
1. The pin overrode newer work.
`openBotCanonicalChat` opened the pin unconditionally. It now prefers the
bot's freshest VISIBLE session — but only AFTER `profiles.list` has
verified through `preferred_session` that the pin is alive and is a real
canonical Bot Chat. That ordering matters: with a dead or unverified pin,
adopting the profile's latest row would claim an unrelated user
conversation as the bot's chat, and the hide sweep would then hide it.
The existing "no pin" / "dead pin" safety tests cover exactly that and
still pass. The pin keeps owning plumbing (creation, hide sweep, DM
delivery); it just stops shadowing newer conversations.
Guards on the candidate (`newerVisibleBotChat`): the canonical chat can
never shadow itself, an empty draft never displaces a real conversation,
and a gateway that omits `message_count` is treated as real history
rather than discarded.
2. The workspace did not follow the bot.
The three `host.openSession` calls on the bot path relied on the SDK
default `keepAllProfilesScope: true`, so `$activeGatewayProfile` stayed on
whatever profile was active before the click. Sessions created afterwards
were then filed under the previous bot's profile — measured: four new
chats started from three different bots all persisted into one profile's
state.db. Clicking a bot IS a profile switch, so these pass `false`.
Note on the call shape: `previewSession` is `bot.preferred_session || last`,
so on a pinned bot it resolves to the PIN (preview identity must match click
identity). Feeding that as the "newer" candidate makes the whole preference
dead code — it always sees the pin and short-circuits on "same id". The
freshest visible session therefore arrives as its own argument. The first
attempt at this fix had that bug and passed its tests, which is why
`bot-row-opens-latest.test.mjs` mirrors the production call site argument for
argument rather than constructing a convenient one.
Tests: 362 pass (was 348). Each new guard was verified by sabotage — reverting
any one of the three behaviours above makes the suite fail (1, 3, and 1 tests
respectively), so none of them is a test that passes either way.
resolve_exec_command wrote the repo hermes script (env-python shebang)
straight into Exec=; spawned by the DE that shebang escapes the venv and
dies on the first import, invisibly (Terminal=false, entry rewritten
every launch). A python-script launcher whose shebang points outside the
running interpreter's env now gets Exec={sys.executable} {script} desktop;
native binaries, bash wrappers, and venv-shebang scripts are untouched.
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.
Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
PaneTab gated its hover close button on two independent inputs: the
onClose verb, and a showCloseButton prop that TreeGroup fed from a
showCloseButton flag on the pane contribution. The middle-click and
Meta-click gestures read only onClose. A tab could therefore close on a
pointer gesture and advertise no control for it.
The flag had no user that hideOnly did not already cover. Both setters
also set hideOnly: true, which removes every close gesture:
- the sessions pane (app/contrib/controller.tsx),
- the Bots pane (plugins/hermes-bots/plugin.js).
The flag was an opt-out marker with no reachable effect, so this change
deletes it instead of teaching it to track the gestures. onClose alone
now decides both shapes. A tab that closes shows the button. A tab
without the verb shows nothing. To make a tab uncloseable, give it no
close verb.
hideOnly and uncloseable keep their meaning. They gate the verb, and
both shapes follow the verb together.
The DialogContent and SheetContent prop of the same name is a different
prop and stays. It has no close verb to derive from, and one caller
changes it while the dialog is open.
Tests: the new tab-close-affordance test renders the real TreeGroup and
asserts that button presence equals middle-click closure. It covers
hideOnly chrome, a plain side pane, the uncloseable workspace, and a
session tile. It reads closure from the layout tree, not from a spy, so
a wired-up mock cannot pass it. A regression that hides the button on a
closeable tab fails two of the four cases. The compiler rejects the
deleted prop, so the test carries no fixture for it. The pane-tab unit
test moves off the deleted prop.
Verified with the full apps/desktop vitest suite, npm run typecheck, and
npm run lint. Two electron process-spawn tests fail on this machine.
They also fail on a clean tree, and they do not touch the pane shell.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
cron_delivery_targets() now also lists machine-local bot-chat:<profile>
entries; the sibling test's exact set-equality assertion predates them.
Scope the platform assertions to gateway entries and pin that bot-chat
entries are always home_target_set.
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.
- cron/scheduler.py: token parsing, target resolution (own profile /
named local profile / unknown -> skipped with warning), subprocess
delivery lane with cron.bot_chat_delivery_timeout_seconds (default
600s), preflight exemption, and bot-chat entries in
cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
must exist on this machine (fail at create, not at 3am); deliver schema
documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
token on the profile-scoped create so Desktop-side aliases can never
name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.
Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
The backend binds to `backend.host` immediately. The bind fails when the
target is not ready, because uvicorn cannot bind a name that does not
resolve, or an address that no interface holds. A unit that starts at boot
loses this race against the daemon that supplies the target, such as
tailscaled.
A bind to a Tailscale MagicDNS name shows the problem. The name is the
correct bind target, because the dashboard refuses each request with a Host
header that is different from the address that the server bound to, and a
shared machine has a different address in each tailnet. But the name does
not resolve until tailscaled is up, so the unit fails at each boot until
`Restart=on-failure` finds the moment when the name works.
A systemd user unit cannot order itself after a system unit. `After=` and
`Requires=` are silent no-ops across that boundary. Thus the wait is a poll,
and not a dependency.
This change adds three options to `services.hermes-agent.backend` on both
the NixOS module and the Home Manager module:
- `waitFor` — `null` (the default, unchanged behavior), `"hostname"`, or
`"interface"`
- `interfaceName` — the interface to take the address from
- `waitTimeout` — the time in seconds before the unit stops
With `waitFor`, ExecStart becomes a launcher that polls for the target and
then execs hermes. `exec` keeps hermes as the MainPID, so the restart logic
of systemd sees the real process. A timeout stops the unit with an error. It
does not bind a fallback address, because a fallback can expose the backend
more widely than the user intends.
The default is not changed. Without `waitFor`, ExecStart is the same
command line as before.
The strip could only be hidden by an undiscoverable double-tap, and once hidden
the zone had no chrome left to click — no tab, no ✕, no menu holding "Show".
This puts it on the same footing as the status bar, whose hide has never
stranded anyone: ⌥⌘T, a ⌘K row, the shell context menu, and the zone menu, which
now prints the keystroke on the row that takes the strip away so the way back is
stated at the moment it matters. All four resolve their target zone the same way
the other tab verbs do (hovered, else focused, else the workspace) and describe
themselves from what is on screen rather than from a stored value, so "toggle"
always means the opposite of what the user is looking at.
Adds an app-wide default alongside it, in Appearance next to Session List
Density — auto, always, or never, matching VS Code's `workbench.editor.showTabs`
and Zed's `tab_bar.show` for people who want one answer everywhere instead of a
per-zone choice they repeat. A zone that has stated its own preference still
wins, and neither value can strand a pane.
`headerHidden` carried two meanings at once. `true` was either "the user hid
this" or "a double-tap nobody meant hid this"; `false` was either "the user
wants a strip" or "insert / tab-cycling / dock-enforce / adoption pinned one to
escape a dead end". Because the layout wrote the same field the user did, a
repair silently overwrote a preference and neither could be read back — and
since hiding also unmounted the tab, the ✕ and the menu offering "Show header",
a zone that got hidden by accident stayed that way across restarts.
Replaces it with `tabStrip?: 'always' | 'never'`, where absent is auto and only
the user ever writes it, and moves the decision into one resolver that TreeGroup
and the store both call, so the strip on screen and the toggle command cannot
disagree. Reachability moves into that resolver as an invariant that outranks an
explicit `never`: a closeable tile keeps its ✕ and a lone tool panel keeps its
chip, because "hide the chrome" is never a request to make a surface
unreachable. With that guarantee held centrally, the four repair writes are
gone. Persisted `headerHidden` is dropped rather than translated — nothing on
disk distinguishes a deliberate hide from an accidental one, and carrying the
accidents forward would re-strand exactly the people who reported being stuck.
The double-tap hide goes with it, along with the synthesized double-tap detector
it was the only consumer of. It fired from ordinary double-clicks on a tab,
nothing announced it, and its undo lived behind the chrome it had just removed.
`data-zone-no-header` goes too: it marked full-page views for a body
double-click toggle that no longer exists, and nothing has read it since.
Supersedes the tab-side half of the fix from abundantbeing and yoniebans, whose
commits this builds on.
Narrows #86278 to exactly the defect. Tabs pass no double-tap context on any press path (generic pane drag, multi-tab selection drag, chrome.tabDrag), so a double-click on a tab can no longer hide the strip; the strip background keeps its documented hide gesture unchanged.
The body double-tap reveal from #86278 is dropped: the zone body deliberately carries no double-click gesture (virtualized content recreates its nodes between clicks, per the standing ruling in tree-group.tsx), and recovery surfaces for a deliberately hidden header are being decided separately across #84458 / #81638 / #89225. The DOUBLE_TAP_MS export is reverted since no consumer remains outside drag-session.
Test file trimmed to the two assertions that pin the grammar: a tab double-tap must not hide the strip (red on main), the strip background double-tap still hides. Taps release on window between presses so the drag-session synthesized double-tap path is the one exercised.
The synthesized double-tap that hides a zone's tab strip rode every tab's
pointerdown (generic pane drag and each pane's tabDrag), so a routine
double-click on a tab (select a title, retry a click) vanished the whole
bar and stranded the zone with no tab, no close X, and no way back but a
right-click. Keep the documented hide gesture on the strip background
only, and add its inverse as recovery: double-tap a hidden zone's body
restores the strip. Regression tests pin both sides of the grammar.
_turn_transcript_messages pre-classified every message with
_is_compressed_summary_message (full content flatten + prefix scan), then
_message_response re-ran the same classifier inside its projection --
2x per non-summary row, 3x per summary row on every run.completed emit.
The outer guard was redundant: _message_response already yields
display_kind hidden for pure handoffs. One projection call per row now.
Surfaced by the post-merge simplify re-review of #91517/#91535.
Artifact receipts live only in memory, so files left behind by a dead
process were unreachable but persisted forever despite the advertised
300s TTL — a retention failure on the surface meant to be ephemeral
(blocker 4 of andrexibiza's #91535 review). A fresh ArtifactStore now
removes every artifact-id-shaped file and stale *.tmp with no index entry
(at construction the index is empty, so all such files are orphans).
Non-artifact-shaped names are untouched. Regression: store -> recreate
store over same root -> orphan+tmp gone, unrelated file kept.
The global broker snapshotted browser.extension_control.developer_mode once
at construction, so flipping it OFF in config did not revoke raw CDP/eval
from already-attached controllers until process restart — a revocation
failure at the highest-privilege browser surface (blocker 3 of
andrexibiza's #91535 review). select() now consults the live config on
every privileged selection (explicit bool still pins for tests); off->on
also unlocks without restart. Regression test drives both directions
against an attached controller. Also drops the dead back-compat
_artifact_store property (zero readers).
The frame handler returns reply dicts (heartbeat/detach acks) that the WS
reader loop sends back; the -> None annotation was the only new ty
diagnostic vs origin/main.
Addresses both merge blockers from @andrexibiza's review of #85351:
1. HTTP-uploaded artifacts could never be consumed by broker dispatch:
artifact_scope_key hashed (principal, session, family), the HTTP routes
store with an EMPTY session (API-key auth has no server session) while
broker validation carries a session-bearing ControllerScope — every
real upload->dispatch journey died with ArtifactScopeMismatch
(reproduced before fixing). Canonical ownership is now
principal/transport-family (documented in the scope-key docstring);
ids stay unguessable server-minted 32-hex and downloads one-shot.
New composition regression: HTTP-shape upload -> registered controller
scope -> broker artifact dispatch, mutation-checked (re-adding session
to the key makes it fail).
2. The 'profile-scoped' artifact store was first-profile-wins process
state: one adapter-level singleton pinned profile B to profile A's
physical root on multiplex listeners (same frozen-handle class as
#88734). Stores are now cached by resolved profile, and the broker
selects the store from the controller scope's profile_id (default-slot
fallback preserves single-profile/test behaviour). New A/B multiplex
regression proves distinct physical roots regardless of touch order.
Also documents the advertised ticket_expires_at as best-effort wall clock
(broker enforces expiry monotonically) per review feedback.
browser_control_enabled()/browser_control_developer_mode() run on every
browser tool call and inside every check_fn evaluation (uncached for bound
sessions). Both are pure reads of nested dicts; load_config()'s defensive
deepcopy (~135us/call) is wasted there. Same pattern as the other read-only
config probes.
Surfaced during review of PR #85351.
- Rename the broker's TicketInvalid to ControllerTicketInvalid: the same
exception name already exists in hermes_cli/dashboard_auth/ws_tickets.py
and BOTH are caught in the same WS auth flow this feature touches — two
unrelated same-named exception types in one blast radius invited a wrong
except clause.
- Import the 'server-internal' sentinel identity from its canonical
definition (ws_tickets.INTERNAL_USER_ID/INTERNAL_PROVIDER) instead of
re-declaring the strings; drift would have silently broken the
internal-peer exclusion in _is_authenticated_identity.
Surfaced during review of PR #85351.
attach/disconnect/detach acquire a per-controller threading.Lock that a
worker-thread dispatch can hold for up to 10s while blocking on the event
loop to transmit its command frame (run_coroutine_threadsafe +
result(timeout=10)). Acquiring that lock synchronously from loop context
(controller WS finally, frame handler, gateway WS teardown) could park the
ENTIRE gateway event loop behind the send bridge — a deterministic
multi-second global stall whenever controller teardown raced an in-flight
command. All loop-context broker calls now go through asyncio.to_thread,
matching the existing offload pattern for _close_sessions_for_transport.
Surfaced during review of PR #85351.
The router treated any server-stamped principal as a bound lane, so with the
flag ON every authenticated dashboard/API session lost the legacy browser
backend even when no extension controller ever registered (scope_for_session
returns None -> ControllerUnavailable, no fallback) — while check_fns still
advertised the tools via the legacy OR-gate.
New broker.lane_registered() distinguishes the two cases:
- lane never registered -> generic callers keep the legacy backend
- lane registered (controller offline/ambiguous) -> fail closed, unchanged —
a control-this-tab session never silently jumps to another browser
Also makes the four non-allowlisted wrapped tools (cdp/console/vision/
get_images) behave correctly for never-registered lanes (legacy backend)
while staying fail-closed for registered lanes.
Surfaced during review of PR #85351.
The feature flag was only documented in cli-config.yaml.example; every other
browser.* key is declared in DEFAULT_CONFIG so config tooling (dashboard
editor, hermes config get) can see it. Defaults unchanged: enabled=False,
developer_mode=False. Surfaced during review of PR #85351.
Generic Hermes callers still use the existing browser backend when extension control is disabled or no server-bound controller identity exists.
Once the gateway binds a controller identity, missing scope, disconnect, or capability loss now fail closed instead of silently switching a control-this-tab request to another local or cloud browser. Covers the schema-build to dispatch disconnect race.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
The full-screen boot surfaces (connecting, onboarding, boot failure, root
crash fallback) paint their backdrop with --ui-chat-surface-background,
which the glass field turns transparent so <body> can be the one painter
(0483133842). That was harmless while glass shipped off; once it shipped
on by default (be3166607e) every boot overlay became a window onto the
shell behind it.
These overlays mask the whole app, so they declare data-glass-opaque —
the existing contract for surfaces that paint over siblings — which pins
the token back to opaque chrome under glass and changes nothing when
glass is off.
The 7-key internal-fields tuple was inlined twice (agent/compaction_display.py
and _project_client_message); a drift between the copies would silently leak
one internal field class through the API projection. Surfaced during review
of PR #85442.
Project client-visible session messages through the canonical compaction classifier. Hide standalone handoffs, unwrap merged carriers to their authentic prior-tail content, strip inherited internal fields, and keep model-facing recovery history unchanged.
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
The Codex OAuth backend (chatgpt.com/backend-api/codex) intermittently
injects prompt_cache_retention into its own upstream call and then rejects
it, returning HTTP 400 invalid_parameter. Hermes never sends that field on
this route (see agent/transports/codex.py::_default_prompt_cache_retention_
for_request, which only sets it for api.meta.ai and bedrock-mantle hosts).
Reproduced live: a minimal 1-message request carrying no cache parameters
at all failed 4/20 (20%) with this error, so the rejection is not
deterministic and retrying the identical request is the correct recovery.
Previously the catch-all in _classify_400 returned format_error/
retryable=False, which tripped the is_client_error abort gate in
conversation_loop and killed the turn on the first attempt - burning an
entire large-context request (~550k tokens) per failure.
Classify these as retryable server_error (should_compress=False - the
request shape was never the problem). The same guard is applied to the
sibling 5xx request-validation branch, where a fronting proxy can surface
the identical rejection.
Deliberately narrow: keyed on parameters we only send on specific routes,
and skipped when the current provider is one that legitimately sends them,
so a genuine client-side bad parameter (max_tokens on GPT-5) still fails
fast as a format_error.
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.
The fix:
- _is_summary_item delegates to the canonical
agent.context_compressor.is_compaction_summary_message provenance check
(not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
through to the pruner, so it can read summary content directly from the
source when the Responses conversion shape is lossy (tool-result carrier
becomes function_call_output, or stale codex_message_items replay shadows
merged content)
Fixes#90975.
Salvage of #90976 by @JoaoMarcos44.
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.
One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
configured/logged_in=True with key_source 'keyless' — flows through
get_auth_status to every status consumer (hermes status, dashboards,
list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
get has_creds=True before any env/pool/auth-store checks — this is
the source for /model, the TUI picker, and the desktop
/api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
providers are kept — there is nothing to 'explicitly configure', and
hiding a zero-setup provider defeats its purpose.
E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
updateGroupChat's inline durable-map builder (the local-mutation persist
path) skips tombstoned rooms and carries roomId — but durableGroupChatRooms,
the SEPARATE builder persistGroupChatRooms uses for the remote-merge path
(every pullGroupChatServerState / gateway-swap sync), has neither.
Two independent gaps in the same function:
1. Tombstone resurrection. Disband sets a runtime-only tombstone
({tombstone: true, log: [], ...}) while a drive may still be mid-turn,
with no roomId. mergeRemoteGroupChatSnapshotIntoRooms spreads
...existing before its explicit field overrides (none of which touch
tombstone), so if a remote gateway hasn't received the delete yet
(plausible now that sync fans out to every reachable default-profile
gateway with independent per-connection backoff) and still has a live
copy of the room, the tombstone flag survives into the merged room.
That merged map is handed straight to persistGroupChatRooms, which
wrote it to storage because durableGroupChatRooms had no tombstone
check. On the next cold hydrate the persisted tombstone reads back as
an empty, non-tombstoned room, resurrecting the original bug
(recreating a room under the same name silently becomes "<name> 2")
through a path the earlier tombstone fix didn't cover.
2. roomId loss. mergeRemoteGroupChatSnapshotIntoRooms correctly carries
roomId into the merged in-memory room, but durableGroupChatRooms never
included it in the persisted snapshot. Every room merged in via the
remote-sync path therefore loses its immutable room identity on the
next cold hydrate (comes back with roomId: null) and falls back to
legacy name-keyed identity — breaking id-based rename/merge resolution
and member-session titling ("Group: <roomId>").
Fix: durableGroupChatRooms now mirrors updateGroupChat's inline map
exactly — skip tombstones, carry roomId.
Tests: durableGroupChatRooms unit tests for both gaps, plus an
end-to-end reachability test (tombstone -> mergeRemoteGroupChatSnapshot-
IntoRooms -> persistGroupChatRooms -> storage) proving the merge really
does forward the tombstone and the fix really does keep it out of
storage. Mutation-verified against pre-fix code (all 3 new tests fail).
Full hermes-bots plugin test suite (60+ files) green.
Review-round residuals: the spinner's user-select guard now beats
[data-selectable-text] regardless of stylesheet order; the will-change
layer hint clears under the global renderer pause and reduced motion so
parked spinners hold no compositor layer; the e2e travel assertion reads
the engine's keyframes (a computed transform always serializes to a
matrix, so the old '%' check could never fail); inline import() type
hoisted for the lint gate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>