Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
The Bots editor's model write (profiles.configure) was the one switch
surface that bypassed the data-policy / expensive-model selection guard:
a guarded pick (e.g. muse-spark contributor tier) was applied silently,
with no confirm flow anywhere — the #95293 remainder after the core
picker's confirm handshake landed in use-model-controls.
Gateway: profiles.configure now answers confirm_required +
confirm_message for a guarded model (same handshake as config.set
model) and writes NOTHING until the client resends with
confirm_expensive_model: true. Other sections still apply; the pending
model section is not reported as failed.
Desktop: the confirm flow is extracted out of use-model-controls into
one shared applier (lib/guarded-model-switch.ts, exported through the
plugin SDK) — warning toast, staleness-guarded Confirm, single
confirmed resend, never a retry loop. The core picker and the Bots
editor now consume the SAME handler; the Bots editor's Confirm resends
only the model section with confirm_expensive_model: true.
Fixes#95293 (Bots surface remainder).
profiles.list opened every profile state.db as a writable SessionDB,
which waits out write-lock patience while that profile's backend is
mid-turn. The desktop RPC timed out and Bot Mode's infinite React
Query retry kept the sidebar on a spinner.
Inspect those DBs read-only and bound roster retries so names still
paint.
A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.
Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.
Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.
No migration shims: this IS the system.
Worker sessions are deny-listed out of every conversation list, so a
profile grinding through a 30-minute kanban task read idle ('3 hr ago')
with no ACTIVE NOW entry the entire run.
- tui_gateway/methods_profiles.py: profiles.list rows gain worker_session
— the newest kanban/tool row (id, source, title, last_active). Workers
heartbeat last_activity_at every <=60s while running (#72016), so the
field stays fresh exactly while work is happening. last_session keeps
its deny-list contract; include_sessions:false omits the field; older
clients ignore it.
- hermes-bots plugin: workerActiveAt() (150s window, one missed heartbeat
of slack) feeds ACTIVE NOW, the row pulse dot ('Working on a task right
now'), and the row age label while a worker runs. Chat semantics are
untouched when no worker is live.
- Tests: 4 new pytest (real SessionDB on temp HERMES_HOME), 2 new node
behavior tests; sabotage-verified.
Session-list visibility of workers (the issue's first half) is left as-is
by design — auto-resume and shared lists must keep excluding workers; the
roster signal was the actionable gap.
`hermes profile rename default <name>` (and the Desktop/dashboard rename
flows) now set a presentation-only `display_name` in profile.yaml instead
of erroring. The canonical id stays "default"; resolution, comparison,
and spawn paths are untouched. Named profiles keep real renames and their
display_name survives the move.
Surfaces: profile list/show/status, /profile (text only — data.profile
stays canonical), dashboard ProfilesPage, TUI-gateway profiles.list, and
Desktop (rail, switcher, Manage page, and the Bot Mode roster via a
displayName fallback so a renamed default shows its name, not "default").
Slimmer redo of the direction in PR #87760 by @yxssxn — thanks; see PR
body for what changed vs that approach.
The BOTS sidebar previewed each profile's most recently active session
(last_session) but clicking the row opened the pinned canonical chat —
two different session identities, so the preview described one
conversation and the click landed in another.
- profiles.list gains an optional preferred_session_ids param
({profile: session_id}): an exact, existence-checked per-profile
lookup that resolves hidden rows and compression lineages to the
live tip (the same resolver session.resume uses) and returns a
preferred_session summary alongside the unchanged last_session.
- The hermes-bots plugin sends its canonical-chat pins with each
roster poll and previews preferred_session ?? last_session.
- openBotCanonicalChat verifies pins through the precise resolver
instead of a paginated, hidden-excluding session.list window that
misjudged real hidden pins as gone; transient lookup failures no
longer clear the pin or mint a replacement chat.
- Grandfathering: a bot with history but no pin adopts the previewed
session on first open instead of minting a new empty chat — the
behavior the design comment already promised.
Closes#88200
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.
- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
(process, home), keyed off the agent's OWN home (not ambient
HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
agent.bot_mode_protocol (default True), stable tier, byte-stable
across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
the bundled plugin gates ALL SOUL protocol writes on it (backfill,
composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere
Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
Follow-up to #86227: _DEFAULT_OFF_TOOLSETS entries (a2a, spotify,
discord, video, x_search, ...) and the region-specific yuanbao are
global opt-ins configured via `hermes tools`/Settings; showing them
unchecked in every per-profile capabilities editor is noise, and
showing a2a at all confuses users who never enabled agent-to-agent
serving (tester report). Enabled ones still show — hiding an ACTIVE
toolset would misrepresent the profile.
Three fixes for what profiles.describe & friends expose (tester
report with screenshot):
1. describe's toolsets used the RAW registry (get_all_toolsets) —
leaking internal platform composites (hermes-discord, hermes-cron,
feishu_drive, discord_admin, desktop_ui, ...) that are gated,
platform-restricted, or deliberately hidden from users — and
reported everything enabled whenever the profile had no pin.
Now: the same filtered universe the `hermes tools` checklist
offers (_get_effective_configurable_toolsets, platform-filtered),
with enablement resolved via _get_platform_tools like the runtime.
2. skills.manage accepts optional `profile`: list/install scoped to
that profile's skills dir via the home override, so editors can
manage a bot's skills (incl. hub installs) from the main window.
search/browse/inspect (the hub catalog) unchanged.
3. New mcp.catalog method: the bundled MCP catalog with per-profile
installed/enabled state + required env keys, so capability UIs can
offer the full menu and route un-setup entries through setup
instead of silently listing dead servers.
E2E: describe now returns only user-facing toolsets with honest
enabled flags; mcp.catalog returns the 5-entry catalog; skills list
scopes to the named profile.
profiles.create inherits the launch profile's provider+model when the
caller doesn't pin one — but the gate was 'config.yaml doesn't exist
yet'. Voice-section mirroring (#85755) runs FIRST and legitimately
creates config.yaml (tts/stt), so inheritance silently skipped for
every non-clone profile since: the bot's editor showed 'Inherit
(launch profile)' while the profile actually had NO model section,
and the first message failed with 'No inference provider configured'
even though the main agent was authenticated and working (Bot Mode
tester report, screenshots).
Gate on what we actually care about: the profile's own raw config
lacking a complete model section (provider+default). Clones bring
their own section and stay untouched; explicit pins unchanged.
E2E: create receipt now model_inherited=true and the fresh profile's
config.yaml carries the launch profile's provider/model.
Three widenings for capabilities UIs (Bot Mode's bot builder):
1. profiles.create share_auth (default false): skip the auth.json
COPY so the new profile reads OAuth/token state through the
existing global-root fallback and refreshes write through to it.
A copy forks token state — the first refresh on either side
invalidates the other for single-use refresh tokens; sharing keeps
ONE live token pool for the main profile and every bot. Static
.env keys still copy (no refresh semantics). Receipt:
mirrored.auth = 'shared'.
2. profiles.describe reports mcp_servers
[{name, enabled, transport}] from the profile's config.
3. profiles.configure accepts enabled_mcp_servers (replace
semantics): toggles via the standard disabled flag; enabling a
server the profile lacks copies its definition from the launch
profile's catalog (names never invented). Launch catalog read
BEFORE the home override flips config resolution.
E2E: describe keys include mcp_servers; create with share_auth ->
mirrored.auth='shared' + no auth.json in the profile dir; configure
applied.mcp_servers=true.
profiles.list's last_session.preview reused list_sessions_rich's
shared preview, which is the session's FIRST user message — right for
session lists (recognition), wrong for a messaging-style roster where
the line under each agent should track the conversation ('Hey, tell
me about yourself!' forever, per user report). Override with the
newest active user/assistant text (same query shape and lock
discipline as SessionDB.latest_message_row_id); best-effort, falls
back to the first-message preview on any failure.
E2E: live profiles.list now shows each bot's latest exchange.
* fix: mirror voice config (stt/tts/voice) into profiles created via profiles.create
Desktop dictation is profile-scoped: /api/audio/transcribe resolves the
stt section inside the TARGET profile's home. Profiles created through
profiles.create got only a model section, so dictation and TTS silently
fell back to defaults (local whisper, often not installed) — 'voice
dictation doesn't work in bot mode but is fine in regular mode'.
Mirror the launch profile's stt/tts/voice sections (key-wise, never
overwriting sections the clone already has) under the same
mirror_credentials flag that gates .env/auth mirroring, and report it
as mirrored.voice in the receipt.
* guard: route voice-config mirror through canonical loaders
read_user_config_raw (write-back round-trip; load_config would merge
DEFAULT_CONFIG and no-op the mirror) + save_config under the target
profile's HERMES_HOME override — same mechanism as _write_profile_model.
Satisfies test_config_read_guard.
ui_meta (#85440) syncs compact roster metadata but is 64KB-capped
because it rides every profiles.list — image avatars stayed per-client.
set_asset writes a validated image (data URL or base64; PNG/JPEG/WebP
by magic bytes, 2MB cap, atomic write) to assets/avatar.<ext> in the
profile dir; get_asset returns it as a data URL on demand; profiles.list
gains a cheap has_avatar flag so rosters know to fetch without probing.
Server-side, so every client machine paints the same profile picture.
* feat: server-side ui_meta on profiles.list/configure
Roster UIs built on profiles.* have per-profile presentation state
(avatar, accent color, display title, pet) with nowhere server-side to
live — client plugin storage paints a different roster on every
machine. profiles.configure now accepts ui_meta (merged key-wise into
profile.yaml's ui_meta block via the existing atomic_yaml_write path,
null deletes a key, 64KB cap since it rides every roster paint) and
profiles.list returns the block per row. Consumers namespace under
their own key. No new files or config; profiles without the block are
unchanged.
* test: stop primary-runtime-restore tests probing live endpoints
_make_agent left the compressor's lazy context-length resolution
unmocked; for reachable base_urls (the nous portal test) the endpoint's
32K answer for the empty test model trips agent_init's 64K floor and
fails the suite on network behavior. Pin get_model_context_length in
the fixture.
profiles.list/create (#85093) let plugins enumerate and create profiles
but not read or modify an existing profile's configuration over ws.
profiles.describe returns the full editor snapshot (description, SOUL.md,
model pin, per-skill enablement via the disabled-list model, per-toolset
enablement via the tools.enabled_toolsets pin); profiles.configure
applies any subset (description via write_profile_meta, soul, model via
_write_profile_model, disabled_skills replace-semantics via
save_disabled_skills, enabled_toolsets replace-semantics with empty-list
clearing the pin) independently and best-effort, reporting per-section
results. Both scoped via the HERMES_HOME override and pool-dispatched.
A profile created through the headless ws door (profiles.create, #85093)
was born with no inference provider: create_profile() seeds a comment-only
.env, never copies auth.json, and a fresh profile has no config.yaml. Its
first message failed with 'No inference provider configured' and the flow
has no interactive setup step to recover with.
New mirror_credentials param (default true): copy the launch profile's
.env (only over the seeded stub — never clobber cloned secrets) and
auth.json (only when absent), both chmod 600, and inherit
model.provider/model.default when the caller gave no explicit pin and no
config was cloned. mirror_credentials:false preserves the old isolated
behavior byte-for-byte. Result gains a mirrored:{env,auth,model_inherited}
receipt. CLI and REST create paths untouched.
Desktop plugins reach the backend exclusively through the generic ws
JSON-RPC door (host.request), but profile enumeration/creation only
existed on the dashboard REST router, which plugins cannot reach — so
anything 'one chat per agent profile'-shaped (bot rosters, profile
pickers, team panes) was impossible to build as a plugin.
- tui_gateway/methods_profiles.py: new @method handlers
* profiles.list — profiles + optional last_session preview per profile
(mirrors session.list's kanban/tool deny-list; best-effort per-profile
state.db probe degrades to null instead of failing the call)
* profiles.create — ws twin of POST /api/profiles (clone_from/clone_all/
no_skills/description), plus optional SOUL.md content and a best-effort
model+provider pin; mirrors the CLI flow (seed skills, safe alias)
Both run on the RPC pool, not the WS reader thread (list_profiles walks
skill trees; create copies bundles).
- SDK: host.openSession(id, { profile, intent }) — open a stored session
the way core surfaces do, soft-swapping to the owning profile's backend
first (ensureGatewayProfile), and host.newChat(profile) — fresh draft in
a named profile (same door as the sidebar's per-profile '+').
- Docs: desktop-plugin-sdk.md gains both surfaces.
First consumer: a Grok Bot-style 'Bots' roster plugin (one persistent
chat per agent profile with a New Agent dialog) built on exactly these
four doors.