`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).
Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.
Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.
Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.
- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
registered when the waiter times out; control.ack/control.error/error
frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
`compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
applies the metadata mirror and emits the same `session.info` a normal
compress does plus the existing `status.update kind=compacted` edge; a
late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
on waiter timeout answer `status: pending` (not 5019) and register the
late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
`status: 'pending'` renders as an info notice, not `error:`; the
`compacted` status edge rehydrates an idle active session's transcript
(mid-turn compaction still defers to the turn settle path).
Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.
Refs #97948
Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
Review follow-up on #93959:
1. Partial-failure window: if the row commits but the transcript copy or
title write fails, the durable-but-empty child defeated the lazy
first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
the renderer fail-latched on a transcript-less session again. The seed
block now compensates: delete just this child so the lazy path can
retry cleanly. Disk-full is exempt — deleting data on a full disk makes
things worse.
2. Silent degradation: the best-effort catch now logs at WARNING with
exc_info instead of DEBUG, so a regression in this user-facing path is
observable without enabling debug logs.
Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.
A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:
- Row created in the PARENT's profile-scoped state.db, stamped with
_branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
from pending_title — the branch lands in the parent's lineage instead
of falling back to a message-preview name.
Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.
Fixes#93959
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.
Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).
NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.
Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.
Fixes#99222
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions
The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.
That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.
- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
coverage for both session kinds
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.
Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
created by older desktop builds that never sent the marker; hidden
non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
per-member room sessions
Fixes#89497
Chat/event-plane quality work for the amended Phase 1 scope of #94484
(maintainer restructure: lean chat/event plane, no control-plane
changes). Three fixes came out of a source-level comparison against
OpenHands, Chainlit, VS Code, Zed, LangGraph, and Goose.
1. Per-turn trace_id + active-turn telemetry: _start_inflight_turn mints
a 12-hex trace_id; _event_frame stamps it on every event frame in the
turn, so a client can correlate the full lifecycle (dispatch -> first
token -> tool calls -> complete) from one identifier — none of the six
surveyed projects has frame-level turn correlation.
session.events.stats now reports active_turns (session_id, trace_id,
elapsed_s, streaming).
2. Transient vs durable events (OpenHands StreamingDeltaEvent pattern):
message.delta / thinking.delta are stamped with seqs (live ordering
holds) but never buffered — one streaming turn emitted hundreds of
delta frames and evicted every durable control event from the
512-slot ring, defeating replay for the exact reconnect window it
exists to cover. The ring now evicts manually and records the highest
DURABLE seq dropped, so truncated means real data loss and delta-only
gaps no longer false-positive.
3. Seq-namespace epoch (Goose stale-cursor recovery): event_replay.EPOCH
(8-hex per boot) is announced in gateway.ready and echoed by
session.events.since; the client drops its seq watermarks when the
epoch changes, so a stale HIGH watermark from a previous gateway
process can no longer suppress replay/gap-detection forever. Legacy
backends without an epoch are unaffected (client keeps watermarks).
Validation: Python 85/85 across replay/entry_ws/keepalive/protocol
(12 replay tests, 3 new); vitest 9/9 shared (2 new epoch tests), 66/66
desktop; full tui_gateway sweep 614/615 (1 known ordering flake, passes
in isolation). Live e2e on this tree: 12-event turn -> seqs contiguous
1..12, 8 durable frames buffered + 4 deltas live-only, single trace_id
on all frames, active_turns elapsed_s matches the real turn duration.
Research provenance: NousResearch/hermes-agent#94484 (comparative-scan
comment); techniques credited to OpenHands (transient split), Goose
(epoch/stale-cursor), per maintainer-restructured plan.
The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.
- server: events_since() now returns bare event objects (the frame's
params), the exact shape the live dispatch path consumes; ring stores
params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
seq-gated afterward — no double dispatch of deltas, no gap-skip from
a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
reset them while clients kept high watermarks (replay forever empty,
truncated=false). New replay_epoch advertised in gateway.ready and
echoed by session.events.since; the client clears watermarks on epoch
change.
- methods_session no longer reaches into event_replay privates
(is_truncated() accessor).
Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.
Server: per-session monotonic seq on every routed event frame, bounded
512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs —
session.events.since (replay newer-than-watermark, reports latest_seq +
truncated so clients detect gaps) and session.events.stats (telemetry).
Client: per-session seq watermarks recorded from live frames; after any
successful reconnect a fire-and-forget fetchReplay() drains missed events
through the normal dispatch path (recordSeq ignores non-increasing seqs,
so stale replay can never regress a watermark); focus-triggered reconnect
nudge in use-gateway-boot for the Electron unfocused case where macOS wake
skips visibilitychange.
Replay failures are swallowed by design: lossless resume is an upgrade
over the previous lossy reconnect, never a new failure mode.
Review batch (3 reviewers) on the final diff surfaced:
- H1: title-based donor matching could adopt AND non-recoverably retire
an UNRELATED default-store conversation (bot titles collide by design;
get_session_by_title has no archived filter/ordering). Donor probe is
now exact-id only — the stranded repro always has the id.
- H2: re-adoption after a partial run could retire a donor that had
accumulated NEWER messages than the profile copy (skip-based
idempotency never merges). New divergence guard compares message
counts and refuses retirement when the donor is ahead (still adopts).
- M1: donor_retired reported True even when every retirement step
failed under suppress. Now per-segment tracked + warn-logged;
True only when all applied.
- M3: adopted=False (e.g. import validation limits) was silent — now
warn-logged with import errors.
- M4: archived donors are never re-adopted (no cross-profile cloning).
- Dead 'from pathlib import Path' dropped; contextlib no longer needed.
5 new red-first-verified regressions (title-collision immunity,
archived-donor immunity, non-vacuous owns_db gating with a real donor
seeded, divergent-donor retirement refusal, donor_retired truthfulness).
tests/tui_gateway: 578 passed. ruff clean.
Pre-#93296, the desktop routed session RPCs by the focused tile, so a
profile bot's turns executed on the default backend and its canonical
session accumulated in the DEFAULT profile's state.db. Post-fix, the
profile backend correctly receives the resume — but its store has never
seen the session, so the same chat 4001s forever (unreachable instead
of misrouted). Live repro: Teknium's Developer bot, session c93770.
- hermes_state_portability: SessionDB.adopt_session_lineage_from() —
composes the existing export_session_lineage()/import_sessions()
primitives; donor rows are archived (never deleted) with
end_reason=adopted_by_profile, which is deliberately NOT in
RECOVERABLE_END_REASONS so canonical-lookup resurrection cannot undo
an adoption. Idempotent (already-present ids skip).
- tui_gateway/methods_session: profile-scoped session.resume falls back
to adoption from the default store right before the 4007; ids unknown
to BOTH stores still 4007 exactly as before, and launch-profile
resumes never consult the fallback.
- tests: 10 new (7 unit on the primitive incl. compression-lineage
unit adoption + non-resurrectable archive; 3 handler-level through
server.handle_request incl. the live repro shape); db-ownership
leak test taught that the shared launch handle probe is by design.
Follow-up to #93296/#93311; part of #93091.
Live WS E2E after the #93361 merge (real web_server + tui_gateway, isolated
HERMES_HOME, 2s grace): drop socket -> re-resume stored id on a new socket
still produced a ws_orphan_reap reclaim. The lazy/unpersisted resume branch
(no state.db row yet -- every fresh Bot Chat) returned the sentinel-parked
live record without rebinding its transport or cancelling the armed reap
Timer, so the storm survived for exactly the Bot Mode sessions the cluster
targeted. The unit-covered paths (_live_session_payload, _reuse_live_response,
_claim_or_reuse_live) were all correct; this branch bypassed them.
Regression test drives the real session.resume RPC against a sentinel-parked
unpersisted record (sabotage-verified: fails without the fix). After the fix
the full live E2E passes 10/10 scenarios including a 4-cycle drop/resume storm
loop with zero reclaim broadcasts.
Storm killer for the reap->broadcast->auto-re-resume feedback loop:
- New _pending_ws_reaps registry (sid -> Timer): _schedule_ws_orphan_reap
registers, _reap pops, and _cancel_ws_orphan_reap(sid) is called from
every resume/reuse/rebind path — the session.resume fast-path reuse
(methods_session.py), _claim_or_reuse_live winners, and the
_live_session_payload live-transport rebind.
- When a resume mints a fresh runtime for stored session id S, any prior
runtime for S still parked on the detached-WS sentinel is claimed under
the resume lock, its reap Timer cancelled, and the record finalized
quietly with end_reason superseded_by_resume — NOT in
_RECLAIM_END_REASONS, so no session.reclaimed broadcast fires and the
client's auto-re-resume can't storm.
- superseded_by_resume added to _RECOVERABLE_END_REASONS in
hermes_state_common.py so canonical Bot Chat resurrection still applies.
Unit tests: resume cancels the reap timer, superseded runtimes finalize
without a reclaimed broadcast, and the normal orphan reap still fires
when nobody re-resumes.
After the existing reconnect grace, route a still-detached running session through the same interrupt mechanism as session.interrupt. Preserve delegation deferral, sidecar teardown, partial history, and single-owner reap semantics.
Verified on upstream main: RED 4 failed/2 passed without production changes; GREEN 603 related gateway/compute-host tests. Ruff and py_compile passed. Momus pass 2: APPROVE.
- Clear the accidental end stamp on resurrection (at the lineage tip):
a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
archive auto-resurrect on the next lookup — the user could never retire
the canonical chat. Test pins the resurrect -> deliberate-archive ->
stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
compressed lineage carries end_reason='compression', so tip-stamped
accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
dm resolution) filtered archived rows out via list_sessions_rich and
still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
session.create intentionally persists no state.db row until the first
prompt, but session.resume only looked in the database — so resuming a
live lazy session by its stored key or pending title hard-404'd. Bot
Mode hits this on every fresh non-default bot: the canonical Bot Chat
is created lazily on the profile, the open/send resumes it, and the
user gets 'session not found' on their first message to that bot.
session.resume now falls back to the in-memory session registry,
matching by stored key or pending title scoped to the SAME profile
home. Cross-profile lookups still fail closed; unknown ids still 404.
The eager session.resume path called _transfer_db_to_agent(agent, db)
unconditionally. With no non-launch profile selected, db resolves to the
SHARED launch handle (_get_db()), so the transfer succeeded on identity
alone — the agent IS holding that handle — and session.close() then
closed the process-wide database under every unrelated session:
subsequent writes failed with "'NoneType' object has no attribute
'execute'" and the Desktop could not open chats until restart (#91610).
This directly violated _transfer_db_to_agent's own contract ("Never
called for the shared launch handle", introduced with the ownership
lifecycle in #81071).
Gate the transfer on owns_db (dedicated handles only), and add defense
in depth: _transfer_db_to_agent now refuses db is _get_db() even when a
caller invokes it incorrectly.
A bot's forever-chat now has exactly one identity: the session titled
"Bot Chat" on that bot's profile. Core UNIQUE(title) makes (profile,
'Bot Chat') an exact registry, and every open consults it directly via
session.list {title, include_hidden}. The stored-id pin
(ui_meta['hermes-bots'].chat) and its entire verification apparatus —
preferred_session_ids resolution, drifted-pin keep branches, last_session
grandfathering, dead-pin recovery re-anchoring, newerVisibleBotChat — are
removed, not deprecated. Legacy ui_meta.chat keys are ignored and dropped
from merges on sight.
Every lost-canonical-chat incident (#88146, #88200, #90524, #90705, and
five hardening waves) traced to that pointer dangling or being stolen,
then later guards welding the wrong session in. A name cannot dangle:
corrupt pins self-heal on first click because the pointer is simply never
read.
Gateway: profiles.list now reports canonical_session per profile row
(registry row resolved server-side by title — hidden rows resolve,
deny-listed sources and archived rows do not, compression lineages
resolve to the live tip), replacing the preferred_session_ids request
contract. The roster preview, activity signals, and the /new→/compact
guard all read canonical_session, so preview identity and click identity
are the same row by construction.
No migration shims: this IS the system.
The #90732 adoption scan used session.list's 200-row recency window. A busy
bot profile (group-chat traffic, routines, or accumulated fork spam) pushes
an older forever-chat past row 200, the scan misses it, and the mint path
re-enters the unique-title-conflict fork loop — same pathology, higher
trigger threshold.
Profile → Named Session is an exact registry (UNIQUE title index), so
consult it exactly:
- session.list gains a `title` param: indexed WHERE title = ? lookup,
window-free, hidden rows resolve, archived/deny-listed do not,
compression lineages resolve to the live tip (resolved_id), mirroring
profiles.list's preferred_session resolver.
- findExistingCanonicalChat sends title: 'Bot Chat'. Older gateways ignore
the unknown param and return the windowed listing — the local scan stays
as the compatibility rung.
- Adoption opens the lineage tip (resolved_id) while pinning the durable id,
same split as the preferred_session path.
Bot Mode's group chats spawned one per-member session per room, and those
"Group: ..." rows (plus canonical Bot Chats when the old eye-toggle pref was
off) flooded the global Sessions sidebar — a 6-bot room dumped six identical
rows into recents (reported with screenshot, Aug 17).
Plugin (apps/desktop/src/plugins/hermes-bots/plugin.js):
- session.create now passes hidden:true UNCONDITIONALLY for both canonical
Bot Chats and group-room member sessions; the $hideBotChats pref, its eye
toggle, and its storage hydrate are removed (Bot Mode sessions are plumbing
or plugin-owned forever-chats, never scratch conversations).
- hideOwnedBotSessions(): idempotent reconciliation sweep over every owned
session id (bot meta canonical chats + each room's member sessions) via
session.set_hidden, run on plugin load and on each gateway reconnect, so
rows born visible under the old pref get cleaned up.
- The Bots session browser and canonical-chat recovery scan pass
include_hidden:true so they still see the rows they own.
Gateway (tui_gateway/methods_session.py):
- session.list honors an include_hidden param (default off — the resume
picker and all global callers keep dropping hidden rows).
- session.set_hidden gains a durable fallback: when no LIVE runtime session
matches, resolve the stored session id in the target profile's state.db
(via resolve_session_id) and flip the flag there. The sweep holds stored
ids for chats that aren't live; the old live-only lookup 4001'd them.
Validated E2E with real imports against a temp HERMES_HOME: born-hidden row
(hidden=1), profile-scoped session.list default vs include_hidden (0 vs 1),
and stored-id sweep on a non-live legacy row (hidden=1). Plugin suite
167/167; new RPC regression tests in tests/tui_gateway/test_session_hidden_rpc.py.
session.workspace.move refused a running live session with 4009
(session busy), but the desktop's Move-to-project flow calls exactly
this RPC — so the UI updated its local grouping while state.db kept the
old cwd and the agent's tools kept running in the old workspace. Two
sources of truth disagreed (#86626).
An explicit move now wins: the stored row and the live session re-anchor
together. In-flight tool calls keep the cwd they were launched with; the
next tool call uses the new workspace.
Follow-up to the #62799 salvage: Desktop sends both defer_history and
omit_messages on a cold resume. Make explicit in the deferred branch that
defer_history supersedes omit_messages — the single history read happens in
the background hydration worker and the synchronous omit_messages read on
the cold-resume default path is skipped entirely, so the transcript is
never loaded twice for one resume.
The Desktop's content-based truncation-target resolution (and reactions)
address persisted turns by row_id, but session.history loaded the
transcript without include_row_ids=True, so _history_to_messages had no
stamp to forward and the projection silently stripped the one durable
address clients can use. Discovered live-testing the #87294 client flow:
resolveDurableRowId saw 0 stamped rows and degraded every edit to a
plain resubmit.