The three /status renderers (hermes_cli/cli_session_mixin.py::_show_session_status,
gateway/slash_commands_status.py::_handle_status_command, tui_gateway/methods_session.py
session.status) each hand-built Session ID / Path / Title / Model (provider) / Created /
Last Activity / Tokens / Agent Running with their own getattr(agent, "model") fallback
chain, their own updated_at/last_updated_at/last_activity_at scan and their own timestamp
format. A fix to one (a new last-activity column, a placeholder change) silently missed the
other two.
hermes_cli/status_report.py::build_status_fields now derives the common facts once and
returns them as structured, display-ready data; status_lines() renders the English
"Label: value" form for the CLI and TUI. The gateway keeps translating through its
existing t("gateway.status.*") catalog keys (no locale change); the CLI keeps reasoning /
approvals / context, the gateway keeps free-tier / context / queue depth / Matrix scope,
the TUI keeps its Project line. tui_gateway/methods_session.py::_status_dt and the
CLI's inline updated_at loop are gone; cli_session_mixin._timestamp_or stays for its
remaining history-timestamp caller.
Behavior change: none intended for populated sessions. Unified edge cases: a
SessionDB row with an unparseable started_at now falls back to now() on the TUI as it
already did on the CLI, and the TUI's fallback on a bad updated_at is the created stamp on
both surfaces.
Test: tests/hermes_cli/test_status_report_contract.py drives the three real renderers with
one session (distinctive model, provider, title, stamps, token count) and asserts each
output carries every common value. Sabotage-verified: builder dropping tokens -> red;
TUI hand-formatting the model line -> red; restored -> green.
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors
The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status
answers has_guest / enabled / carries_inference / notice_pending with zero network, and
free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check
reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and
billing.state carry free_tier (billing answers the free tier locally instead of a portal call that
can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or
locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector
transfer and returns its code and consent URL; the poller waits for the transfer before the token
grant, persists the account, runs settle_after_upgrade, and the poll response gains reason,
account_email and model.
* feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog
The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch
intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding
overlay opens on a ready screen when the free tier carries inference, else a one-time strip above
the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model /
Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the
tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that
drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens;
Done settles billing, model options, providers and re-homes a session still on nous/welcome. The
picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes.
* fix(desktop): free_tier.status starts the free tier's background setup when no identity exists
A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free
tier was never set up on the desktop: no connectors, no notice strip. The first status read now
starts the same one-attempt background setup; the call itself never waits.
* fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected
The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in).
The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free
tier, instead of Nous Portal · Connected.
* fix(desktop): Settings > Providers never files the free tier under Connected
* fix(desktop): the intro's shape is keyed on the route, not on the identity
free_tier.status reports available (an identity exists and the tier is on); whether inference
runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready
screen shows when that route is the free tier; the composer strip when the user's own provider
carries inference. An own-key install used to get the ready screen.
* docs(desktop): say what the free-tier chip is keyed on
* fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds
* fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro
Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the
transfer wait, after the token grant, and once more under the session lock together with the
save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an
attempt generation that every continuation checks after each await, so a poll from a closed
attempt cannot publish over the one on screen (and its backend session is cancelled). The ready
screen comes down only after the backend recorded the acknowledgement. A composer still mounted
takes over the notice claim when its owner unmounts. One thin test per hole.
Two independent reviews of the seeded-create change found three more
places where the newly durable hidden row, or the new create-time copy,
was not handled by the same rule as the rest of the path:
- Message search (dashboard search and the session_search tool) had no
display_kind filter, so a hidden opening row matched a query the
person never saw. The shared search predicate now skips hidden rows.
- _seed_row left the fresh session row behind when the transcript copy
failed after the row was committed. The first prompt's retry copies
the whole seed, so a kept partial copy would be duplicated. The row
is now deleted when the copy did not complete, the compensation
_persist_branch applies to branch children; the first prompt then
starts clean.
- _live_session_payload (a resume that reuses a live session) reported
message_count as the raw history length while its messages array was
filtered. It now follows _resume_response: the stored size when
messages are omitted, else the wire count.
Tests: the two seeded-create tests now drive the first-submit path
through _persist_session_row_for_submit, the function prompt.submit
calls, and assert search and the reuse-live count; a third test pins
the rollback (no row after a failed copy, one copy after the retry).
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once
session.create accepts opening messages. Three defects sat in that path:
- A seeded session without a parent was never persisted at create, so a
restart before the first prompt lost it and session.resume answered
4007. Only branch children (#93959) were persisted up front. The
same rationale applies to any seeded create: seeded content is
intent, not an abandoned draft. Parentless seeds now persist their
row, transcript and client title at create; empty drafts stay lazy.
- _coerce_seed_history dropped display_kind, so a seeded row tagged
"hidden" (model-facing scaffolding) rendered as a user bubble. The
coercion keeps "hidden" and only "hidden"; every other kind is
stamped by the gateway at turn time and is not accepted from the wire.
- A branch child's seed was written twice: _seed_branch_row copied it at
create but never marked it persisted, so the first prompt's
_persist_branch_seed appended the copy again. The create path now
sets _branch_seed_persisted, and the gate is a create-time `seeded`
stamp instead of parent_session_id, so a resumed session (whose
history comes from the DB) can never re-append its transcript.
Two invariant tests, both red on main: a parentless seed survives a
gateway restart with the hidden row kept out of the wire transcript and
not re-written by the first-submit path; a branch child's seed is stored
exactly once. The reasoning-fields fixture stamps `seeded`, the flag
session.create sets.
* fix(tui-gateway): a hidden seed row stays out of the list preview and the create count
Live-testing the seeded create on every surface showed two places where
the newly durable hidden row (display_kind="hidden") still surfaced:
- session.list built a session's preview from its first user row with no
display_kind filter, so a hidden opening row (model-facing scaffolding
the gateway never paints) became the sidebar preview. The preview
predicate now skips hidden rows, in every listing query that shares it.
- session.create reported message_count as the raw seed length while its
messages array already filtered the hidden row (2 vs 1). It now counts
what is on the wire, the same rule session.resume applies.
Both are covered by the existing seeded-create test: the create count
equals the wire transcript, and the preview of a session whose first
user row is hidden is its first visible user row.
* fix(tui-gateway): a live unpersisted resume counts the wire transcript
session.resume on a live session that has no row yet reported message_count as
the raw history length while its messages array was already filtered, the same
mismatch the previous commit fixed on session.create. Count the wire, as the
cold, deferred and reuse-live resume paths already do.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
A session held exactly one transport, and prompt.submit, session.resume, session.activate, and the queued-prompt drain all rebound that slot. A second client therefore took the stream away from the first: the earlier client stopped receiving the turn it was already rendering, and either client disconnecting parked the whole session on the drop sentinel.
FanoutTransport goes in the same slot and satisfies the same Transport protocol, so write_json and every other reader of the slot are unchanged. It delivers each frame to a snapshot of its peers, concurrently when more than one peer is attached and the caller is not on an event loop, and prunes any peer that returns False or raises. A dead client is dropped; a slow one costs the emitter at most one write timeout per frame rather than one per peer. Request/response RPCs are unaffected: they still answer on the request's context-bound transport, so a client only ever sees replies to its own calls.
The rebind sites become attach sites through _attach_session_transport, whose ladder keeps the single-client shape identical. The same object already in the slot is a no-op; an empty, stdio, or parked slot is taken outright; only the arrival of a second live client wraps both. The queued-prompt drain is included because it pinned the drained turn to the queuer and silenced everyone else. A non-peer newcomer such as stdio or the drop sentinel never displaces a live client, so an activate dispatched without a bound websocket cannot silence the socket that owns the session.
Disconnect detaches first. A session that retains another client keeps streaming and is neither parked nor reaped, and only the clientless ones follow the existing close_on_disconnect and park-sentinel path, so a single-client disconnect, the orphan reaper, and its grace window behave as before. _ws_session_is_orphaned is unchanged: it still asks whether the drop sentinel is in the slot, and a fan-out is never the sentinel, so a session that still has a peer is never reported as orphaned.
Attach performs no entitlement check: any authenticated peer may mirror any session.
The fan-out architecture follows the approach in #40822 by @OmarB97.
What the slot's later history forces. _close_sessions_for_transport drops the #83716 rebind-to-the-most-recent-surviving-viewer, which fan-out membership subsumes — a pop-out window is a peer, so a session that still shows in one is never returned as clientless — and keeps the #77129 revalidation before parking, now expressed as a liveness check under _session_transport_lock so it is race-free against attach and detach. _transport_is_live_peer defers its last answer to _transport_is_dead: a socket that already latched _closed is a departed client, and admitting it would keep a session out of both the park and the reap. _transport_is_dead also learns the fan-out: a FanoutTransport with no live peer is dead, so a session whose peers were all pruned by failed writes cannot outlive the TTL and LRU reapers. The upstream test that pinned the #83716 rebind, test_close_transport_rebinds_session_to_remaining_viewer, is re-expressed in fan-out terms: both windows attached, the pop-out closes, the session stays with the main window unparked and still receiving frames.
The pin from the previous commit lives in _SESSION_STATE but nothing cleared it, so a CLI
/new, /resume or /branch (same AIAgent, reset_session_state + _invalidate_system_prompt)
replayed the previous session's git snapshot into the new session's prompt. Clear it in
reset_session_state next to the other session anchors; one invariant test (red without it).
The TUI/Desktop session.context_breakdown RPC ran the prompt builder on the RPC thread
with no session cwd bound, so it re-probed against the backend's cwd and overwrote the
session's pin — one /context between compactions restored the divergence this fix removes.
Bind the session context around the build like the live rebuild in server.py does.
Also: trim _coding_parts' docstring to the WHY, drop the isinstance/len guard on a value only
this function writes, and remove the tests' assertions on the private pin shape.
_resume_live_unpersisted hand-rolled the same transport + viewers + reap-cancel
sequence _rebind_live_transport now owns. Route it through the helper; the
stdio case (no current transport) keeps cancelling the reap as before.
The salvaged fix correctly put the activate guard + transport rebind under the
process-wide resume lock, but it dragged _live_session_payload in with them. The
Desktop passes omit_messages=true (cheap), the Ink TUI does not: every TUI session
switch then read the full persisted history from the profile DB while holding the
lock that serializes every resume, disconnect and reap Timer.
Extract _rebind_live_transport from _live_session_payload; activate does guard +
rebind under the lock (the part that must be atomic with grace expiry) and builds
the payload after releasing it.
The 4007/4009 guard the salvaged fix added to prompt.submit, session.activate,
_resume_live_unpersisted and _resume_reuse_live_locked was the same six lines
four times. One helper in session_lifecycle.py (next to _cancel_ws_orphan_reap,
whose contract it mirrors) so the next reattach path cannot drift from the others.
Behaviour unchanged; the 19 race tests cover all four call sites.
Re-scope #98106 onto the activity-based orphan policy from #100504. Keep timer ownership across callbacks and continuations, serialize reconnect paths against interrupt claims, and avoid recursive eager-resume locking. Leave cleanup polling armed when concurrent cold reuse is rejected.
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).
Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.
Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.
Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.
- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
registered when the waiter times out; control.ack/control.error/error
frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
`compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
applies the metadata mirror and emits the same `session.info` a normal
compress does plus the existing `status.update kind=compacted` edge; a
late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
on waiter timeout answer `status: pending` (not 5019) and register the
late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
`status: 'pending'` renders as an info notice, not `error:`; the
`compacted` status edge rehydrates an idle active session's transcript
(mid-turn compaction still defers to the turn settle path).
Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.
Refs #97948
Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
Review follow-up on #93959:
1. Partial-failure window: if the row commits but the transcript copy or
title write fails, the durable-but-empty child defeated the lazy
first-prompt fallback (_ensure_session_db_row is INSERT OR IGNORE), so
the renderer fail-latched on a transcript-less session again. The seed
block now compensates: delete just this child so the lazy path can
retry cleanly. Disk-full is exempt — deleting data on a full disk makes
things worse.
2. Silent degradation: the best-effort catch now logs at WARNING with
exc_info instead of DEBUG, so a regression in this user-facing path is
observable without enabling debug logs.
Tests: compensation deletes the half-written row and preserves
pending_title; disk-full keeps the row and surfaces the WARNING.
Desktop branch creation hung on an infinite spinner and lost the branch
on restart. Root cause: the renderer branches via session.create with
parent_session_id + a seeded transcript, but session.create defers the
DB row to the first prompt (the draft-hygiene contract). The renderer's
post-create resume then re-fetches the fresh child through REST and
defer_history hydration — both read the DB. An unpersisted child 404s
and hydrates empty, the client fail-latch (sessionShouldHaveTranscript +
empty messages) refuses to bind a "transcript-less" session, and the
user sees a spinner forever; on restart the rowless child vanishes and
the optimistic "Draft: Branch N" entry disappears with it.
A seeded branch is explicit user intent, not an abandoned draft.
session.create now persists the child immediately when both
parent_session_id AND seeded history are present:
- Row created in the PARENT's profile-scoped state.db, stamped with
_branched_from + parent_session_id (same shape as TUI /branch).
- Seeded transcript copied via append_messages_batch so REST prefetch
and defer_history hydration find it on the first read.
- Title assigned from get_next_title_in_lineage(parent) and cleared
from pending_title — the branch lands in the parent's lineage instead
of falling back to a message-preview name.
Persistence is best-effort: a broken DB logs and lets create succeed,
leaving the lazy first-prompt path as fallback. Plain drafts keep the
lazy-row contract unchanged.
Fixes#93959
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.
Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).
NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.
Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.
Fixes#99222
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions
The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.
That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.
- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
coverage for both session kinds
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.
Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
created by older desktop builds that never sent the marker; hidden
non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
per-member room sessions
Fixes#89497