Two sites dropped profile ownership on the tile-transcript refresh path:
- tui_gateway/server.py _sessions_sig statted only the launch home's
state.db, so a turn landing in a served sibling profile's store never
produced sessions.changed. The watcher now also probes every profile
home _profile_home() has resolved for this backend (empty set on
single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
transcript unscoped; it now passes the tile's ownerRoute scope, the
same way reconcileActiveTranscript already does for the main pane, and
keys the change signature by owner.
Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).
Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.
bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.
On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.
Closes#100523
Supersedes #100544, #100542
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).
Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.
Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.
Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.
- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
registered when the waiter times out; control.ack/control.error/error
frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
`compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
applies the metadata mirror and emits the same `session.info` a normal
compress does plus the existing `status.update kind=compacted` edge; a
late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
on waiter timeout answer `status: pending` (not 5019) and register the
late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
`status: 'pending'` renders as an info notice, not `error:`; the
`compacted` status edge rehydrates an idle active session's transcript
(mid-turn compaction still defers to the turn settle path).
Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.
Refs #97948
Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
Review point from keeltrace, and it is a real gap: secret redaction and
prompt omission are different contracts, and only the first one is
pattern-shaped.
redact_sensitive_text removes credentials. A provider 4xx that quotes the
request back carries the user's own prose - a paragraph about a person, a
file pulled in by an @ reference - which matches no credential pattern and
so passed through untouched into cause=. The record's stated contract is
that prompt content is not logged, and the previous commit only enforced
the half of it that a regex can see. The existing prompt test could not
catch this: its provider error does not echo the prompt, so it proves the
prompt is not logged directly, not that it cannot arrive by being quoted.
_strip_prompt_echo closes the quoted path directly. Anything the message
shares with the submitted prompt for 24 characters or more becomes
<prompt>. Shingle-set matching rather than a diff, so cost is linear in
both strings on a path that runs for every failed turn and can face an
@-expanded prompt of arbitrary size; the JSON-escaped form of the prompt is
shingled too, because a provider handing back its own request body often
hands it back escaped. The prompt is captured after @-expansion on purpose:
an injected file's contents are exactly the material an echo would carry,
and they are not in the submitted text.
Ordering is load-bearing. The strip runs after the whitespace collapse, so
a re-wrapped quote still matches, and before the length cap, so a quote
cannot survive by being cut mid-run.
What this does not claim: verbatim echo is what it stops. A paraphrase, a
summary, or a re-encoding would survive it. The alternative keeltrace
raised - log only structured provider metadata and drop the message body -
is airtight but costs the diagnosis this PR exists to enable, since the
reporter needed to tell a 402 from a crashed finalizer. Happy to switch if
maintainers prefer the stricter contract.
Tests: the non-secret sentinel keeltrace asked for (a benign phrase present
only in the prompt, echoed by the provider error, asserted absent from the
record), plus guards that a message sharing nothing with the prompt is
untouched, that an overlap below the window is not treated as an echo, that
a prompt shorter than the window cannot blank the message, that a
JSON-escaped echo is stripped, that the strip precedes the length cap, and
that whitespace shape does not hide an echo. The three that cover the new
path fail with the strip removed; the guards pass either way.
Fixes#89117
The whole of #89117 is two log lines:
tui_turn finished: ui_session=0dfcee58 status=error error_retained=True duration=0.9s
A provider 4xx, a budget wall, a billing block and a crashed finalizer all
produce exactly those characters, so an intermittent failure cannot be
triaged from the one record that is guaranteed to exist.
The bookend came from #86865, which added it to trace compression
rotations across #86647 -- identities and a coarse status were the job, and
content was deliberately excluded. What that leaves is a returned-error
path (provider 4xx, budget, billing) which writes no other log line at all.
The exception path at least prints `[gateway-turn] <Type>: <msg>` to
stderr, so the failures that go unlogged are exactly the sub-second ones
this issue is about.
Both failure paths now stash a one-line cause, and the bookend appends it.
The record keeps its shape when nothing failed: a successful turn gains no
new fields.
The cause is redacted with `redact_sensitive_text(force=True)` and capped at
240 characters with a visible ellipsis, because a 4xx body routinely quotes
the request that produced it -- adding the cause without redacting it would
write an Authorization header the user never chose to log. Redaction fails
closed: if the redactor cannot run, the fragment reads `<unredactable>`
rather than the raw message. Whitespace is collapsed so a multi-line
provider body cannot split the record, which is the only property that
makes it greppable for a bug like this one.
12 regression tests. Four mutations proven: disabling the helper fails 9,
dropping redaction fails 2, dropping truncation fails 1, wiring only the
exception path fails 4.
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
A scoped projects.tree / projects.project_sessions response is built from
ONE profile's state.db, so the request scope is authoritative even for
legacy rows whose persisted profile_name is NULL. Without the stamp those
rows reach the renderer ownerless, and every owner lookup off them — the
branch path included — falls back to whichever backend is active.
Co-authored-by: evan-bradford <evan-bradford@users.noreply.github.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Seven test files asserted sub-2s wall-clock bounds (elapsed < 0.5/1.0s,
stop(timeout=1.0), event waits of 0.5-2s). Under CI load these fired on
healthy code: main run 33455779041 alone flaked 6 of them in one pass
(observed 1.01s vs 0.5, 1.20s vs 1.0, 3.61s vs 3.0, 1.55s vs 1.0,
stop(1.0) returning False, lease TTL 0.1s expiring before the authority
change was observed).
Per the AGENTS.md flake policy (waits >= 2s), bounds are raised to 5s+
while keeping their teeth: every hang path they guard blocks for 10s+
(release.wait holds), so the loosened bounds still distinguish bounded
from unbounded behavior. The authority-loss test gets a 30s lease TTL so
lease expiry can no longer preempt the authority-change assertion.
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
config.set matches an explicit key list and answers 4002 for anything
else, so a renderer mirroring an unlisted key wrote nothing at all. The
reactions toggle shipped that way: every write was rejected into a
swallowed .catch(), and react_to_message stayed dark no matter what the
user picked.
Adds the display booleans as a recognized group so the toggle reaches
the config of whichever gateway the app is actually talking to, which is
the only place a check_fn can read it.
Pin that session.interrupt retires the original and rotated marker keys
on ACK, and that a Stop arriving before the disk write cannot leave a
resume-able marker behind.
Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
Desktop opens the registry id, then session.resume walked the legacy
unmarked-child fallback, so Open Chat still landed in a side chat after
the title lookup was already strict. Recoverable-archive resurrection
uses the same helper.
Co-authored-by: Gille <4317663+helix4u@users.noreply.github.com>
A shared dashboard's launch HERMES_HOME is not the selected profile. model.options now runs under @_profile_scoped, and global-remote REST keeps ?profile= even for the primary label.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Follow-ups on top of the #94595 cherry-pick, implementing the maintainer
review's two blockers:
Blocker 1 (turn-admission chokepoint): _run_prompt_submit itself now runs
the ownership admission, so synthesized turns that never pass through the
prompt.submit RPC handler (crash auto-continue from cold session.resume,
wake-ups) are fenced too. Auto-continue additionally checks ownership
BEFORE emitting message.start and leaves the marker in place, closing the
#94778 shape where backend B resumed a session backend A was actively
running and auto-continued A's fresh interrupted-turn marker into a
duplicate concurrent turn.
Blocker 2 (fail-closed registry semantics): try_acquire_active_session no
longer converts an unreadable/corrupt registry into an untracked go-ahead.
Ownership uncertainty is a distinct typed refusal —
SESSION_COORDINATION_UNAVAILABLE — because "could not prove ownership" must
never be collapsed into "no owner exists". The TUI gateway claim helper
fails closed on claim exceptions for every surface, not just desktop.
Also: empty session ids short-circuit to a no-op lease (nothing to fence,
and the strict registry schema rejects empty ids), and the existing
fail-open tests were updated to assert the new fail-closed contract.
Pin gateway re-tagging of idle/preflight lifecycle lines as compacting,
and assert the TUI keeps that status until compacted rather than
restoring the busy bar after 4s.
config.get and config.set ignored the focused profile on a shared
app-global backend, so reads and persistent writes used the launch
config.yaml. Bind the existing @_profile_scoped decorator and write
_save_cfg through the request home override.
Fixes#95760
Follow-ups on the salvaged #97744 runner:
- tui_gateway/hosted_room_driver.py: HostedRoomRuntime.cancel() treated its
initial status read as truth, so a task transitioning queued->running (or
settling) between the read and the state call surfaced a transient
'running work requires acknowledged two-phase cancellation' /
StaleTaskError to the caller and failed groups.disband. Deterministic
repro on the PR head: test_client_event_id_cannot_squat_disband_receipt
failed 5/5 locally. cancel() now re-reads and re-routes on every
race-shaped failure (bounded retries), returns already-cancelled tasks
idempotently, and rejects truly terminal states honestly.
- methods_groups conflict resolution keeps both method sets: the replication
surface from #99047 (groups.replicate/replica_state/promote/demote) and
the runner surface from this layer (groups.stop/retry/approve).
- test_groups_replication_methods.py updated to the runner's stricter
create contract (2-6 profile-backed members, live worker service).
Every participant gateway can now keep a durable copy of a hosted room's
ordered log and continue the room when its authority host is gone:
- gateway/hosted_room_replicas.py: replica store in root state.db.
ingest_page() persists authority-stamped groups.log pages idempotently,
refusing sequence gaps and authority-epoch regressions. promote_replica()
continues the room locally at epoch+1 with a lineage-proving
authority.claimed event; the stale owner is fenced everywhere the claim
replicates. demote_room() lets a returning stale authority fence itself
(authority.lost) upon observing a newer epoch, killing split-brain writes.
- tui_gateway/methods_groups.py: groups.replicate / groups.replica_state /
groups.promote / groups.demote RPC surface. Promotion requires
confirm=true — storage decides HOW takeover is atomic and provable, the
caller (user action now, lease/quorum driver later) decides WHEN it is
safe, matching the boundary blessed on #97681.
Validation: 20 new tests incl. a full failover round-trip (A hosts, B
replicates incrementally, A dies, B promotes with complete history, A
returns demoted and fenced); 69 total across the hosted-rooms area; E2E
with two real gateway stores and real install identities.
Resume was attaching an unused store as {todos: [], revision: 0} and the desktop rejected tool.start updates that have no revision. A merge:true start after reconnect never patched the list until complete.
Skip unused empty snapshots. Apply unversioned updates without moving the watermark so a later todo.updated can still win.
The gateway's session-backed MCP OAuth flow (mcp.servers.oauth.start) binds
its browser-callback listener on the BACKEND machine's 127.0.0.1. When the
Desktop app connects to a remote backend (SSH/Tailscale), the user's browser
resolves that loopback to the user's machine, the redirect dies, and every
OAuth catalog server (ClickUp, Hospitable, ...) fails in-app with no working
path — the exact topology from the 'MCP Recurring erros' support thread.
Fix mirrors the Desktop's native gateway login (native-oauth-login.ts):
- gateway: mcp.servers.oauth.start accepts client_redirect_uri (loopback-only,
RFC 8252-style validation); when supplied no gateway listener is bound and
the OAuth redirect_uri pins to the client's listener.
- gateway: new mcp.servers.oauth.callback RPC relays the client-captured
code/state into the flow; state verification stays in
DashboardOAuthFlow.deliver_callback (constant-time compare, replay-safe).
- desktop: mcp-oauth-callback-ipc.ts hosts a one-shot 127.0.0.1 listener in
the main process (hermes:mcp-oauth:listen/wait/cancel via preload bridge).
- desktop: hermes-bots mcp-setup.tsx prefers the client listener for local
AND remote backends, falling back to the legacy gateway-listener flow on
older gateways (feature-detect via start rejection).
- docs: remote-host MCP OAuth section documents the automatic Desktop path.
Validation: 19 new gateway tests (validator allowlist, listener skip, relay
accept/reject/replay) — sabotage-verified; 5 new desktop tests against a real
ephemeral listener; E2E through the real session registry + flow bridge with
a stubbed provider probe; tsc electron+renderer builds clean.
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions
The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:
1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
the coding-brief sentence in agent/coding_context.py now names
todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
held 'tour' — post-hook ownership would have double-emitted for
gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
contract under the legacy defer:[] override, or assert on granted
tool names instead of visible schemas.
Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.
649 tests green locally across all touched files, both orderings.
Flip the salvaged --start-now behavior (PR #97958) into the unconditional
default: /loop's first iteration is due the moment the loop is set, then
recurs on the normal cadence. The flag is dropped — it was never released,
so there is nothing to deprecate.
- LoopManager.set(): next_due_at = now for both cadence modes
- drop --start-now parsing, the persisted LoopState.start_now field, and
the flag from help text; confirmation now always says the first wakeup
fires now
- tests updated to pin the new default (incl. the TUI not-due test, which
now has to push next_due_at out explicitly)
- docs: quick-start and command table describe the immediate first run
_runtime_model_config merges the agent's current identity onto the row's
existing model_config JSON. For model and provider it only SET the key
when the agent attribute was truthy, while base_url/api_mode/
reasoning_config/service_tier already deleted stale values when falsy.
When an agent rebuilt with an empty provider (inheriting the profile
default) was persisted, the previous provider/endpoint survived in
model_config while _persist_live_session_runtime updated the model
column separately. Resume then read the fresh model from the column but
the STALE provider from model_config, silently routing the resumed chat
to the wrong endpoint (e.g. a VeniceAI/empero route under a model that
should run on the profile default).
Apply the same delete-on-falsy rule to model and provider, mirroring the
or-None deletion the CLI path (_persist_model_switch_to_session) already
uses, so a stale session state can never survive into a resume override.
Existing desynced rows self-heal on the next live metadata persist.
Adds regression tests: merge drops stale provider/model when the agent
attribute is falsy, a truthy provider overwrites the stale value, resume
overrides fall back to the billing provider instead of the stale
endpoint, a real-DB round trip heals an already-desynced row, and a
first write (existing=None) reflects only the agent's current identity.
Drives the REAL session.resume -> _make_agent -> AIAgent path (eager_build)
through tui_gateway.server.handle_request against a real seeded state.db in
an isolated HERMES_HOME — no mocks. Three scenarios from the live report:
deleted provider falls back to default, renamed provider heals, and a legacy
canonical Bot Chat (no follow_profile_config marker) follows the profile's
current config.
A/B verified: all 3 FAIL on origin/main with the reported
"resume failed: Unknown provider '<name>'"; 3/3 pass with the salvaged
fixes + backfill.
Bot Chats created before the follow_profile_config marker existed carry no
contract in model_config, so they would stay pinned to a stale stored
provider until deleted — the exact shape of the live reports (#89497,
#94818). Mirror the plugin's own identity rule (the profile's session
titled exactly 'Bot Chat') as a legacy fallback in
_stored_session_runtime_overrides, matching the room-plumbing legacy
'Group:' title fallback.
Follow-up to the salvaged #90343 (@curator8888) and #96111 (@lorzl).
Bot-Mode canonical chats (the ONE forever DM per bot) and room plumbing
sessions are plugin-owned scratch conversations. They are now created with
an explicit follow_profile_config contract, persisted in the session row's
model_config, so session.resume rebuilds from the member profile's CURRENT
config instead of restoring the stored model/provider pin from an old row.
That stale pin is what left bot DMs stuck on a dead provider (e.g. 'out of
Nous credits' after the profile was switched to ollama-cloud) while the
same bot worked fine in rooms — the mirror image of the room-plumbing bug
(#89497 class). Normal 1:1 user chats keep the stored-runtime restore:
opening an older chat must show the model it actually used.
- tui_gateway/methods_session.py: accept follow_profile_config on session.create
- tui_gateway/server.py: persist the marker in the row; skip stored-runtime
overrides on resume when present
- apps/desktop/src/plugins/hermes-bots/plugin.js: send the contract from
createCanonicalChat and ensureGroupChatSession
- tests: backend override + row-persist coverage; desktop source-contract
coverage for both session kinds
Room member sessions in Bot Mode are per-member scratch conversations
inside a group chat. session.resume restored their stored model/provider
pin from the row's model_config, so a room bot stayed stuck on whatever
provider was pinned when the row was first written — even after the
profile was switched. Every room message then failed on the stale
provider (e.g. 'out of Nous credits' after switching a profile from
Nous to ollama-cloud) while the same bot worked fine in DMs.
Add an explicit room_plumbing contract:
- session.create accepts room_plumbing: true, persisted in model_config
- _stored_session_runtime_overrides() returns {} for marked rows, so a
room session always rebuilds from the member profile's CURRENT config
- hidden + 'Group:' title shape is kept as a legacy fallback for rows
created by older desktop builds that never sent the marker; hidden
non-room chats keep the stored-runtime restore
- Desktop Bot Mode sends room_plumbing: true when creating the hidden
per-member room sessions
Fixes#89497
A session row persists the provider identity a chat actually used. When that
provider is later renamed or removed (e.g. a custom_providers:/providers:
entry deleted, or a provider renamed oldone->newone), Desktop/TUI resume
restores the stale name into agent init and dies with:
agent init failed: Unknown provider '<name>'
while the CLI resumes the same session fine with the configured default.
- runtime_provider: add is_routable_provider() (full resolution chain:
built-in -> providers: -> custom_providers: -> models.dev)
- _stored_session_runtime_overrides: heal a non-routable provider via
canonical_custom_identity (base_url -> model -> configured provider),
drop to the configured default when unrecoverable, and clear the stale
base_url after healing so a dead endpoint cannot override the registry URL
- _start_agent_build: gate deferred-resume overrides on provider routability;
when the stored provider is gone, prefer the model the user picked for THIS
session, else the configured default
- tests: is_routable_provider cases, heal/fallback round-trips, gate checks
Refs #75128
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the
same `tour(action='targets')` discovery call: one bubble with an arrow, for a
sentence that would be clearer with a finger on the thing it's about. Dimming
the whole app to say "the model name is a button" is the wrong weight.
Fire-and-forget rather than a round-trip, because a tip is not a question and
blocking the turn on one would stall the reply it belongs to. The renderer
enforces the user's opt-out itself, so a stale config read can never put a
bubble on a screen that asked for none.
Automatic Desktop ends (ws_orphan_reap, disconnect, idle, LRU, shutdown)
now drop the local runtime but keep the state.db row and durable-key
delegations when another live lease still owns the session.
Co-authored-by: metamindedu <metamind@kakao.com>
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.
Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.
Credit: @Finn763