_live_visible_history backs the payload a tab switch repaints from. Reading
the active-only projection made a compacted chat collapse to its summary on
switch while REST still served the whole thing.
Two sites dropped profile ownership on the tile-transcript refresh path:
- tui_gateway/server.py _sessions_sig statted only the launch home's
state.db, so a turn landing in a served sibling profile's store never
produced sessions.changed. The watcher now also probes every profile
home _profile_home() has resolved for this backend (empty set on
single-profile installs — behavior byte-identical there).
- use-background-sync.ts reconcileTileTranscripts read the tile's
transcript unscoped; it now passes the tile's ownerRoute scope, the
same way reconcileActiveTranscript already does for the main pane, and
keys the change signature by owner.
Reimplemented minimal from PR #99333 (the PR head's commit identity does
not match the GitHub author).
Co-authored-by: StodsEcho5 <250208229+StodsEcho5@users.noreply.github.com>
`_find_live_session_by_key` matched live runtimes by bare stored session id.
Stored ids are timestamp-based and can exist in more than one profile's
store, so `session.resume` for profile B (fast path, post-build re-check, and
`_claim_or_reuse_live`) could hand back profile A's live runtime — the turn
then ran with A's persona/tools and wrote A's memory (#100029).
Give the lookup an optional `profile_home` (default: any profile, unchanged
for callers that have no profile to scope by) using the same string compare
`_find_live_unpersisted` already uses, and pass the resolved home at every
resume/claim site. `_claim_parked_runtimes` gets the same scope so a resume
under B never finalizes A's parked runtime of the same id.
Reimplements the profile-scope half of #100213 by @Finn763; the Group-title
capability-sync change from that PR is intentionally not carried.
Co-authored-by: Finn763 <165816600+finn763@users.noreply.github.com>
Compression rotates a conversation's tip id while tiles stay keyed by
whichever segment id they were opened with. focusOpenSession and
openSessionTile tested exact ids, so right after a rotation the same
chat read as 'not open' and opened again in a second tab — and a tile
keyed to a MIDDLE segment (the tip when it was opened) could no longer
prove it names the conversation at all, rendering as an untitled ghost.
The projected list row now carries the full chain
(SessionDB.get_compression_chain, served as _lineage_ids by
list_sessions_rich and the sidebar tree row), lineageAliases indexes
every segment, sessionMatchesStoredId accepts membership, and the tab
focus/open paths dedupe through the lineage instead of the exact id.
Older gateways omit the field and degrade to today's root/tip pairing.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
Manual /compress on a compute-host (turn_isolation) session blocked its RPC
waiter for a hard-coded 120s, answered error 5019, and then DROPPED the
host's late `control.ack`: HostSupervisor.control() popped the pending
queue in `finally`, so `_handle_host_frame` had nothing to deliver to. The
host kept compressing, succeeded minutes later, rotated the session — and
the gateway session never mirrored the new session_key/history_version and
the desktop never refreshed its transcript.
- host_supervisor: `control(..., on_late_ack=)` leaves a one-shot handler
registered when the waiter times out; control.ack/control.error/error
frames for that request_id fire it (bounded: 30min TTL, cap 64). A host
crash fails outstanding handlers with a synthetic control.error.
- server: `_compute_host_compress_wait_seconds()` derives the wait from
`compression.context_total_ceiling_seconds` (+30s slack, floor 120s,
cap 630s) instead of the literal 120. `_adopt_late_compute_host_compress_ack`
applies the metadata mirror and emits the same `session.info` a normal
compress does plus the existing `status.update kind=compacted` edge; a
late error goes out through the existing `error` event.
- session.compress / slash.compress (methods_tools + _mirror_slash_side_effects):
on waiter timeout answer `status: pending` (not 5019) and register the
late-ack handler.
- desktop: SESSION_COMPRESS_TIMEOUT_MS 120s -> 660s (above the gateway cap);
`status: 'pending'` renders as an info notice, not `error:`; the
`compacted` status edge rehydrates an idle active session's transcript
(mid-turn compaction still defers to the turn settle path).
Minimal extraction of the design in #99630 by @vsd2807 (design trace by
@andrexibiza and @JoaoMarcos44 in the #97948 thread); no new DB tables,
modules, or polling protocol.
Refs #97948
Co-authored-by: VVV <vaibhavdahiya28@gmail.com>
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.
- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
workers by exact delegation_id (heuristic shape/time grouping kept for
older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
returned by the dispatch and used for cache/delegation/live/<id>/.
Review point from keeltrace, and it is a real gap: secret redaction and
prompt omission are different contracts, and only the first one is
pattern-shaped.
redact_sensitive_text removes credentials. A provider 4xx that quotes the
request back carries the user's own prose - a paragraph about a person, a
file pulled in by an @ reference - which matches no credential pattern and
so passed through untouched into cause=. The record's stated contract is
that prompt content is not logged, and the previous commit only enforced
the half of it that a regex can see. The existing prompt test could not
catch this: its provider error does not echo the prompt, so it proves the
prompt is not logged directly, not that it cannot arrive by being quoted.
_strip_prompt_echo closes the quoted path directly. Anything the message
shares with the submitted prompt for 24 characters or more becomes
<prompt>. Shingle-set matching rather than a diff, so cost is linear in
both strings on a path that runs for every failed turn and can face an
@-expanded prompt of arbitrary size; the JSON-escaped form of the prompt is
shingled too, because a provider handing back its own request body often
hands it back escaped. The prompt is captured after @-expansion on purpose:
an injected file's contents are exactly the material an echo would carry,
and they are not in the submitted text.
Ordering is load-bearing. The strip runs after the whitespace collapse, so
a re-wrapped quote still matches, and before the length cap, so a quote
cannot survive by being cut mid-run.
What this does not claim: verbatim echo is what it stops. A paraphrase, a
summary, or a re-encoding would survive it. The alternative keeltrace
raised - log only structured provider metadata and drop the message body -
is airtight but costs the diagnosis this PR exists to enable, since the
reporter needed to tell a 402 from a crashed finalizer. Happy to switch if
maintainers prefer the stricter contract.
Tests: the non-secret sentinel keeltrace asked for (a benign phrase present
only in the prompt, echoed by the provider error, asserted absent from the
record), plus guards that a message sharing nothing with the prompt is
untouched, that an overlap below the window is not treated as an echo, that
a prompt shorter than the window cannot blank the message, that a
JSON-escaped echo is stripped, that the strip precedes the length cap, and
that whitespace shape does not hide an echo. The three that cover the new
path fail with the strip removed; the guards pass either way.
Fixes#89117
The whole of #89117 is two log lines:
tui_turn finished: ui_session=0dfcee58 status=error error_retained=True duration=0.9s
A provider 4xx, a budget wall, a billing block and a crashed finalizer all
produce exactly those characters, so an intermittent failure cannot be
triaged from the one record that is guaranteed to exist.
The bookend came from #86865, which added it to trace compression
rotations across #86647 -- identities and a coarse status were the job, and
content was deliberately excluded. What that leaves is a returned-error
path (provider 4xx, budget, billing) which writes no other log line at all.
The exception path at least prints `[gateway-turn] <Type>: <msg>` to
stderr, so the failures that go unlogged are exactly the sub-second ones
this issue is about.
Both failure paths now stash a one-line cause, and the bookend appends it.
The record keeps its shape when nothing failed: a successful turn gains no
new fields.
The cause is redacted with `redact_sensitive_text(force=True)` and capped at
240 characters with a visible ellipsis, because a 4xx body routinely quotes
the request that produced it -- adding the cause without redacting it would
write an Authorization header the user never chose to log. Redaction fails
closed: if the redactor cannot run, the fragment reads `<unredactable>`
rather than the raw message. Whitespace is collapsed so a multi-line
provider body cannot split the record, which is the only property that
makes it greppable for a bug like this one.
12 regression tests. Four mutations proven: disabling the helper fails 9,
dropping redaction fails 2, dropping truncation fails 1, wiring only the
exception path fails 4.
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.
- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
registry; piper/kittentts loaders extracted so warm-up and synthesis
share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.
Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
config.set matches an explicit key list and answers 4002 for anything
else, so a renderer mirroring an unlisted key wrote nothing at all. The
reactions toggle shipped that way: every write was rejected into a
swallowed .catch(), and react_to_message stayed dark no matter what the
user picked.
Adds the display booleans as a recognized group so the toggle reaches
the config of whichever gateway the app is actually talking to, which is
the only place a check_fn can read it.
A confirmed Desktop Stop left the crash-recovery marker on disk until the
run thread finished. If the backend exited in that window, resume treated
the leftover as a crash and auto-continued the turn the user had stopped.
Co-authored-by: Jaime Chieng <164842890+buddhaholic420@users.noreply.github.com>
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the
maintainer-review revision points folded in during the rebase:
- the lease engages UNCONDITIONALLY: try_acquire_active_session no longer
returns a disabled no-op lease when max_concurrent_sessions is unset;
the concurrency cap stays an orthogonal, optional policy checked second
- ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE)
instead of degrading to an untracked go-ahead: a corrupt/unreadable
registry must not be collapsed into 'no owner exists' (review blocker 2)
- the ownership admission sits at the _run_prompt_submit chokepoint that
EVERY fresh turn source crosses, and crash auto-continue acquires (or
bails) BEFORE emitting message.start — closing the #94778 bypass where
backend B's auto-continue ran a duplicate turn while backend A was live
(review blocker 1)
- the TUI gateway claim helper fails closed on claim exceptions for every
surface, not just desktop
- CLI and messaging-gateway call sites pass live_session_id metadata so
the (pid, live id) re-entrancy identity protects them from self-fencing
on a leaked lease
Co-authored-by: teknium1 <teknium1@users.noreply.github.com>
Follow-up to the salvaged #98571: forward the interrupt to the compute
host whenever the parent 'running' mirror is stale, but only for
sessions that actually have hosted activity — HostSupervisor.interrupt()
calls start(), so an unconditional forward would spawn a compute-host
child just to deliver an interrupt for an idle lazy session.
Adds a regression test asserting the idle-lazy-session no-spawn path.
Refs #92916
Idle and preflight compaction arrived as lifecycle status without the
"Compacting context" marker, so TUI never entered a compacting state.
Re-tag those lines and freeze the busy FaceTicker on "compacting" for
the whole pause instead of restoring "running…" after 4s.
The salvaged #98948 change returned False for every db=None, which also fired in deliberately store-less/degraded contexts (no _db_error), regressing six prompt.submit tests. Gate the loud failure on _db_error being set — the actual #98924 symptom — and keep the pinned best-effort contract otherwise.
Companion to #98935, which fixes _fts_table_probe itself. This covers the
surfaces that PR does not touch:
- web_server._open_session_db_at_path: the one-writable-open heal only
caught sqlite3.DatabaseError; a raw UnicodeDecodeError (pysqlite failing
to decode SQLite's own error message over corrupt file bytes) bypassed
it, so the heal documented for malformed schema never fired (#98924
Failure 1). Both catches widened; decode errors dispatch to the heal.
- SessionSchemaMixin._recover_stale_fts_locked: drop-and-recreate skipped
vtables whose probe raised UnicodeDecodeError, the same too-narrow
catch the issue identified in the probe.
- TUI gateway: _ensure_session_db_row returned silently when the store
could not open, so prompt.submit streamed the turn while persisting
nothing (#98924 Failure 2). It now returns False and prompt.submit
fails the RPC with code 5072 so desktop maps it to a toast, mirroring
the disk-full/5070 convention. session.create stays silent per its
pinned degraded-mode contract.
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).
NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.
Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.
Fixes#99222
config.get and config.set ignored the focused profile on a shared
app-global backend, so reads and persistent writes used the launch
config.yaml. Bind the existing @_profile_scoped decorator and write
_save_cfg through the request home override.
Fixes#95760
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
Resume was attaching an unused store as {todos: [], revision: 0} and the desktop rejected tool.start updates that have no revision. A merge:true start after reconnect never patched the list until complete.
Skip unused empty snapshots. Apply unversioned updates without moving the watermark so a later todo.updated can still win.
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
avg_latency_s, avg_tps (reads the same per-call deque history from
agent/conversation_loop.py; keys omitted when no data — Codex
app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
segments (breakpoints 96/104/110 cols, lowest priority — they shed
first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
filters TUI segments too: cache_hit, latency, tps, duration,
compressions, bg_tasks, bg_subagents, voice, battery, title,
context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
The eager _read_persisted_todo_state(db, target) added a second
get_messages_as_conversation call on every resume, breaking the
one-lineage-SELECT contract pinned by
test_session_resume_uses_parent_lineage_for_display. Derive the
snapshot from the history each resume path already loaded instead;
deferred (defer_history) resumes cache it in the hydration worker once
the transcript arrives.