Prose false positives (Branch A trailing boundary, Branch D leading
boundary), quoted-multiline data payloads, and the fail-closed
unbalanced-quote fallback.
Follow-up to @BrunoBza's #93062:
1. Set failed=True only for the new repeated_outer_errors exit reason.
Previously the error exit left failed=False, so finalize_turn reported
completed=True for a turn that actually failed — incorrect.
2. Don't append_message the assistant response at the break. A thinking-
prefill or interim assistant may already be the tail, and appending
would create assistant→assistant role-alternation violation.
finalize_turn (lines 341-353) handles this safely by checking
_tail_role != 'assistant' before appending.
3. Update test to assert failed=True and completed=False for the
repeated_outer_errors exit.
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.
Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.
The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.
Fixes#92450
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.
Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
(matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
predicates now delegate to the shared home (tool_executor previously
matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
shutdown signal and abandons the turn — one log warning, no print,
no traceback, no debug dump, no retry; outer handler gets the same
guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
(set by the background-review fork) instead of bypassing it.
Refs #55924#58720 (same class in cron delivery), adjacent to #90683.
Test-helper + regression test from PR #92633 by @DavidMetcalfe:
mock _is_supervised_gateway_process instead of setting the raw env
var, and add a test proving a CLI agent session with inherited
_HERMES_GATEWAY=1 but no PID ownership is no longer blocked.
Follow-up to the #91701 salvage: persistent registrations survive a routine
unload-all, but must not outlive their plugin.
- Targeted unload (plugin disable/uninstall) now gathers persistent rows
from the ownership ledger and disposes them.
- Unload-all parks live persistent handles in _persistent_carryover;
discover_and_load(force=True) evicts the ones whose plugin did not
re-register the same (kind, key) — superseded handles are dropped
without disposal so a same-object re-registration stays live.
The dashboard auth registry is process-global, but a bundled auth provider
was registered under the per-home plugin manager's scope and enrolled in that
manager's reverse-order teardown. A per-home manager is unloaded routinely
(profile-scoped dashboard activity, forced re-discovery), and that teardown
disposed the registration — emptying the auth registry for the whole process
and permanently disabling sign-in until restart.
Register dashboard-auth providers in the process-global slot as persistent
host-owned registrations kept out of per-home manager teardown, so a routine
unload can no longer disable authentication process-wide. Registration upserts,
so a forced re-discovery (e.g. a password change) still rotates the provider in
place. The test-only manager reset now clears the auth registry too, since
persistent registrations deliberately survive unload.
Fixes#91701
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.
Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.
The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
- Clear the accidental end stamp on resurrection (at the lineage tip):
a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
archive auto-resurrect on the next lookup — the user could never retire
the canonical chat. Test pins the resurrect -> deliberate-archive ->
stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
compressed lineage carries end_reason='compression', so tip-stamped
accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
dm resolution) filtered archived rows out via list_sessions_rich and
still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
- Drop the false fairness claim from acquire_turn_lock's docstring (LOCK_NB
probe + sleep retry gives no arrival-order guarantee; only the budget is).
- logger.debug once when the lock degrades to a no-op on fcntl-less
platforms so silent serialization loss stays diagnosable.
- Document the real worst-case deliver handler hold (120s lock wait + 600s
turn = ~720s) where clients tune their timeouts against it.
- Pin non-reentry: local_delivery_command must stay a raw 'hermes -p' argv —
wrapping it in --run-delivery would make the child contend with its
parent's own flock and fail every relay delivery with target_busy.
- De-flake: the cross-profile test's upper-bound wall-time assert tolerates
loaded CI runners; the wait-duration message assert matches ~Ns generally.
Two CI failures: (1) the target_busy test's global subprocess.run patch
recorded unrelated gateway-init git calls (rev-parse/ls-remote) as the
delivery spawn — now local_delivery_command is monkeypatched to a
sentinel argv so only the real delivery path counts; (2) the new
bot_mode config section is single-field, tripping the dashboard
no-single-field-categories rule — folded into the agent tab via
_CATEGORY_MERGE (same as #93102's fix).
Final-diff pass: pin both envelope mtimes via os.utime relative to the
watermark (write_text alone is wall-clock/FS dependent), and reset
relayDrainRerun in stopBotRelay so a rerun remembered mid-drain can't
leak one stale drain into the next start/stop cycle.
Review follow-ups:
- A push signal landing while drainRelayOutboxes is mid-flight hit the
relayDrainBusy early-return and was gone forever — the gateway signature
is monotone (one event per new envelope, never re-broadcast), so the
envelope waited out the full 4s poll, exactly the latency the push path
removes. relayDrainRerun remembers the race and schedules one debounced
follow-up pass after the drain finishes.
- test_new_envelope_after_drain_fires_pending_again pins the untested half
of the monotone contract: the watermark must not eat genuinely NEW
envelopes (write -> drain -> write-newer fires twice). Mutation-checked:
a stale-signature regression fails it while the other three still pass.
Cross-connection DMs were pure polling: the Desktop drains every gateway's
bot_relay outbox on a 4s interval, so each hop eats up to 4s outbound plus
4s for the reply leg (#92760 'bots reply slowly').
Emission point: the gateway's existing change watcher (_CHANGE_WATCHES in
tui_gateway/server.py). Envelopes are written by the AGENT process
(message_agent -> tools.bot_relay.enqueue_envelope), not the gateway, so no
gateway RPC is on the enqueue path and an in-process emit is impossible.
That is exactly the situation the change watcher already solves for the
pairing store (pairing.changed: 'written by a different process; the files
are the only shared signal') - so a new bot_relay.outbox.pending entry in
the existing watch table is the smallest correct diff: one cheap 1s-interval
stat probe folded into the existing 0.5s watcher tick, no new thread, no new
RPC, and _broadcast_global_event fans it to every connected WS client for
free. The signature is monotone (newest envelope mtime ever seen) so a
drain emptying outbox/ never re-fires the event.
Desktop (hermes-bots plugin): subscribe via the existing host.onEvent tap
(feature-detected - older shells lack it) and run drainRelayOutboxes through
a 250ms trailing debounce so a burst of signals collapses to one drain.
The 4s interval poll is intentionally UNCHANGED as the backstop: the event
tap only hears the active gateway socket, so per-connection push detection
would be complex and wrong to trade the poll against - push simply makes
the common case near-instant while older backends keep working exactly as
before.
Tests: 3 new watcher contracts (fires on enqueue, monotone across drain,
silent with no outbox) and a new relay-push-drain.test.mjs (debounce burst
-> one drain, re-arm after window, disposed no-op, poll backstop intact).
Review follow-up: bare \b401\b / \b402\b / \b429\b / \b5xx\b matched any
3-digit token in error text ('line 502', 'took 429 ms'), and server_error
misfires feed AUTO_RETRYABLE — a supervisor could auto-retry a permanent
local failure. Numeric rules now require an 'error code:'/'status:'/'http'
prefix; phrase alternatives (rate limit, server error, overloaded, out of
funds) unchanged. Adds parametrize rows for the false-positive guards and
the previously untested branches (bare 'status: 401', 'upstream server
error', 'model_not_found').
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).
_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.
Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
session.create intentionally persists no state.db row until the first
prompt, but session.resume only looked in the database — so resuming a
live lazy session by its stored key or pending title hard-404'd. Bot
Mode hits this on every fresh non-default bot: the canonical Bot Chat
is created lazily on the profile, the open/send resumes it, and the
user gets 'session not found' on their first message to that bot.
session.resume now falls back to the in-memory session registry,
matching by stored key or pending title scoped to the SAME profile
home. Cross-profile lookups still fail closed; unknown ids still 404.
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).
- update_inventory: restart_via becomes a machine-readable mechanism id
(systemd|launchd|desktop|manual) — THE policy table as data; display
derived via describe_restart_mechanism. match_runtime_outcomes()
reconciles every planned runtime against the restart phase's
bookkeeping (restarted/stopped/failed/unaccounted);
report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
land in the receipt (runtime_outcomes); any unaccounted runtime
escalates exactly like a STALE/DOWN fleet row (exit 1).
Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).
_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.
That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.
Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.
A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.
Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.
tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
Fixes#88654.
After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:
restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
if restart_mode is None:
continue
A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:
cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'
with no operator signal of any kind.
Two changes.
_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.
When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
A /p/<profile>/ URL prefix on a gateway with multiplex_profiles off was
silently ignored: the request was handled as the gateway-owning profile,
so /p/lokaj/v1/toolsets reported the OWNER's platform_toolsets (and every
other profile-owned config read — skills, capabilities, model options,
agent-run toolset resolution — resolved from the owner too). That is the
exact repro in #91583 defect 2: enabling computer_use with
'hermes -p lokaj tools enable computer_use --platform api_server' showed
enabled in lokaj's config while /p/lokaj/v1/toolsets stayed false, and
enabling it on the owner profile flipped it true.
Per-profile capability isolation is the intended design (ruling on
a different profile's config. Multiplexed gateways were already correct —
the profile-prefix middleware enters _profile_runtime_scope and every
canonical config loader honors the HERMES_HOME override contextvar
(verified empirically for load_config, get_config_path and
_load_gateway_config) — the leak was only the non-multiplex fallthrough.
Fix at the one seam both adapters share: _resolve_request_profile now
rejects (404) a prefix naming any profile other than the one the gateway
actually serves. A self-referential prefix (/p/default/ on the default
gateway, /p/lokaj/ on a gateway launched for lokaj) still falls through
so existing well-formed clients keep working. Same change in the webhook
adapter, which had the identical fallthrough. New shared helper
hermes_cli.profiles.profile_matches_home does the home comparison,
fail-closed.
Tests: tests/gateway/test_multiplex_toolsets_profile_isolation.py —
E2E-style with two real profile homes + config.yamls under a temp
HERMES_HOME, real aiohttp routing through the profile-prefix middleware:
per-profile /p/<x>/v1/toolsets isolation for both owner and secondary
(the #91583 repro asserts computer_use true under /p/lokaj only),
cross-profile key rejection, and the fail-closed non-multiplex prefix
for both adapters. Sabotage-verified: reverting the adapter change fails
the 3 fail-closed tests.
Fixes#91583 (defect 2). Repro and live validation by @kubaboski.
The prefix is an address: the caller is naming WHICH agent the request is
for. With gateway.multiplex_profiles off, _resolve_request_profile ignored
the prefix entirely — "don't 404 a would-be valid route" — so a request
explicitly addressed to one agent was silently answered by a different one.
Observed live (Aug 2026): `hermes peer dm mini/researcher` was answered by
the mini's DEFAULT agent with no error on either side, because that host
runs one LaunchDaemon per profile and only the default daemon hosted an
api_server. A wrong-agent answer is strictly worse than an error: the
sender believes the addressee got the message.
With multiplexing off the process serves exactly one profile, so the prefix
is honored when it names that profile (peers address single-profile daemons
this way without knowing the host's topology — get_active_profile_name() is
the same identity the file already uses for model resolution) and rejected
otherwise through the existing _PROFILE_REJECTED path (404). A process that
cannot resolve its own identity rejects too: if it cannot prove who it is,
it must not answer as anyone.
Unprefixed requests are untouched, and multiplexed hosts are untouched —
the change is confined to the prefix-present, multiplexing-off branch that
previously discarded the caller's addressing.
Widen the DM tempfile-leak fix (#91902/#92407) to the sibling sites
PR #92784 introduced:
- tools/bot_relay.py: expose the 6h stale sweep as
cleanup_bot_relay_artifacts() (cleanup_*_cache contract) and wire it
into gateway housekeeping — previously it ran only when the Desktop
drained the outbox, so plaintext envelopes/replies queued while the
Desktop was away could sit on disk forever.
- tui_gateway/methods_bot_relay.py: move the payload write inside the
try/finally so a failed write no longer leaks hermes-relay-dm-*.txt.
- tools/bot_mode_dm.py: _spawn_delivery takes dm_file=None for relay
waiter deliveries, which have no plaintext DM tempfile to reclaim.
The in-band sweep in _write_dm_file only runs when another DM is
written — a gateway that never sends one keeps orphans forever. Expose
the sweep as cleanup_bot_dm_cache() with the same contract as the other
cleanup_*_cache helpers (returns files removed) and wire it into the
gateway housekeeping loop on the hourly media-cache cadence. Also sweeps
legacy hermes-dm-*.txt and hermes-relay-dm-*.txt orphans in the OS temp
root.
Folded in from #92407 (mehmetkr-31), adapted to the runner-owned
cleanup design salvaged from #91902.
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).
Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.
- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
— bounded, interrupt-safe wait over pending notify_on_complete
sessions; reconciles orphaned-pipe exits (#17327) each pass so a
wedged reader cannot burn the full bound; KeyboardInterrupt aborts
the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
session flush / cleanup (covers -q and -Q, i.e. the DM recipient
shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.
Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.
Fixes#90879
Bot Mode always hides canonical 'Bot Chat' sessions, but _find_bot_chat's
GET /api/sessions listing used the default include_hidden=False path, so
the existing hidden row was invisible, _ensure_bot_chat tried to create a
duplicate, and the peer DB's UNIQUE(title) guard rejected it — DM failed.
- api_server: GET /api/sessions now accepts an exact-title lookup
(?title=...) and honors include_hidden=1 ONLY alongside a title filter,
so canonical hidden rows resolve without exposing a blanket hidden
listing on the client surface. The title needle is pushed into SQL
(search_query) so old hidden rows outside the recency window are found.
- peer dm client: _find_bot_chat sends title + include_hidden=1; older
peers ignore the unknown params and degrade to today's behavior.
- Clear diagnosable error on the older-peer duplicate-create rejection,
naming the hidden canonical chat and the PATCH hidden:false workaround.
- Unit tests (hidden resolution, no duplicate create, older-peer error,
older-peer visible fallback) + real-gateway E2E over a real state.db.
Root-cause analysis and regression recipe by @kubaboski in #91583.
Fixes#91583
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
CI has no faster-whisper and lazy installs disabled, so the 'local'
provider resolves through a different unavailability branch than a dev
box — the reason string differs but the relay verdict (and no key in the
payload) is the actual contract.
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.
Backend:
- tools/voice_client_config.py: single resolver; per-provider client
wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
elevenlabs-tts). Server-host-only providers (local whisper, edge,
command/plugin) and missing credentials resolve to {mode: relay}.
xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).
Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
profile) with 60s TTL, provider-direct STT + TTS calls, sentence
cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
below it. Barge-in via the same stopVoicePlayback sequence bump.
Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.
Docs: voice-mode.md client-direct section ships in this PR.