6 Commits

Author SHA1 Message Date
teknium1 dd1baee0e4 refactor(secrets): drop scope-aware env shims; runtime_provider and the voice/xai tools read the canonical getters
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.

tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.

Behavior change: none.
2026-09-13 05:07:50 -07:00
Teknium 62732c7d8b simplify(compat): tools/transcription_tools — drop 47 re-exports/aliases, repoint 4 callers + 9 test files 2026-09-03 13:17:25 -07:00
Teknium 8703bc9465 refactor(tools): voice_mode/voice_client_config — suppress() for swallow-blocks, compact docstrings, fold one-shot locals 2026-09-02 22:35:56 -07:00
Teknium b2e9f0c962 refactor(tools): voice_client_config — compact module docstring, banners and layout 2026-09-02 22:16:40 -07:00
Teknium 2ef1e8e4e0 refactor(tools/voice): extract tts delivery + wake_word engines; dedupe transcription/voice_mode helpers; compact tts providers 2026-09-02 14:45:15 -07:00
Teknium 10f0d2278b feat(desktop): client-direct voice — use the active profile's STT/TTS keys from the desktop, no audio relay
Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.

Backend:
- tools/voice_client_config.py: single resolver; per-provider client
  wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
  elevenlabs-tts). Server-host-only providers (local whisper, edge,
  command/plugin) and missing credentials resolve to {mode: relay}.
  xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
  same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).

Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
  profile) with 60s TTL, provider-direct STT + TTS calls, sentence
  cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
  first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
  startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
  below it. Barge-in via the same stopVoicePlayback sequence bump.

Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.

Docs: voice-mode.md client-direct section ships in this PR.
2026-08-23 03:54:52 -07:00