4e3feb8bbb
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts now hold a lease on the TTS engine. Acquiring pre-loads the configured provider (piper/kittentts model into the same LRU slot synthesis reads; lazily-installed cloud SDKs), so the first spoken reply no longer pays the model load as dead air. Releasing the last lease across surfaces unloads resident local models. - tools/tts_tool.py: warm_tts_provider / release_tts_provider / acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES registry; piper/kittentts loaders extracted so warm-up and synthesis share one resolution path. - web_server: POST /api/audio/tts-lease (profile-scoped, off-loop, failures reported in body never as HTTP errors). - tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease. - desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest intent wins) driven from useComposerVoice; setTtsLease API client. - docs: features/tts.md section. Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold → 92ms after the toggle warmed the engine; release drops the model.