feat: add STT voice transcription for all messaging channels (#28)
* feat: add STT voice transcription for all channels Automatically transcribes audio/voice messages (Telegram, WeChat, Slack, etc.) into text before the agent sees them. Enabled via config, off by default. Changes: - EvoScientist/stt.py: new STT engine using faster-whisper with lazy model loading and per-language model selection (zh/en/auto) - EvoScientist/channels/base.py: hook in _enqueue_raw() to transcribe audio files and prepend transcript to message text; removes the raw [voice: ...] annotation after successful transcription so the agent does not attempt further audio processing - EvoScientist/config/settings.py: stt_enabled (default False), stt_language (default "auto") - pyproject.toml: optional [stt] dependency group (faster-whisper>=1.0) - tests/test_stt.py: unit tests covering all backends and channel integration Usage: pip install 'EvoScientist[stt]' EvoSci config set stt_enabled true EvoSci config set stt_language zh # zh / en / auto Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: remove unused imports (ruff F401) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: address PR #28 reviewer feedback Changes per SemiGlassFace review (CHANGES_REQUESTED): 1. Cache config at channel __init__ — no longer calls load_config() on every incoming message; STT settings stored as instance attributes (_stt_enabled, _stt_language, _stt_model, _stt_device, _stt_compute_type) set once during Channel.__init__(). 2. Replace deprecated asyncio.get_event_loop() with get_running_loop() to avoid DeprecationWarning on Python 3.12+. 3. Annotation removal now uses exact path matching instead of substring search — checks fp == a or a.endswith(f": {fp}]") so only the correct annotation is removed after transcription. 4. Expose stt_model, stt_device, stt_compute_type as config fields so users can override the HuggingFace model id, inference device, and quantisation without touching code. transcribe_file() forwards all three to the engine. Also: _engines dict replaced with single _engine + _engine_key tuple (model_id, device, compute_type) — reuses cached model unless settings change, simpler than a dict. Tests: 19 STT-specific tests all pass; total 1105 tests green, ruff clean. * fix: resolve ruff lint errors (UP037, I001, PT006) * style: apply ruff format --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -63,6 +63,7 @@ discord = ["discord.py>=2.3"]
|
||||
slack = ["slack-sdk>=3.27", "aiohttp>=3.9"]
|
||||
wechat = ["pycryptodome>=3.20"]
|
||||
qq = ["qq-botpy>=1.0"]
|
||||
stt = ["faster-whisper>=1.0"]
|
||||
oauth = ["ccproxy-api>=0.2.4"]
|
||||
all-channels = [
|
||||
"python-telegram-bot>=21.0",
|
||||
|
||||
Reference in New Issue
Block a user