fix(stt): kill faster-whisper silence hallucinations at the source
Local faster-whisper called model.transcribe with bare {'beam_size': 5}:
no VAD, cross-window conditioning on, no confidence filtering. Pure
silence produced hallucinated tokens (E2E: 5s anullsrc WAV -> 'You',
no_speech_prob=0.705) and noisy clips could produce runs of junk, often
in other languages.
Three-layer class fix, one shared owner for every local-whisper call
site (build_local_transcribe_kwargs):
1. Silero VAD filter (bundled with faster-whisper) on by default —
silence never reaches the model. stt.local.vad: false restores the
raw behavior for music/ambient transcription.
stt.local.vad_min_silence_ms tunes chunk splitting (default 500).
2. condition_on_previous_text=False — one hallucinated token can no
longer seed a self-reinforcing run; negligible cost for
voice-note-length audio.
3. Segment confidence gate (_join_confident_segments): drop a segment
only when no_speech_prob > 0.6 AND avg_logprob < -1.0 (openai-whisper's
own heuristic shape; both must hit so quiet-but-real speech survives).
Config: stt.local.no_speech_prob_threshold / logprob_threshold.
The WHISPER_HALLUCINATIONS blocklist in voice_mode.py stays as
last-resort defense but should now almost never fire.
E2E (real faster-whisper 'base', CPU int8):
silence.wav before 'You' -> after ''
noise.wav before '' -> after ''
speech.wav before/after 'Hello World, this is a test of the
transcription system.' (unchanged)
Docs (EN + zh-Hans), DEFAULT_CONFIG, cli-config.yaml.example updated;
19 unit tests (kwargs contract, off-switch, confidence gate incl.
quiet-speech survival, _transcribe_local wiring), sabotage-verified.
This commit is contained in:
@@ -1138,6 +1138,12 @@ stt:
|
||||
model: "base" # tiny | base | small | medium | large-v3 | turbo
|
||||
# language: "" # auto-detect; set to "en", "es", "fr", etc. to force
|
||||
# initial_prompt: "" # Optional faster-whisper prompt, e.g. bias Chinese output to simplified Chinese
|
||||
# --- Anti-hallucination hardening (whisper decodes junk from silence without these) ---
|
||||
# vad: true # Silero VAD filter (default on) — silence never reaches whisper.
|
||||
# # Set false to restore raw behavior (e.g. transcribing music/ambient audio).
|
||||
# vad_min_silence_ms: 500 # min silence (ms) that splits speech chunks when vad is on
|
||||
# no_speech_prob_threshold: 0.6 # drop a segment only if no_speech_prob > this...
|
||||
# logprob_threshold: -1.0 # ...AND avg_logprob < this (both must hit — quiet real speech survives)
|
||||
language: "en" # GLOBAL language hint for every STT provider (per-provider language wins). Set "" for auto-detect.
|
||||
# groq:
|
||||
# model: "whisper-large-v3-turbo"
|
||||
|
||||
Reference in New Issue
Block a user