fix(stt): kill faster-whisper silence hallucinations at the source

Local faster-whisper called model.transcribe with bare {'beam_size': 5}:
no VAD, cross-window conditioning on, no confidence filtering. Pure
silence produced hallucinated tokens (E2E: 5s anullsrc WAV -> 'You',
no_speech_prob=0.705) and noisy clips could produce runs of junk, often
in other languages.

Three-layer class fix, one shared owner for every local-whisper call
site (build_local_transcribe_kwargs):

1. Silero VAD filter (bundled with faster-whisper) on by default —
   silence never reaches the model. stt.local.vad: false restores the
   raw behavior for music/ambient transcription.
   stt.local.vad_min_silence_ms tunes chunk splitting (default 500).
2. condition_on_previous_text=False — one hallucinated token can no
   longer seed a self-reinforcing run; negligible cost for
   voice-note-length audio.
3. Segment confidence gate (_join_confident_segments): drop a segment
   only when no_speech_prob > 0.6 AND avg_logprob < -1.0 (openai-whisper's
   own heuristic shape; both must hit so quiet-but-real speech survives).
   Config: stt.local.no_speech_prob_threshold / logprob_threshold.

The WHISPER_HALLUCINATIONS blocklist in voice_mode.py stays as
last-resort defense but should now almost never fire.

E2E (real faster-whisper 'base', CPU int8):
  silence.wav  before 'You'                        -> after ''
  noise.wav    before ''                           -> after ''
  speech.wav   before/after 'Hello World, this is a test of the
               transcription system.' (unchanged)

Docs (EN + zh-Hans), DEFAULT_CONFIG, cli-config.yaml.example updated;
19 unit tests (kwargs contract, off-switch, confidence gate incl.
quiet-speech survival, _transcribe_local wiring), sabotage-verified.
This commit is contained in:
Teknium
2026-07-28 22:51:32 -07:00
parent aac753dd05
commit bf8004e3a8
6 changed files with 294 additions and 12 deletions
+6
View File
@@ -1138,6 +1138,12 @@ stt:
model: "base" # tiny | base | small | medium | large-v3 | turbo
# language: "" # auto-detect; set to "en", "es", "fr", etc. to force
# initial_prompt: "" # Optional faster-whisper prompt, e.g. bias Chinese output to simplified Chinese
# --- Anti-hallucination hardening (whisper decodes junk from silence without these) ---
# vad: true # Silero VAD filter (default on) — silence never reaches whisper.
# # Set false to restore raw behavior (e.g. transcribing music/ambient audio).
# vad_min_silence_ms: 500 # min silence (ms) that splits speech chunks when vad is on
# no_speech_prob_threshold: 0.6 # drop a segment only if no_speech_prob > this...
# logprob_threshold: -1.0 # ...AND avg_logprob < this (both must hit — quiet real speech survives)
language: "en" # GLOBAL language hint for every STT provider (per-provider language wins). Set "" for auto-detect.
# groq:
# model: "whisper-large-v3-turbo"