fix(stt): consume lazy whisper segments inside the CUDA→CPU retry guard

faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2
dlopens the CUDA runtime on the FIRST encode, which happens while the segments
are iterated in `_join_confident_segments()` — outside the try/except that
implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an
NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux
`libcublas.so.12`) the model loads fine, the error escapes the guard, and every
voice note fails with "Local transcription failed: Library cublas64_12.dll is
not found or cannot be loaded" until the user pins `stt.local.device: cpu`.

Materialize the segments inside the guarded block (first attempt and CPU retry)
so the dlopen failure reaches the existing evict-and-retry-on-CPU path.

Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class).
Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`,
`cudnn64_`, `cudart64_`) — the Windows message already matches the existing
"cannot be loaded" marker, proven by the live probe with the reporter's exact
string; the 6-test file was reduced to 2 invariant tests in the existing suite.

Fixes #111929
Fixes #105295
Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate)

Co-authored-by: atmaksri <sri.atmakur@gmail.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
This commit is contained in:
Sahil Vishnalya
2026-09-15 11:41:24 -07:00
committed by Teknium
parent 0ff9941b76
commit afe9e25c57
2 changed files with 64 additions and 0 deletions
+7
View File
@@ -356,6 +356,12 @@ def _transcribe_local(
if v})
try:
segments, info = model.transcribe(file_path, **transcribe_kwargs)
# faster-whisper's transcribe() is lazy: the decode (and with it the
# dlopen-on-first-use of the CUDA runtime on Windows) happens while
# ITERATING segments, after this call has already returned (#103793).
# Consume inside the guard so a first-use cuBLAS/cuDNN load failure
# retries on CPU exactly like a load-time failure does.
segments = list(segments)
except Exception as exc:
# CUDA libs can fail at dlopen-on-first-use, AFTER loading: evict the poisoned
# cached model, reload on CPU and retry once, else every later message fails.
@@ -365,6 +371,7 @@ def _transcribe_local(
"evicting cached model and retrying on CPU (int8).", exc)
model = _replace_cached_model_on_cpu(model_name)
segments, info = model.transcribe(file_path, **transcribe_kwargs)
segments = list(segments)
transcript = _join_confident_segments(segments, local_cfg)
logger.info("Transcribed %s via local whisper (%s, lang=%s, %.1fs audio)",
Path(file_path).name, model_name, info.language, info.duration)