fix(stt): consume lazy whisper segments inside the CUDA→CPU retry guard
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2 dlopens the CUDA runtime on the FIRST encode, which happens while the segments are iterated in `_join_confident_segments()` — outside the try/except that implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux `libcublas.so.12`) the model loads fine, the error escapes the guard, and every voice note fails with "Local transcription failed: Library cublas64_12.dll is not found or cannot be loaded" until the user pins `stt.local.device: cpu`. Materialize the segments inside the guarded block (first attempt and CPU retry) so the dlopen failure reaches the existing evict-and-retry-on-CPU path. Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class). Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`, `cudnn64_`, `cudart64_`) — the Windows message already matches the existing "cannot be loaded" marker, proven by the live probe with the reporter's exact string; the 6-test file was reduced to 2 invariant tests in the existing suite. Fixes #111929 Fixes #105295 Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate) Co-authored-by: atmaksri <sri.atmakur@gmail.com> Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com> Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>
This commit is contained in:
@@ -356,6 +356,12 @@ def _transcribe_local(
|
||||
if v})
|
||||
try:
|
||||
segments, info = model.transcribe(file_path, **transcribe_kwargs)
|
||||
# faster-whisper's transcribe() is lazy: the decode (and with it the
|
||||
# dlopen-on-first-use of the CUDA runtime on Windows) happens while
|
||||
# ITERATING segments, after this call has already returned (#103793).
|
||||
# Consume inside the guard so a first-use cuBLAS/cuDNN load failure
|
||||
# retries on CPU exactly like a load-time failure does.
|
||||
segments = list(segments)
|
||||
except Exception as exc:
|
||||
# CUDA libs can fail at dlopen-on-first-use, AFTER loading: evict the poisoned
|
||||
# cached model, reload on CPU and retry once, else every later message fails.
|
||||
@@ -365,6 +371,7 @@ def _transcribe_local(
|
||||
"evicting cached model and retrying on CPU (int8).", exc)
|
||||
model = _replace_cached_model_on_cpu(model_name)
|
||||
segments, info = model.transcribe(file_path, **transcribe_kwargs)
|
||||
segments = list(segments)
|
||||
transcript = _join_confident_segments(segments, local_cfg)
|
||||
logger.info("Transcribed %s via local whisper (%s, lang=%s, %.1fs audio)",
|
||||
Path(file_path).name, model_name, info.language, info.duration)
|
||||
|
||||
Reference in New Issue
Block a user