afe9e25c57
faster-whisper's `model.transcribe()` returns a lazy generator; ctranslate2 dlopens the CUDA runtime on the FIRST encode, which happens while the segments are iterated in `_join_confident_segments()` — outside the try/except that implements the CUDA → CPU fallback in `_transcribe_local`. On a host with an NVIDIA driver but no CUDA runtime (Windows `cublas64_12.dll`, Linux `libcublas.so.12`) the model loads fine, the error escapes the guard, and every voice note fails with "Local transcription failed: Library cublas64_12.dll is not found or cannot be loaded" until the user pins `stt.local.device: cpu`. Materialize the segments inside the guarded block (first attempt and CPU retry) so the dlopen failure reaches the existing evict-and-retry-on-CPU path. Salvaged from #103848 by @Sahilvishnaliya (earliest fix of this class). Trimmed during salvage: the `_CUDA_LIB_ERROR_MARKERS` additions (`cublas64_`, `cudnn64_`, `cudart64_`) — the Windows message already matches the existing "cannot be loaded" marker, proven by the live probe with the reporter's exact string; the 6-test file was reduced to 2 invariant tests in the existing suite. Fixes #111929 Fixes #105295 Part of #103793 (the CPU fallback now fires; GPU-wheel install is separate) Co-authored-by: atmaksri <sri.atmakur@gmail.com> Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com> Co-authored-by: isoenthusiast <287677567+isoenthusiast@users.noreply.github.com>