Specified provider 'CUDAExecutionProvider' is not in available provider names (ONNX Runtime)
You have the CPU-only onnxruntime package, or its CUDA/cuDNN requirements are unmet — so inference silently runs on CPU. Install onnxruntime-gpu alone, and match the CUDA versions its release notes require.
Updated
The error
UserWarning: Specified provider 'CUDAExecutionProvider' is not in available provider names. Available providers: 'AzureExecutionProvider, CPUExecutionProvider'
Or, with the GPU package present but its dependencies unmet:
[E:onnxruntime:Default, provider_bridge_ort.cc:1862 TryGetProviderInfo_CUDA] ... Failed to load library libonnxruntime_providers_cuda.so with error: libcublasLt.so.12: cannot open shared object file: No such file or directory
What it means
ONNX Runtime executes models through execution providers — pluggable backends like CPU, CUDA and TensorRT. You asked for the CUDA provider and the installed runtime does not offer it. Being a warning, inference continues on CPU: the model runs, many times slower, and nothing else complains. The second variant means the GPU provider exists but could not load NVIDIA's libraries at the versions it was built against.
Why it happens
The pip packaging splits the product: onnxruntime is CPU-only; GPU support lives in onnxruntime-gpu. Installing the plain package (or a dependency pulling it in alongside the GPU one — the two conflict) leaves you CPU-bound. The library-loading variant is version drift: each onnxruntime-gpu release targets specific CUDA and cuDNN major versions, and a machine with a different CUDA generation cannot satisfy it.
How to fix it
1. Install exactly one package: the GPU one.
pip uninstall -y onnxruntime onnxruntime-gpu
pip install onnxruntime-gpuUninstall both first — coexisting copies shadow each other unpredictably. Then verify:
import onnxruntime as ort
print(ort.get_available_providers())['TensorrtExecutionProvider', 'CUDAExecutionProvider', 'CPUExecutionProvider']
2. Match CUDA/cuDNN to what your onnxruntime-gpu release expects. The requirements matrix lives in the ONNX Runtime documentation per release — check it rather than guessing, since the targeted CUDA and cuDNN majors change across versions. The missing-library name in the error (libcublasLt.so.12 → CUDA 12 family) tells you which generation it wants. Installing the matching CUDA/cuDNN via pip's nvidia packages or conda usually settles it.
3. Request providers with a CPU fallback, and log what you got.
sess = ort.InferenceSession(
"model.onnx",
providers=["CUDAExecutionProvider", "CPUExecutionProvider"],
)
print(sess.get_providers()) # confirm CUDA is FIRST in the listThe session silently falls back when CUDA fails to initialise; printing the resolved list is the difference between knowing and assuming.
4. Check the driver layer beneath, when the provider loads but errors at run time. nvidia-smi must work and its CUDA ceiling must cover the runtime's needs — the standard checks from the driver page.
How to prevent it
Pin onnxruntime-gpu (never both packages) in requirements, note the CUDA/cuDNN pairing in the README, and add the get_providers() assertion to service startup so a CPU fallback fails deployment loudly instead of shipping 10x latency.
Related errors
- Exporting the operator to ONNX is not supported — the export-time counterpart
- torch.cuda.is_available() returns False — the same silent-CPU trap in PyTorch
- CUDA driver version is insufficient
- Could not load dynamic library libcudart