Error database

The installed version of bitsandbytes was compiled without GPU support

bitsandbytes needs a CUDA GPU, and your install (or machine) does not provide one. Upgrade to a recent version on CUDA machines — on Mac or CPU-only boxes, drop 4-bit loading and use a GGUF runtime instead.

The message you saw
The installed version of bitsandbytes was compiled without GPU support

By Updated

The error

Output
The installed version of bitsandbytes was compiled without GPU support. 8-bit optimizers, 8-bit multiplication, and GPU quantization are unavailable.

Transformers raises a related refusal when you try 4-bit loading anyway:

Output
RuntimeError: No GPU found. A GPU is needed for quantization.

What it means

bitsandbytes is the library behind load_in_4bit / load_in_8bit quantization in Hugging Face. Its fast kernels are CUDA code. The installed copy either lacks those kernels for your platform, or there is no usable CUDA GPU for them to run on. Either way, quantized loading through this path is off the table until that changes.

Why it happens

Three distinct situations share the message:

  • No NVIDIA GPU at all — a Mac, or a CPU-only server. No bitsandbytes version can fix hardware.
  • A CUDA machine with a stale bitsandbytes. Older releases had no Windows wheels and shaky detection; a years-old pin produces this on perfectly good GPU boxes.
  • A broken CUDA setup underneath — torch itself cannot see the GPU, and bitsandbytes inherits the blindness.

How to fix it

1. Establish whether PyTorch sees a GPU.

python
import torch
print(torch.cuda.is_available())

False on a machine with an NVIDIA card means the problem is below bitsandbytes — fix that first via the related pages.

2. On a CUDA machine, upgrade bitsandbytes.

bash
pip install -U bitsandbytes

Current releases ship proper Linux and Windows wheels (Windows support landed in 0.43). Then run the built-in self-check:

bash
python -m bitsandbytes

It prints what it detected and whether the GPU kernels loaded.

3. On Mac or CPU-only machines, change approach rather than fight. Remove the quantization config and load small models in full precision:

python
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct",
                                             dtype="auto")

Or run quantized models through a runtime built for CPU and Apple Silicon: llama.cpp (and tools built on it, like Ollama or LM Studio) with GGUF-format models delivers exactly the small-memory experience bitsandbytes cannot give you there.

4. Do not pin ancient versions from old tutorials. bitsandbytes==0.38 era pins are a common source of this error on machines that would work fine with the current release.

How to prevent it

Know which quantization stack belongs to which hardware: bitsandbytes for NVIDIA CUDA, GGUF/llama.cpp for CPU and Apple, plus GPTQ/AWQ files for pre-quantized CUDA inference. Choosing the stack that matches the machine avoids this whole category.