Error database

Killed while loading checkpoint shards (Hugging Face)

The operating system killed your process for exhausting system RAM during model load. Load with device_map="auto" so shards stream to the GPU, or quantize at load time.

The message you saw
Killed while loading checkpoint shards (Hugging Face)

By Updated

The error

Output
Loading checkpoint shards:  33%|███▎      | 2/6 [00:41<01:22, 20.65s/it]
Killed

No traceback, no exception — the process is gone. In a notebook, the kernel dies at the same spot instead.

What it means

Killed is not Python. It is the Linux OOM killer — the kernel mechanism that terminates the biggest memory consumer when system RAM runs out. Your model's weight files (the "shards") were being read into ordinary RAM, RAM ran out, and the kernel shot the process. This is a system RAM problem, not a GPU one; it happens before the weights ever reach the GPU.

Confirm it if you like:

bash
dmesg | grep -i "out of memory"

Why it happens

The naive loading path builds the full model in RAM first, then moves it to the GPU. A 7B-parameter model in float16 needs about 14 GB of RAM for that staging step; bigger models need proportionally more. Machines with 16 GB of RAM die partway through the shards — at 33%, at 67% — which makes it look flaky. It is not flaky. It is arithmetic.

How to fix it

1. Load with device_map="auto". This streams each shard directly to its destination instead of staging everything in RAM.

bash
pip install accelerate
python
from transformers import AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mistral-7B-Instruct-v0.3",
    dtype=torch.bfloat16,
    device_map="auto",
)

device_map="auto" places weights on GPU, spilling to CPU only when the GPU is full. Peak RAM drops to roughly one shard at a time.

2. Quantize at load to shrink everything. Quantization stores weights in fewer bits:

python
from transformers import BitsAndBytesConfig

model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mistral-7B-Instruct-v0.3",
    quantization_config=BitsAndBytesConfig(load_in_4bit=True),
    device_map="auto",
)

A 7B model lands around 4 GB in 4-bit — friendly to both RAM and GPU.

3. Pick a smaller model. A 3B or 1B model that loads beats a 7B model that dies. For learning and prototyping, small instruct models are underrated.

4. Add swap as a cushion, not a cure. A few GB of swap lets a borderline load squeak through, slowly. If the model only fits via heavy swapping, inference will crawl — prefer fixes 1-3.

5. On Colab or notebooks: same fixes, plus check the RAM gauge. The free tier's RAM dies on exactly this. "Your session crashed after using all available RAM" is this error in Colab clothing.

How to prevent it

Before loading any model, do the arithmetic: parameters × bytes per parameter, and compare against free RAM and GPU memory. Make device_map="auto" your default for anything above a couple of billion parameters. Close other RAM-heavy processes before loading — the OOM killer picks the biggest target, which is you.