Killed while loading checkpoint shards (Hugging Face)
The operating system killed your process for exhausting system RAM during model load. Load with device_map="auto" so shards stream to the GPU, or quantize at load time.
Updated
The error
Loading checkpoint shards: 33%|███▎ | 2/6 [00:41<01:22, 20.65s/it] Killed
No traceback, no exception — the process is gone. In a notebook, the kernel dies at the same spot instead.
What it means
Killed is not Python. It is the Linux OOM killer — the kernel mechanism that terminates the biggest memory consumer when system RAM runs out. Your model's weight files (the "shards") were being read into ordinary RAM, RAM ran out, and the kernel shot the process. This is a system RAM problem, not a GPU one; it happens before the weights ever reach the GPU.
Confirm it if you like:
dmesg | grep -i "out of memory"Why it happens
The naive loading path builds the full model in RAM first, then moves it to the GPU. A 7B-parameter model in float16 needs about 14 GB of RAM for that staging step; bigger models need proportionally more. Machines with 16 GB of RAM die partway through the shards — at 33%, at 67% — which makes it look flaky. It is not flaky. It is arithmetic.
How to fix it
1. Load with device_map="auto". This streams each shard directly to its destination instead of staging everything in RAM.
pip install acceleratefrom transformers import AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-Instruct-v0.3",
dtype=torch.bfloat16,
device_map="auto",
)device_map="auto" places weights on GPU, spilling to CPU only when the GPU is full. Peak RAM drops to roughly one shard at a time.
2. Quantize at load to shrink everything. Quantization stores weights in fewer bits:
from transformers import BitsAndBytesConfig
model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-Instruct-v0.3",
quantization_config=BitsAndBytesConfig(load_in_4bit=True),
device_map="auto",
)A 7B model lands around 4 GB in 4-bit — friendly to both RAM and GPU.
3. Pick a smaller model. A 3B or 1B model that loads beats a 7B model that dies. For learning and prototyping, small instruct models are underrated.
4. Add swap as a cushion, not a cure. A few GB of swap lets a borderline load squeak through, slowly. If the model only fits via heavy swapping, inference will crawl — prefer fixes 1-3.
5. On Colab or notebooks: same fixes, plus check the RAM gauge. The free tier's RAM dies on exactly this. "Your session crashed after using all available RAM" is this error in Colab clothing.
How to prevent it
Before loading any model, do the arithmetic: parameters × bytes per parameter, and compare against free RAM and GPU memory. Make device_map="auto" your default for anything above a couple of billion parameters. Close other RAM-heavy processes before loading — the OOM killer picks the biggest target, which is you.
Related errors
- CUDA out of memory — the GPU-side sibling, with the deeper memory-budget discussion
- MemoryError: Unable to allocate N GiB
- The kernel appears to have died — the same event seen from Jupyter