AI glossary

VRAM

In one sentence VRAM is the GPU's own on-board memory — the hard limit deciding which models you can run and train at all.

By Updated

VRAM (video random-access memory) is the memory physically attached to a GPU — and whatever you are computing must fit inside it.

Think of a workshop's workbench. The warehouse next door (your computer's ordinary RAM and disk) can store any amount of material, but work happens only on the bench, and the bench has fixed dimensions. A cupboard too large for the bench cannot be built on it — not slowly, not cleverly, not at all. VRAM is the bench: when a model does not fit, the job does not run, and you meet deep learning's most famous error message: CUDA out of memory.

Budgeting it is simple arithmetic worth doing before every download. Model weights: parameters × bytes per parameter — a 7B model needs ~14 GB at 16-bit, ~4 GB at 4-bit quantization. On top of that, inference adds the growing KV cache (long chats, many users — gigabytes). Training adds far more: gradients plus optimizer state can quadruple the weights' footprint, which is why a card that runs a model often cannot train it.

consumer cards : 8-24 GB   (RTX 4060 → 4090)
workstation    : 48 GB
datacentre     : 80-192 GB (A100, H100, B200)

When you hit the wall, the escape ladder is standard: quantize, shrink the batch-size, use LoRA instead of full fine-tuning, use gradient checkpointing (recompute activations instead of storing them), offload layers to CPU (slow), or split the model across several GPUs. Apple-silicon Macs blur the picture with unified memory shared between CPU and GPU — the reason 64 GB MacBooks became popular local-LLM machines.

Where to go next