GPU
In one sentence A GPU does thousands of simple calculations at once, which is exactly the shape of work neural networks need — hence its central role in AI.
Updated
A GPU (graphics processing unit) is a chip built to run thousands of simple calculations simultaneously — the workload deep learning happens to be made of.
The CPU-versus-GPU contrast is four maths professors versus five thousand schoolchildren. The professors (CPU cores) are brilliant, flexible, and handle any problem in sequence — perfect for running your operating system, where each step depends on the last. But multiplying two huge grids of numbers is not professor work; it is millions of independent small multiplications. Hand one each to the schoolchildren and the room finishes together, thousands of times faster than four professors ever could.
Neural networks are almost entirely that second kind of work: multiplying tensors, where every output element can be computed independently. GPUs were built to render game pixels — also millions of independent little calculations — and the discovery that the same silicon trains networks (AlexNet, 2012, two gaming cards) started the deep learning era. Modern AI GPUs added tensor cores, units specialised further for matrix multiplication in the low-precision formats training uses.
The numbers that matter when reading a spec sheet: FLOPS (raw arithmetic rate), VRAM capacity (decides what fits at all — usually the binding constraint for LLMs), and memory bandwidth (how fast weights stream to the compute — the true bottleneck for generation). Software-wise, NVIDIA's CUDA platform is why its chips dominate; PyTorch hides the details behind model.to("cuda"). No GPU? Rent by the hour (Colab's free tier included), and remember plenty of classical ML and small-model work runs fine on CPU.
Where to go next
- Full lesson: PyTorch basics
- Related terms: vram, tensor, flops, batching