FLOPS
In one sentence FLOPS measure how many arithmetic operations per second hardware can perform — the standard currency for compute in AI.
Updated
FLOPS are floating-point operations per second — the count of additions and multiplications on decimal numbers a machine performs each second.
A floating-point operation is one piece of arithmetic on a decimal number: one multiply, one add. Counting them is like measuring a kitchen in "chopping strokes per minute" — crude, ignoring skill and recipe, yet genuinely useful for comparing kitchens, because in the end the vegetables must all get chopped. Deep learning is billions of such operations arranged in tensor multiplications, so raw arithmetic rate is a meaningful axis.
The scales involved take a moment to absorb:
1 GFLOPS = 10⁹ ops/sec late-90s desktop
1 TFLOPS = 10¹² ops/sec a phone chip today
modern AI GPU ~10³-10⁴ TFLOPS at low precision
frontier training run ~10²⁵-10²⁶ total operations, months of thousands of GPUsTwo usages hide under one acronym, and context disambiguates. FLOPS-per-second rates hardware (spec sheets, with the fine print that low-precision numbers — FP16, FP8, INT8 — inflate the figure several-fold versus FP64). Total FLOPs (lowercase s, a count) measures work done: training cost ≈ 6 × parameters × training tokens is the standard estimate, and it is the compute axis of scaling-laws. Regulators have even used total training FLOPs as a threshold for oversight — an EU AI Act criterion sits at 10²⁵.
The honest caveat: peak FLOPS are a ceiling, not a promise. Real workloads often stall on memory bandwidth — data cannot reach the arithmetic units fast enough — which is why utilisation ("we achieved 40% of peak") and memory specs matter as much as the headline number.
Where to go next
- Full lesson: vLLM
- Related terms: gpu, scaling-laws, parameter, throughput