GPUs: Memory, Scheduling and Cost
Choosing a GPU for inference
The right GPU for serving a model is decided mostly by memory — does it fit, with room left over for real traffic — not by which GPU is fastest on paper.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Choosing a GPU for serving a model is mostly a memory question. Will the model actually fit, with room left over? It is not mainly a question of which GPU is fastest.
The analogy you have already lived
You do not hire a truck to move one chair. You do not try to move a houseful of furniture with an auto-rickshaw either. The vehicle has to be the right size for the load. Big enough that it fits and the trip is not a struggle. Not so big that you pay for empty space you never use.
Picking a GPU works the same way. It has to be big enough to hold the model. Space must also remain for the work of actually answering requests. "Biggest and fastest" is not automatically "best." It is often only the most expensive answer to a question that did not need it.
Why it exists
Every GPU has a fixed amount of memory. If a model's weights, plus everything needed to answer a request, do not fit in that memory, the model will not run there. Not slowly, not badly — not at all.
Picking a GPU means answering one question honestly first. How much memory does this specific model, serving real traffic, actually need? Everything else — speed, cost, availability — only matters once that question has a real answer.
What actually needs to fit
The model's weights. The single largest, most predictable piece. A bigger model needs more memory, in a way you can calculate before ever renting anything.
Room to actually answer a question. The GPU needs working space beyond the weights alone. Memory used while a request is actually processed grows further, if you want to answer more than one question at a time.
A safety margin. Running a GPU right up to its very last byte of memory is fragile. A slightly bigger request than usual, and the whole thing can crash instead of only slowing down.
How it works
how much memory does the MODEL'S WEIGHTS need?
|
v
how much EXTRA memory does answering real requests need,
at the traffic level you actually expect?
|
v
add a safety margin
|
v
pick the smallest GPU whose memory comfortably covers all of thatA real example you have seen
Phone apps run a small AI model directly on your phone — a keyboard's next-word suggestion, a camera's face detection. They use tiny models on purpose. They are sized to fit a phone's much smaller memory. The same idea, at a much bigger scale, decides which GPU a company rents. It serves a much larger model to millions of people.
The honest part
The fastest, most expensive GPU available is rarely the right answer. Most of what you pay for on a top-tier GPU is raw speed you will not use, if your model is small. Or memory you will not fill, if your traffic is modest. The right GPU is the smallest one that comfortably does your actual job — figured out honestly, not guessed at.
Remember this
- The main question is does the model fit, with real room to spare — not which GPU is fastest.
- Memory needs come from the model's weights plus the work of answering real requests.
- The biggest GPU is rarely the right one. The right-sized one is.
What to learn next
- How much memory the KV cache eats — the part of the memory budget that grows fastest under real traffic.
- Reading GPU utilisation honestly — checking whether the GPU you chose is actually being used well.
- GPU cost per request — turning a memory-sized choice into a real cost number.
Developer — Code and libraries.
Setup
pip install torch transformersMeasuring real memory use, not guessing at it
import torch
from transformers import AutoModelForCausalLM
torch.cuda.empty_cache()
torch.cuda.reset_peak_memory_stats()
model = AutoModelForCausalLM.from_pretrained("distilgpt2", dtype=torch.float16)
model = model.to("cuda")
n_params = sum(p.numel() for p in model.parameters())
theoretical_mb = n_params * 2 / (1024 ** 2) # float16 = 2 bytes per parameter
actual_mb = torch.cuda.memory_allocated() / (1024 ** 2)
print(f"parameters: {n_params:,}")
print(f"theoretical weight size (params x 2 bytes): {theoretical_mb:.1f} MB")
print(f"actual GPU memory allocated after loading: {actual_mb:.1f} MB")
props = torch.cuda.get_device_properties(0)
print(f"\nGPU detected: {props.name}, {props.total_memory / (1024**3):.1f} GB total memory")parameters: 81,912,576 theoretical weight size (params x 2 bytes): 156.2 MB actual GPU memory allocated after loading: 158.9 MB GPU detected: NVIDIA RTX A6000, 45.0 GB total memory
This ran on the real GPU in this machine. Your numbers will differ if you run this on different hardware. The GPU name and total memory line will print whatever card you actually have. The close match between the theoretical and actual size (156.2 MB estimated, 158.9 MB measured) is the useful takeaway. A simple calculation — parameter count times bytes per parameter — gets you most of the way to a real memory estimate, before you rent anything.
Line-by-line walkthrough
dtype=torch.float16. Loading in 16-bit precision instead of the default 32-bit roughly halves memory use for the same model. See number formats for LLMs for what these formats actually are, and what changing them costs in accuracy.
n_params * 2. Two bytes per parameter, because each one is stored in float16. Swap in torch.float32 and the multiplier becomes 4; a quantised 8-bit format brings it to 1. This single number is the most reliable estimate you can make before ever loading anything.
The small gap between theoretical and actual memory. The extra roughly 2.7 MB comes from small buffers and framework bookkeeping that are not, strictly speaking, "weights." For a small model like this it is a rounding error. It is worth remembering it exists though, and stays proportionally smaller as models get larger.
Reference: real published GPU memory sizes
Real, publicly documented specifications, not live prices, which change constantly and vary by provider and region.
| GPU | VRAM |
|---|---|
| NVIDIA T4 | 16 GB |
| NVIDIA A10 | 24 GB |
| NVIDIA RTX 4090 | 24 GB |
| NVIDIA A100 | 40 GB or 80 GB |
| NVIDIA RTX A6000 | 48 GB |
| NVIDIA H100 | 80 GB |
Common mistakes
Sizing only for the model's weights, and forgetting request traffic needs memory too. A model that barely fits at rest can run out of memory the moment real, concurrent requests start arriving. Leave real headroom, not none.
Comparing GPUs by speed alone. A faster GPU that cannot fit your model is not a usable option at any price. Memory is a hard requirement; speed is a trade-off to optimise only after memory is settled.
Assuming bigger is always safer. An oversized GPU is not wrong, but it is often a wasted budget. Paying for capacity that never gets used is its own kind of mistake — covered in GPU cost per request.
Not testing with realistic traffic before committing. A GPU that comfortably serves one request at a time can behave very differently under ten concurrent ones. Test at the concurrency you actually expect, not at the easiest possible case.
Try it yourself
Repeat the measurement with torch.float32 instead of torch.float16 and compare the actual memory used. Confirm it roughly matches the doubled theoretical estimate.
What to learn next
- How much memory the KV cache eats — the part of the memory budget that grows fastest under real traffic.
- Reading GPU utilisation honestly — checking whether the GPU you chose is actually being used well.
- GPU cost per request — turning a memory-sized choice into a real cost number.
Researcher — Mathematics and papers.
A full inference memory budget
Total GPU memory required to serve a transformer model decomposes as
$$M_{\text{total}} = M_{\text{weights}} + M_{\text{kv-cache}} + M_{\text{activations}} + M_{\text{overhead}}$$
- $M_{\text{weights}} = N \cdot b$, where $N$ is parameter count and $b$ is bytes per parameter (4 for fp32, 2 for fp16/bf16, ~1 for int8, less for 4-bit formats).
- $M_{\text{kv-cache}}$ scales with batch size, sequence length, and the number of concurrent sequences being served — this is frequently the dominant term at real serving scale, not the weights, and is covered directly in how much memory the KV cache eats.
- $M_{\text{activations}}$ is the temporary memory used during a forward pass; for inference without gradient tracking, this is much smaller than during training and usually a minor term relative to the KV cache at nontrivial batch sizes.
- $M_{\text{overhead}}$ covers the CUDA context, framework allocator fragmentation, and any compiled kernel workspace — typically a few hundred MB to low GB, roughly constant regardless of model size.
Why weights alone underestimate real requirements
At small batch size and short sequence length, weights dominate and the simple estimate in the demo is close to sufficient. As concurrent traffic and context length grow, $M_{\text{kv-cache}}$ grows linearly with both and can exceed $M_{\text{weights}}$ by a wide margin for long-context, high-concurrency serving — a GPU sized purely off model weight size will appear to work in testing and then fail under real production load, precisely because testing rarely exercises full concurrency and long context simultaneously.
Compute versus memory-bound regimes
GPUs are characterised by both peak FLOPs and memory bandwidth; the ratio between them (arithmetic intensity required to be compute-bound) determines which resource actually limits a given workload. Autoregressive decoding at low batch size is almost always memory-bandwidth-bound, not compute-bound — the GPU spends more time moving weights and KV cache through memory than performing arithmetic on them. This is why decode throughput scales much more closely with a GPU's memory bandwidth (GB/s) than with its advertised peak FLOPs. Choosing a GPU purely by FLOPs benchmarks systematically misleads for LLM serving specifically — as opposed to training, where large-batch compute is closer to the limiting factor.
Reading
- NVIDIA, H100 Tensor Core GPU Architecture whitepaper — memory bandwidth and FLOPs figures for a concrete, current reference point
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
What to learn next
- How much memory the KV cache eats — the part of the memory budget that grows fastest under real traffic.
- Reading GPU utilisation honestly — checking whether the GPU you chose is actually being used well.
- GPU cost per request — turning a memory-sized choice into a real cost number.