GPUs: Memory, Scheduling and Cost
GPU cost per request
The true cost of one request is the GPU's price divided by how many requests it actually serves — and that number depends far more on utilisation than on the GPU's hourly price.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The real cost of one request is the GPU's price, divided by how many requests it actually served. Utilisation matters more to that number than the GPU's price does.
The analogy you have already lived
An auto-rickshaw driver's real cost per ride is not just the fuel for that one trip. It includes the fuel, the vehicle's upkeep, and every minute spent waiting for the next passenger, with the meter not running. A driver who sits idle most of the day earns far less per ride than one whose auto is constantly full. Even if both pay exactly the same for fuel and maintenance.
A GPU serving requests works the same way. Its price per hour is only half the story. How many requests it actually answers in that hour is the other, often bigger, half.
Why it exists
It is tempting to look at a GPU's hourly rental price and treat that as "the cost." It is not. It is the cost of renting the GPU, not the cost of answering one request.
Two GPUs at the exact same hourly price can have wildly different costs per request. One might be busy the whole time; the other might sit mostly idle, waiting for requests that rarely arrive.
How it works
GPU price per hour
|
divided by
|
v
requests actually served in that hour
= real cost per requestThe top number — price — is fixed once you have chosen a GPU. The bottom number — requests actually served — depends on how busy the GPU genuinely is. That is a separate question from how fast it is capable of going.
A real example you have seen
A half-empty flight still costs the airline almost the same to fly as a full one. Same fuel, same crew, same plane. The airline's real cost per passenger on that half-empty flight is roughly double what it is on a full one. A GPU serving requests below its real capacity is in exactly that position.
The honest part
A GPU's maximum throughput is how many requests it could serve if it were always busy. That is often very different from how many requests it actually serves in real traffic, which has quiet periods. Using the maximum number to estimate cost per request paints a rosier picture than reality. Be honest about which number you are actually using.
Remember this
- Cost per request is price divided by requests actually served, not requests theoretically possible.
- Utilisation — how busy the GPU genuinely is — often matters more than the GPU's price tag.
- A GPU's maximum throughput and its real, measured throughput under actual traffic are different numbers. Use the real one.
What to learn next
- When a CPU is enough — the question to ask before this whole calculation, for a genuinely small workload.
- Dynamic batching — the most direct lever for raising the throughput this lesson's cost figure depends on.
- Reading GPU utilisation honestly — measuring the real utilisation number this whole calculation runs on.
Developer — Code and libraries.
Setup
pip install torch transformersMeasuring real throughput, then real cost per request
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained("distilgpt2", dtype=torch.float16).to("cuda")
enc = tok("Tell me about the weather today in", return_tensors="pt").to("cuda")
with torch.no_grad(): # warm-up
model.generate(**enc, max_new_tokens=20, do_sample=False, pad_token_id=tok.eos_token_id)
torch.cuda.synchronize()
N_REQUESTS = 30
start = time.perf_counter()
with torch.no_grad():
for _ in range(N_REQUESTS):
model.generate(**enc, max_new_tokens=20, do_sample=False, pad_token_id=tok.eos_token_id)
torch.cuda.synchronize()
elapsed = time.perf_counter() - start
requests_per_hour = (N_REQUESTS / elapsed) * 3600
print(f"measured: {N_REQUESTS} requests in {elapsed:.2f}s -> {N_REQUESTS/elapsed:.2f} requests/sec")
print(f"at full utilisation: {requests_per_hour:,.0f} requests/hour")
GPU_PRICE_PER_HOUR = 1.50 # illustrative example rate, not a live price
print(f"\nGPU price: ${GPU_PRICE_PER_HOUR:.2f}/hour\n")
for utilisation in [1.0, 0.5, 0.1, 0.02]:
effective = requests_per_hour * utilisation
cost = GPU_PRICE_PER_HOUR / effective
print(f"utilisation {utilisation*100:5.0f}% -> {effective:12,.0f} req/hr served -> ${cost:.6f} per request")measured: 30 requests in 2.28s -> 13.18 requests/sec at full utilisation: 47,443 requests/hour GPU price: $1.50/hour utilisation 100% -> 47,443 req/hr served -> $0.000032 per request utilisation 50% -> 23,721 req/hr served -> $0.000063 per request utilisation 10% -> 4,744 req/hr served -> $0.000316 per request utilisation 2% -> 949 req/hr served -> $0.001581 per request
The throughput number, 13.18 requests/second, is a real, measured figure from this exact machine, model and generation length. Yours will differ with a different model, GPU or output length. The pattern underneath it is the whole point. With the identical GPU at the identical price, real cost per request rose nearly 50x between full utilisation and 2% utilisation. The GPU's price never changed. Only how busy it was did.
Line-by-line walkthrough
N_REQUESTS = 30, run back to back with no gaps. This measures maximum throughput — the best case, if the GPU were never idle. It is a ceiling, not a promise about real traffic.
The utilisation loop. Multiplying the maximum throughput by a utilisation fraction estimates how many requests are actually served in real conditions. The GPU is not always busy. This is the honest number that should drive a cost decision, not the ceiling above it.
Why low utilisation is common, not rare. A GPU sized for a traffic peak sits well below its maximum throughput most of the day, by design. See autoscaling for the standard fix, and reading GPU utilisation honestly for measuring your real number, rather than assuming one.
Common mistakes
Quoting cost per request using maximum throughput. This systematically understates real cost, often by a wide margin. Real traffic is rarely as steady and constant as a back-to-back benchmark loop.
Ignoring that batching changes throughput, not just latency. Serving several requests together (see dynamic batching) can substantially raise effective throughput per hour, directly lowering cost per request. Measure throughput at realistic batch sizes, not only one request at a time as in this simplified demo.
Comparing two GPUs by price alone. A more expensive GPU with much higher real throughput can have a lower cost per request than a cheaper one running at low utilisation. Always compare the final cost-per-request number, not the sticker price.
Not remeasuring after a model or traffic change. A new model version, a longer typical response, or a change in traffic pattern all change real throughput. And therefore real cost per request. Treat this as a number to track over time, not calculate once.
Try it yourself
Change max_new_tokens from 20 to 100 and rerun. Compare how much the requests-per-second figure drops, and what that does to cost per request at the same GPU price.
What to learn next
- When a CPU is enough — the question to ask before this whole calculation, for a genuinely small workload.
- Dynamic batching — the most direct lever for raising the throughput this lesson's cost figure depends on.
- Reading GPU utilisation honestly — measuring the real utilisation number this whole calculation runs on.
Researcher — Mathematics and papers.
Cost per request as a function of arrival process, not just service rate
Let $\mu$ be the GPU's maximum service rate (requests/hour, measured as in the demo) and $\lambda$ the real arrival rate of requests. Utilisation is $\rho = \lambda / \mu$. Real cost per request is
$$C_{\text{request}} = \frac{P_{\text{hour}}}{\lambda}= \frac{P_{\text{hour}}}{\rho \cdot \mu}$$
This makes explicit what the demo shows empirically: cost per request is inversely proportional to $\rho$, holding price and maximum throughput fixed. Critically, $\rho$ is a property of traffic, not of the GPU — the same hardware serving a steadier arrival process achieves a materially lower real cost per request than one serving the identical average volume in a bursty pattern, because bursty traffic forces provisioning for the peak while paying for capacity during every quiet gap between bursts.
The provisioning dilemma
Choosing $\mu$ (via GPU count and type) to keep $\rho$ high enough for good unit economics, while keeping queueing delay acceptable, is a direct application of the same queueing result discussed in model serving: as $\rho \to 1$, expected queueing delay grows without bound ($\propto \rho/(1-\rho)$ for an M/M/1-type system). This means the cheapest cost-per-request configuration (very high $\rho$) is generally in direct tension with acceptable tail latency, and the two must be optimised jointly, not independently — a pure cost-per-request minimisation with no latency constraint will recommend running the GPU dangerously close to saturation.
Batching's effect on the underlying rate
Dynamic and continuous batching (see dynamic batching and continuous batching for LLMs) raise $\mu$ itself, rather than changing $\rho$ for a fixed arrival pattern — the GPU processes more requests per unit time by exploiting parallelism across requests that were arriving anyway. This is why batching is frequently the single highest-leverage lever for cost per request specifically: it improves the numerator's denominator ($\mu$) directly, compounding multiplicatively with any utilisation improvement achieved separately through better traffic shaping or autoscaling.
Amortising fixed overhead correctly
For a full cost-per-request figure, $P_{\text{hour}}$ itself should include more than the raw GPU rental rate — networking, storage for weights and logs, and orchestration overhead (see the one-box production stack) all belong in the numerator for an honest number, echoing the same completeness argument made in unit economics of an AI feature.
What to learn next
- When a CPU is enough — the question to ask before this whole calculation, for a genuinely small workload.
- Dynamic batching — the most direct lever for raising the throughput this lesson's cost figure depends on.
- Reading GPU utilisation honestly — measuring the real utilisation number this whole calculation runs on.