GPUs: Memory, Scheduling and Cost

GPU cost per request

The true cost of one request is the GPU's price divided by how many requests it actually serves — and that number depends far more on utilisation than on the GPU's hourly price.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

The real cost of one request is the GPU's price, divided by how many requests it actually served. Utilisation matters more to that number than the GPU's price does.

The analogy you have already lived

An auto-rickshaw driver's real cost per ride is not just the fuel for that one trip. It includes the fuel, the vehicle's upkeep, and every minute spent waiting for the next passenger, with the meter not running. A driver who sits idle most of the day earns far less per ride than one whose auto is constantly full. Even if both pay exactly the same for fuel and maintenance.

A GPU serving requests works the same way. Its price per hour is only half the story. How many requests it actually answers in that hour is the other, often bigger, half.

Why it exists

It is tempting to look at a GPU's hourly rental price and treat that as "the cost." It is not. It is the cost of renting the GPU, not the cost of answering one request.

Two GPUs at the exact same hourly price can have wildly different costs per request. One might be busy the whole time; the other might sit mostly idle, waiting for requests that rarely arrive.

How it works

   GPU price per hour
              |
              divided by
              |
              v
   requests actually served in that hour

   =  real cost per request

The top number — price — is fixed once you have chosen a GPU. The bottom number — requests actually served — depends on how busy the GPU genuinely is. That is a separate question from how fast it is capable of going.

A real example you have seen

A half-empty flight still costs the airline almost the same to fly as a full one. Same fuel, same crew, same plane. The airline's real cost per passenger on that half-empty flight is roughly double what it is on a full one. A GPU serving requests below its real capacity is in exactly that position.

The honest part

A GPU's maximum throughput is how many requests it could serve if it were always busy. That is often very different from how many requests it actually serves in real traffic, which has quiet periods. Using the maximum number to estimate cost per request paints a rosier picture than reality. Be honest about which number you are actually using.

Remember this

  • Cost per request is price divided by requests actually served, not requests theoretically possible.
  • Utilisation — how busy the GPU genuinely is — often matters more than the GPU's price tag.
  • A GPU's maximum throughput and its real, measured throughput under actual traffic are different numbers. Use the real one.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Measuring real throughput, then real cost per request

cost_per_request.py
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained("distilgpt2", dtype=torch.float16).to("cuda")
enc = tok("Tell me about the weather today in", return_tensors="pt").to("cuda")

with torch.no_grad():  # warm-up
    model.generate(**enc, max_new_tokens=20, do_sample=False, pad_token_id=tok.eos_token_id)
torch.cuda.synchronize()

N_REQUESTS = 30
start = time.perf_counter()
with torch.no_grad():
    for _ in range(N_REQUESTS):
        model.generate(**enc, max_new_tokens=20, do_sample=False, pad_token_id=tok.eos_token_id)
torch.cuda.synchronize()
elapsed = time.perf_counter() - start

requests_per_hour = (N_REQUESTS / elapsed) * 3600
print(f"measured: {N_REQUESTS} requests in {elapsed:.2f}s -> {N_REQUESTS/elapsed:.2f} requests/sec")
print(f"at full utilisation: {requests_per_hour:,.0f} requests/hour")

GPU_PRICE_PER_HOUR = 1.50  # illustrative example rate, not a live price
print(f"\nGPU price: ${GPU_PRICE_PER_HOUR:.2f}/hour\n")
for utilisation in [1.0, 0.5, 0.1, 0.02]:
    effective = requests_per_hour * utilisation
    cost = GPU_PRICE_PER_HOUR / effective
    print(f"utilisation {utilisation*100:5.0f}%  ->  {effective:12,.0f} req/hr served  ->  ${cost:.6f} per request")
Output
measured: 30 requests in 2.28s -> 13.18 requests/sec
at full utilisation: 47,443 requests/hour

GPU price: $1.50/hour

utilisation   100%  ->        47,443 req/hr served  ->  $0.000032 per request
utilisation    50%  ->        23,721 req/hr served  ->  $0.000063 per request
utilisation    10%  ->         4,744 req/hr served  ->  $0.000316 per request
utilisation     2%  ->           949 req/hr served  ->  $0.001581 per request

The throughput number, 13.18 requests/second, is a real, measured figure from this exact machine, model and generation length. Yours will differ with a different model, GPU or output length. The pattern underneath it is the whole point. With the identical GPU at the identical price, real cost per request rose nearly 50x between full utilisation and 2% utilisation. The GPU's price never changed. Only how busy it was did.

Line-by-line walkthrough

N_REQUESTS = 30, run back to back with no gaps. This measures maximum throughput — the best case, if the GPU were never idle. It is a ceiling, not a promise about real traffic.

The utilisation loop. Multiplying the maximum throughput by a utilisation fraction estimates how many requests are actually served in real conditions. The GPU is not always busy. This is the honest number that should drive a cost decision, not the ceiling above it.

Why low utilisation is common, not rare. A GPU sized for a traffic peak sits well below its maximum throughput most of the day, by design. See autoscaling for the standard fix, and reading GPU utilisation honestly for measuring your real number, rather than assuming one.

Common mistakes

Quoting cost per request using maximum throughput. This systematically understates real cost, often by a wide margin. Real traffic is rarely as steady and constant as a back-to-back benchmark loop.

Ignoring that batching changes throughput, not just latency. Serving several requests together (see dynamic batching) can substantially raise effective throughput per hour, directly lowering cost per request. Measure throughput at realistic batch sizes, not only one request at a time as in this simplified demo.

Comparing two GPUs by price alone. A more expensive GPU with much higher real throughput can have a lower cost per request than a cheaper one running at low utilisation. Always compare the final cost-per-request number, not the sticker price.

Not remeasuring after a model or traffic change. A new model version, a longer typical response, or a change in traffic pattern all change real throughput. And therefore real cost per request. Treat this as a number to track over time, not calculate once.

Try it yourself

Change max_new_tokens from 20 to 100 and rerun. Compare how much the requests-per-second figure drops, and what that does to cost per request at the same GPU price.

What to learn next

Researcher — Mathematics and papers.

Cost per request as a function of arrival process, not just service rate

Let $\mu$ be the GPU's maximum service rate (requests/hour, measured as in the demo) and $\lambda$ the real arrival rate of requests. Utilisation is $\rho = \lambda / \mu$. Real cost per request is

$$C_{\text{request}} = \frac{P_{\text{hour}}}{\lambda}= \frac{P_{\text{hour}}}{\rho \cdot \mu}$$

This makes explicit what the demo shows empirically: cost per request is inversely proportional to $\rho$, holding price and maximum throughput fixed. Critically, $\rho$ is a property of traffic, not of the GPU — the same hardware serving a steadier arrival process achieves a materially lower real cost per request than one serving the identical average volume in a bursty pattern, because bursty traffic forces provisioning for the peak while paying for capacity during every quiet gap between bursts.

The provisioning dilemma

Choosing $\mu$ (via GPU count and type) to keep $\rho$ high enough for good unit economics, while keeping queueing delay acceptable, is a direct application of the same queueing result discussed in model serving: as $\rho \to 1$, expected queueing delay grows without bound ($\propto \rho/(1-\rho)$ for an M/M/1-type system). This means the cheapest cost-per-request configuration (very high $\rho$) is generally in direct tension with acceptable tail latency, and the two must be optimised jointly, not independently — a pure cost-per-request minimisation with no latency constraint will recommend running the GPU dangerously close to saturation.

Batching's effect on the underlying rate

Dynamic and continuous batching (see dynamic batching and continuous batching for LLMs) raise $\mu$ itself, rather than changing $\rho$ for a fixed arrival pattern — the GPU processes more requests per unit time by exploiting parallelism across requests that were arriving anyway. This is why batching is frequently the single highest-leverage lever for cost per request specifically: it improves the numerator's denominator ($\mu$) directly, compounding multiplicatively with any utilisation improvement achieved separately through better traffic shaping or autoscaling.

Amortising fixed overhead correctly

For a full cost-per-request figure, $P_{\text{hour}}$ itself should include more than the raw GPU rental rate — networking, storage for weights and logs, and orchestration overhead (see the one-box production stack) all belong in the numerator for an honest number, echoing the same completeness argument made in unit economics of an AI feature.

What to learn next