GPUs: Memory, Scheduling and Cost

When a CPU is enough

A GPU is not free, and not every model needs one. For a small model with modest traffic and a forgiving latency budget, the CPU you already have can be the right, cheaper answer.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A GPU is not free, and not every model needs one. For a small model with modest traffic, the CPU you already have can be the right answer.

The analogy you have already lived

You do not call a moving truck to carry one chair up one flight of stairs. A truck can certainly do it, faster in some sense than carrying it yourself. But hiring one for that single chair is real, unnecessary cost and complexity — for a job your own two hands were already enough for.

A GPU is the moving truck of computing. Genuinely large jobs need it. Plenty of jobs that reach for one anyway did not.

Why it exists

Every earlier lesson in this section assumed a GPU was the right tool. It worked out how to choose, share, schedule and pay for one well. This lesson asks the question that should come before all of them: does this job need a GPU at all?

A GPU adds real infrastructure. A machine to rent or own. Drivers and CUDA versions that have to match. Memory to size correctly. Cost per hour or per request to track. All of that is worth it when a model genuinely needs GPU speed. It is pure overhead when it does not.

How it works

   how BIG is the model, really?
              |
   how MUCH traffic does it actually need to handle?
              |
   how much LATENCY can this request tolerate?
              |
              v
   run it on the CPU you already have, and MEASURE the real numbers
              |
              v
   fast enough, at a traffic level the CPU can handle?
        |                                    |
       yes                                   no
        |                                    |
   stay on CPU.                      NOW a GPU is worth
   No GPU cost, no GPU               its real, added cost
   complexity, at all.               and complexity.

A real example you have seen

Spam filters, simple recommendation systems, and many classic machine learning models run inside ordinary web applications. They run on perfectly ordinary CPU servers, every day, at real scale. Nobody rents a GPU for them, because a GPU was never the bottleneck to begin with.

The honest part

It is tempting to assume "GPU" automatically means "fast," and reach for one out of habit rather than evidence. The only way to know for certain is to actually measure your specific model, on your specific CPU, at your specific traffic level. A real number beats an assumption every time, in either direction.

Remember this

  • A GPU is real infrastructure with real cost, not a free upgrade.
  • The right question is whether the CPU you already have is fast enough for your actual traffic and latency needs.
  • Measure your real model before assuming either answer. Do not guess.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Measuring CPU and GPU honestly, on the same small model

cpu_vs_gpu.py
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token

def bench(device, n=20, max_new_tokens=15):
    model = AutoModelForCausalLM.from_pretrained("distilgpt2").to(device)
    enc = tok("The quick brown fox", return_tensors="pt").to(device)
    with torch.no_grad():
        model.generate(**enc, max_new_tokens=max_new_tokens, do_sample=False, pad_token_id=tok.eos_token_id)  # warm-up
        if device == "cuda":
            torch.cuda.synchronize()
        start = time.perf_counter()
        for _ in range(n):
            model.generate(**enc, max_new_tokens=max_new_tokens, do_sample=False, pad_token_id=tok.eos_token_id)
        if device == "cuda":
            torch.cuda.synchronize()
        elapsed = time.perf_counter() - start
    return elapsed / n * 1000  # ms per single request

cpu_ms = bench("cpu")
gpu_ms = bench("cuda")

print("single small request, one at a time, 15 new tokens:")
print(f"  CPU: {cpu_ms:.1f} ms/request")
print(f"  GPU: {gpu_ms:.1f} ms/request")
print(f"  GPU is {cpu_ms/gpu_ms:.2f}x the CPU's speed for this one-at-a-time workload")
Output
single small request, one at a time, 15 new tokens:
  CPU: 258.1 ms/request
  GPU: 59.6 ms/request
  GPU is 4.33x the CPU's speed for this one-at-a-time workload

Real numbers, on this exact machine, for this small model. The GPU genuinely is faster here — 4.33x, measured, not assumed. The honest question this lesson is actually about is not "which is faster." It is "is 258 ms fast enough for what this request needs." For a background job with no one watching the clock, it comfortably is. For a live user waiting mid-conversation, it might not be. Same measurement, two different correct conclusions — depending entirely on the actual requirement.

Line-by-line walkthrough

bench("cpu") and bench("cuda") run the identical model, identical prompt, identical settings. The only thing that changes between the two calls is device. This is what makes the comparison honest — nothing about the model or workload was changed to favour either side.

The warm-up call before timing. Both CPU and GPU pay a one-time cost the first time a new shape of computation runs. Timing only after warm-up measures steady-state performance, not that one-time setup cost. See timing GPU code correctly for why this matters specifically for GPU code.

No batching in this comparison. This measures one request answered at a time, which is a realistic pattern for genuinely low, spread-out traffic. A workload that can batch many requests together shifts this comparison meaningfully in the GPU's favour. See when the GPU is not faster for where batch size changes the picture.

Common mistakes

Assuming GPU speed automatically translates into a better decision. The GPU being faster does not automatically mean it is the right choice. It means it is faster, at a real infrastructure cost the CPU option does not carry. Whether that speed is needed is a separate, situation-specific question.

Never actually measuring the CPU option. Many teams skip straight to provisioning a GPU. They never first check how the same model performs on hardware they already have. That measurement takes minutes and can save real, ongoing infrastructure cost.

Ignoring that CPUs handle concurrent requests differently. A CPU server can often serve several independent requests across its multiple cores at once. That changes the real throughput picture for concurrent traffic, in ways a single-request timing comparison, like the one above, does not capture on its own.

Forgetting model size changes everything. This comparison used a genuinely small model. A much larger model can flip this conclusion entirely, sometimes becoming impractically slow on CPU rather than only slower. Rerun this exact kind of comparison with your real model — never assume the small-model result transfers.

Try it yourself

Increase max_new_tokens from 15 to 100 and rerun. Watch whether the CPU-to-GPU speed ratio grows, shrinks, or stays about the same, as the amount of work per request increases.

What to learn next

Researcher — Mathematics and papers.

The real decision variables

Whether CPU inference is viable is a function of several independent factors, each worth checking explicitly rather than assumed together:

  • Model size and architecture. Small models with modest FLOPs per forward pass (classical ML, small transformers, distilled models) can run acceptably on CPU. Large transformers scale far worse on CPU. Raw compute is one reason; CPUs also lack the specialised matrix-multiply units, Tensor Cores, that give GPUs their advantage on the dense matrix operations transformers are built from.
  • Batch size and concurrency. GPUs derive much of their advantage from parallelism across a batch. At batch size 1, much of a GPU's theoretical advantage over a well-optimised, multi-core CPU implementation goes unused. The demo's result — 4.33x, not orders of magnitude — reflects exactly this regime.
  • Latency tolerance. A batch job running overnight cares about total throughput, not per-request latency, and can tolerate CPU speeds a live chat interface cannot.
  • Numeric precision and library optimisation. CPU inference can be substantially accelerated with quantisation (int8; see quantisation in practice), and with CPU-specific optimised runtimes like ONNX Runtime, Intel's OpenVINO, or torch.compile with a CPU backend. An unoptimised CPU baseline, as in the demo, understates what well-tuned CPU inference can achieve.

Where the crossover point actually sits

There is no universal crossover model size or batch size at which a GPU becomes strictly necessary. It depends on the specific architecture, the specific CPU's core count and instruction set — AVX-512 support materially changes CPU inference throughput — and the specific GPU being compared against. Published benchmarks comparing CPU and GPU inference (MLPerf Inference results are a useful, standardised reference point) consistently show this crossover moving, as both hardware generations and software optimisation improve. A conclusion reached two years ago about "needing" a GPU for a given model size is worth re-testing, not assumed to still hold.

The total cost of ownership argument

Beyond raw speed, CPU inference removes an entire category of operational cost this section's other lessons address directly. No CUDA/driver version matching (see CUDA, drivers and container images). No GPU scheduling complexity (see scheduling GPUs on Kubernetes). Access to a much larger, cheaper, more available pool of general-purpose compute. For a workload where CPU throughput is genuinely sufficient, this operational simplification is a real, ongoing saving independent of the raw per-request speed comparison.

Reading

  • MLPerf Inference benchmark results (mlcommons.org) — standardised, regularly updated CPU-versus-GPU inference comparisons across model families
  • Intel, OpenVINO Toolkit documentation — CPU-specific inference optimisation techniques referenced above

What to learn next