Batching and Concurrency

Concurrency in a Python inference server

Python threads can wait for many things at once, but only one thread runs Python code at any instant — so a CPU-heavy model needs separate processes, not more threads, to use more than one core.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Python threads take turns on one CPU core. Only separate processes can truly use more than one core at the same time.

The analogy you have already lived

Picture one kitchen with one shared chopping counter and one knife. You can invite four cooks into that kitchen. They can all be "working" — but only one of them can actually be holding the knife and cutting at any instant. The others wait their turn, even though they are all present and eager.

Now picture four separate kitchens, each with its own counter and its own knife. Four cooks, each in their own kitchen, really can chop at the same moment.

Python threads are the first kitchen. Python processes are the second.

Why it exists

Python has something called the GIL — the Global Interpreter Lock. It is a rule inside the Python interpreter: only one thread may run Python bytecode at any instant. That holds no matter how many CPU cores the machine has.

This sounds like a flaw, and people argue about it endlessly. It exists because Python's memory management is not built to be safely touched by two threads at once without it. The GIL is the lock that keeps that safe.

The GIL does not block waiting. A thread that is waiting for a network reply, a disk read, or a database query releases the GIL while it waits, letting another thread run. It only blocks two threads from doing CPU work — actual Python computation — at the same instant.

How it works

CPU-bound work (real computation, e.g. running a model in pure Python):
   thread 1: ##----##----##----##----   <- takes turns holding the one "knife"
   thread 2: --##----##----##----##--
   (adding more threads does not add more knives)

I/O-bound work (waiting on a network call, a database, a disk):
   thread 1: ##....................##  <- "." is waiting, GIL is released
   thread 2: ....##................    <- another thread can run during that wait
   (adding more threads DOES help here)

A real example you have seen

An app that feels instant while doing several things at once — a spinner, a profile picture, a payment check — is usually juggling many small waits. It is not running many heavy computations. That is exactly the case where Python's threads shine.

The honest part

This trips up almost every beginner writing their first Python server. A CPU-heavy model wrapped in "more threads" often shows zero speedup. That is confusing the first time you see it, because everything about threads suggests it should help. Read the developer section below twice if the first pass does not stick — this genuinely surprises people who already know how to code.

Remember this

  • Python's GIL lets only one thread run Python computation at a time, on any one process.
  • Threads help when your server spends time waiting — for the network, disk, or another service.
  • Processes are needed when your server spends time doing real CPU computation, because each process gets its own GIL.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]"

Nothing else is needed — this uses only Python's built-in concurrent.futures.

Proving it to yourself

concurrency.py
import time
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor


def cpu_bound_predict(n: int) -> float:
    """Stands in for a model doing real arithmetic in pure Python --
    no sleeping, this genuinely keeps the CPU busy the whole time."""
    total = 0.0
    for i in range(n):
        total += (i * i) % 97
    return total


def run(executor_cls, workers, n_tasks, work_size):
    start = time.perf_counter()
    if executor_cls is None:
        for _ in range(n_tasks):
            cpu_bound_predict(work_size)
    else:
        with executor_cls(max_workers=workers) as ex:
            list(ex.map(cpu_bound_predict, [work_size] * n_tasks))
    return time.perf_counter() - start


if __name__ == "__main__":
    N_TASKS = 8
    WORK_SIZE = 2_000_000
    WORKERS = 4

    sequential = run(None, 1, N_TASKS, WORK_SIZE)
    threaded = run(ThreadPoolExecutor, WORKERS, N_TASKS, WORK_SIZE)
    multiproc = run(ProcessPoolExecutor, WORKERS, N_TASKS, WORK_SIZE)

    print(f"sequential:              {sequential:6.2f}s")
    print(f"threads,   {WORKERS} workers: {threaded:6.2f}s  (speedup {sequential/threaded:.2f}x)")
    print(f"processes, {WORKERS} workers: {multiproc:6.2f}s  (speedup {sequential/multiproc:.2f}x)")
Output
sequential:                0.75s
threads,   4 workers:      0.75s  (speedup 0.99x)
processes, 4 workers:      0.32s  (speedup 2.34x)

These are real measurements from this exact run, on this exact machine. Both the absolute times and the speedup ratios will differ on your machine — core count, clock speed and what else is running all move the numbers. What should hold on any multi-core machine is the pattern: threads give almost no speedup on CPU-bound work, processes give a real one.

Line-by-line walkthrough

cpu_bound_predict does a real Python loop with real arithmetic — no time.sleep, because sleeping releases the GIL and would hide the exact effect this demo is meant to show.

ThreadPoolExecutor gives every task its own thread, inside the same process, sharing one GIL — so they take turns, exactly like the one-knife kitchen.

ProcessPoolExecutor gives every task its own operating-system process, each with its own Python interpreter and own GIL. These are genuinely separate kitchens, at the cost of a slower startup and no shared memory by default.

Common mistakes

"More workers" as a reflex fix for a slow model. If the model itself is CPU-bound Python code, adding thread workers changes nothing. Reach for ProcessPoolExecutor, or move the hot path into a library that releases the GIL internally (NumPy and most C-extension-backed libraries do, during their C code).

Running uvicorn with one worker for a CPU-heavy model. uvicorn --workers 4 starts four separate processes, each with its own model copy in memory. That is the process-based fix applied at the server level, not the thread-pool level — and it is usually the simpler place to apply it.

Forgetting each process needs its own copy of the model. Four worker processes means four copies of your model in RAM. For a large model, that memory cost is real and needs planning, unlike threads which share memory for free.

Blaming Python for a fundamentally CPU-bound workload. If a model's forward pass genuinely needs raw compute, no amount of Python-level concurrency trickery fixes that. Buy more cores, use a GPU, or make the model cheaper instead — see batching and quantised inference for that side of the problem.

Try it yourself

Change cpu_bound_predict to instead call time.sleep(0.05) — simulating a slow network call to another service, not real computation. Rerun and compare: threads should now show a real speedup, because there is no CPU work for them to fight over.

What to learn next

Researcher — Mathematics and papers.

Where the GIL actually blocks you

The GIL is released around blocking system calls (I/O, time.sleep) and, in CPython's implementation, is preemptively yielded roughly every 5 milliseconds of continuous bytecode execution (configurable via sys.setswitchinterval), so long CPU-bound loops in one thread do not starve the others of scheduling — they still starve them of progress, because whichever thread holds the GIL at that instant is the only one making forward progress on Python code.

C-extension code — NumPy's array operations, most of PyTorch's tensor operations, hashlib, zlib — can explicitly release the GIL while doing the actual C or CUDA work, then reacquire it before returning to Python. This is why NumPy-heavy code can show real thread-level speedup despite the GIL: the parallel part never runs interpreted Python bytecode.

asyncio sidesteps the GIL question differently: it runs a single thread, and cooperative coroutines yield control explicitly at await points. It is excellent for I/O-bound fan-out (many concurrent network calls) and offers no help at all for CPU-bound work, since there is still only one thread doing the computing.

The subinterpreters and no-GIL efforts

PEP 684 (Python 3.12) introduced per-interpreter GILs, letting one process run multiple interpreters each with independent locks — a step toward real intra-process parallelism, though the ecosystem's C-extension compatibility is still catching up as of this writing.

PEP 703 defines an optional free-threaded CPython build (3.13+) that removes the GIL entirely, at a measured single-threaded performance cost reported by the CPython team in the roughly 5–15% range depending on workload, in exchange for real multi-core scaling within one process. It ships as an alternate build, not the default, and most third-party C extensions require updates to support it correctly.

Practical implication for serving

The standard production pattern remains: N worker processes (via uvicorn --workers N, Gunicorn, or a process manager) matched roughly to physical core count, each internally using asyncio or a thread pool to overlap I/O-bound work — network calls to a feature store, a database, another microservice — around whatever CPU-bound model calls it makes. This gets process-level parallelism for compute and thread- or coroutine-level concurrency for waiting, without needing a no-GIL build.

References

  • PEP 703, Making the Global Interpreter Lock Optional in CPython, 2023 — peps.python.org/pep-0703
  • PEP 684, A Per-Interpreter GIL, 2022 — peps.python.org/pep-0684
  • Beazley, D., Understanding the Python GIL, PyCon 2010 — the original deep-dive talk that most later explanations, including this one, build on.

What to learn next