Latency, Load Testing and Capacity

Tuning CPU inference

A math library like NumPy already uses several CPU threads per call, so running many of your own workers on top of it can oversubscribe the machine and make things slower, not faster.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A math library already uses several CPU threads per call, so adding more workers on top of it can make things slower, not faster.

The analogy you have already lived

You have cooked in a small kitchen with one counter. One cook, using the whole counter, works efficiently. Squeeze in three more cooks, each trying to use the same counter at once. Everyone starts bumping elbows, waiting for space, and getting in each other's way. More cooks did not mean more cooking — past a point, it meant less.

A CPU has the same limited counter: its cores. A busy math library can already fill that counter on its own.

Why it exists

Libraries like NumPy do not run on one CPU core alone. Underneath, a matrix multiplication is handed to a BLAS library — highly optimised code for exactly this kind of maths. That BLAS library often spreads the work across many CPU threads by itself, without you asking it to.

This is normally a good thing. The trouble starts when your own code also adds parallelism on top. That means running several model calls at once, each one internally trying to grab several CPU threads for its own matrix maths. On a machine with a limited number of cores, this oversubscribes the CPU. More threads are asking for cores than the machine actually has. They spend real time fighting over that limited counter space, instead of working.

How it works

A small container with 2 CPUs:

Untuned (BLAS grabs many threads per call, AND you run 8 calls at once):
   call 1: #### #### #### ####   <- wants a big chunk of the 2 CPUs
   call 2: #### #### #### ####   <- wants the SAME 2 CPUs, at the same time
   ...(6 more, all fighting for the same 2 CPUs)
   result: everyone waits, switches, waits again -- wasted time

Tuned (each call uses only 1 thread; let YOUR workers provide the concurrency):
   call 1: #     call 2: #     call 3: #     call 4: #     (2 CPUs, used cleanly)
   result: no fighting over the same cores

A real example you have seen

A phone that gets noticeably slower, not faster, right after you open five heavy apps at once is the same idea. Too many things are competing for the same limited processor, and each one gets less than it needs.

The honest part

This effect is easy to miss on a big development machine with dozens of cores, where there is enough room for everyone to avoid fighting. It shows up hardest on the machine sizes most real services actually run on: a container with two or four CPUs. That is exactly why it catches people by surprise the first time they deploy.

Remember this

  • A math library can use several CPU threads per call, on its own, without asking.
  • Running many of your own workers on top of that can oversubscribe the CPU and slow things down.
  • The fix is deciding, deliberately, who provides the parallelism — the library, or your own worker pool — not both at once.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy psutil threadpoolctl

Reproducing the oversubscription, on a machine pinned to 2 CPUs

Most inference containers get 2–4 vCPUs, not a workstation's full core count. This script pins itself to 2 CPUs with psutil so the demo reflects that reality, rather than the sandbox machine's actual core count.

cpu_tuning.py
import time
from concurrent.futures import ThreadPoolExecutor

import numpy as np
import psutil
from threadpoolctl import threadpool_limits

psutil.Process().cpu_affinity([0, 1])   # pin this process to 2 CPUs

SIZE = 300
N_CALLS = 80
A = np.random.rand(SIZE, SIZE)
B = np.random.rand(SIZE, SIZE)


def one_matmul(_):
    return (A @ B).sum()


def run_concurrent(n_python_workers, blas_threads):
    with threadpool_limits(limits=blas_threads, user_api="blas"):
        start = time.perf_counter()
        with ThreadPoolExecutor(max_workers=n_python_workers) as pool:
            list(pool.map(one_matmul, range(N_CALLS)))
        return time.perf_counter() - start


if __name__ == "__main__":
    t_default = run_concurrent(n_python_workers=8, blas_threads=None)
    print(f"8 workers, BLAS default (unrestricted) threads, on 2 CPUs: {t_default:6.2f}s")

    t_tuned = run_concurrent(n_python_workers=8, blas_threads=1)
    print(f"8 workers, BLAS limited to 1 thread,             on 2 CPUs: {t_tuned:6.2f}s")

    t_matched = run_concurrent(n_python_workers=2, blas_threads=1)
    print(f"2 workers, BLAS limited to 1 thread,             on 2 CPUs: {t_matched:6.2f}s")
Output
8 workers, BLAS default (unrestricted) threads, on 2 CPUs:   0.28s
8 workers, BLAS limited to 1 thread,             on 2 CPUs:   0.16s
2 workers, BLAS limited to 1 thread,             on 2 CPUs:   0.16s

Real measurements from this machine, pinned to 2 CPUs for this run. Absolute times will differ on your hardware, but the pattern is worth trusting: leaving BLAS unrestricted while also running 8 concurrent workers was 75% slower than restricting BLAS to one thread per call. Matching worker count to CPU count (2 workers, 2 CPUs) performed as well as 8 tuned workers — extra workers past your core count bought nothing here.

Line-by-line walkthrough

psutil.Process().cpu_affinity([0, 1]) restricts this entire process to CPUs 0 and 1, simulating a small container — the realistic environment where this problem actually bites.

threadpool_limits(limits=blas_threads, user_api="blas") is the tuning knob itself: it caps how many threads the BLAS library underneath NumPy is allowed to use for the duration of the block, without needing an environment variable set before the process starts.

ThreadPoolExecutor(max_workers=n_python_workers) provides Python-level concurrency across separate matmul calls, competing for the same 2 CPUs as BLAS's own internal threads unless BLAS is told to stay small.

Common mistakes

Never checking what BLAS backend NumPy is actually using. threadpoolctl.threadpool_info() reports it directly — OpenBLAS and MKL both default to using many threads, and assuming otherwise is a common source of surprise.

Setting thread counts with environment variables after the process starts. OMP_NUM_THREADS and OPENBLAS_NUM_THREADS are read once, at library load time. Setting them inside a running Python process usually has no effect — set them before the process starts, or use threadpool_limits as shown above.

Assuming more workers is always safer. Past your CPU count, extra workers do not add capacity — they add contention. The 2-worker and 8-worker tuned runs above performed identically, because 2 CPUs cannot usefully run more than 2 things at once.

Forgetting this changes under a container CPU limit. A Kubernetes cpu: "2" limit throttles a process to roughly 2 cores' worth of CPU time, even on a much bigger physical node. NumPy sees the node's core count when it picks a thread count by default, not the container's limit — the exact mismatch this lesson is about.

Try it yourself

Remove the psutil.Process().cpu_affinity([0, 1]) line and rerun on your own machine. If your machine has many cores, the effect should shrink or disappear — this problem is specifically about scarce cores, not slow ones.

What to learn next

Researcher — Mathematics and papers.

Why oversubscription costs more than idle waiting

Beyond simple queueing for CPU time, oversubscribed threads compete for shared cache lines. Each context switch between competing threads can evict useful data from L1/L2 cache, so a thread resuming after a switch frequently pays a cache-miss penalty on top of the scheduling delay itself — the combined cost is measurably worse than what a naive queueing-delay estimate alone would predict.

The relevant environment variables

VariableControls
OMP_NUM_THREADSOpenMP-based libraries generally, including MKL's OpenMP backend
OPENBLAS_NUM_THREADSOpenBLAS specifically
MKL_NUM_THREADSIntel MKL specifically
NUMEXPR_NUM_THREADSnumexpr, used internally by some pandas operations

All are read at library initialisation. threadpoolctl (used above) works around the initialisation-time limitation by locating loaded native thread pools in-process and adjusting them directly through their runtime APIs, rather than relying on an environment variable read at import time.

Where the right number actually comes from

For a service handling one request at a time per Python worker process, a defensible default is one BLAS thread per worker, with worker count matched to the container's CPU limit — turning all available parallelism into Python-level worker concurrency, and none into per-call BLAS parallelism, which is the configuration the "tuned" rows above represent.

For a service handling large individual matrix operations one at a time — a single large batch inference call rather than many small concurrent ones — the reverse can be correct: fewer Python workers, more BLAS threads per call, since there is nothing competing with BLAS for those cores in that scenario. There is no universal answer independent of the actual traffic and workload shape; profile the real service, per profiling inference code, rather than applying either configuration by default.

References

  • OpenBLAS documentation, Faq — github.com/OpenMathLib/OpenBLAS/wiki/faq, covering thread count behaviour under nested parallelism.
  • threadpoolctl documentation — github.com/joblib/threadpoolctl, the library used above to control BLAS thread counts at runtime.
  • scikit-learn documentation, Parallelism, resource management, and configuration — a practical treatment of the same oversubscription problem from a widely used library that sits directly on top of BLAS.

What to learn next