TensorFlow and Keras

Stopping TensorFlow taking the whole GPU

By default TensorFlow grabs nearly all GPU memory the moment it starts — two lines of configuration make it take only what it uses.

Read these first

On this page 5
  1. Why the default is greedy on purpose
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The moment TensorFlow touches a GPU, it reserves nearly all of that GPU's memory — even for a tiny model — unless you tell it not to.

Think of a guest who checks into a hotel and books every room in the building, "in case I need them later". One person, forty rooms. Anyone arriving after — another program, a second TensorFlow script, your own notebook restarted — finds the hotel full.

That guest is TensorFlow's default behaviour on a GPU. The fix is called memory growth: take one room now, book more only when actually needed.

Why the default is greedy on purpose

The greed is not a bug. Asking the GPU for memory is slow, and doing it thousands of times during training costs real speed. Worse, claiming and releasing again and again splinters the memory into scattered fragments. Picture a shelf where no gap is big enough for a new book, even though total free space is plenty. That splintering is called fragmentation.

Reserving everything once avoids both problems. For a dedicated training server running one job, the default is the right call. For your laptop, a shared lab machine, or two experiments side by side, it is a menace.

How it works

default:          start ──▶ reserve (almost) ALL GPU memory
                            second program ──▶ "out of memory" ✗

memory growth:    start ──▶ reserve a little
                            need more? ──▶ take a little more
                            second program ──▶ fits alongside ✓

One rule matters: the choice must be made before TensorFlow first uses the GPU. Once the big booking has happened, it cannot be undone for that run — the configuration lines belong at the top of your script.

A real example you have seen

Shared computers at a college lab or an office. The first student to run a training script silently books the whole GPU. Every later student sees "out of memory" and assumes the machine is broken. The machine is fine; the first booking took everything. Whole lab-hours vanish to this one default, term after term.

Remember this

  • Default: TensorFlow reserves almost all GPU memory at first touch.
  • Memory growth = take only what is used, growing on demand.
  • Set it before the GPU is first used, at the very top of the script.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tensorflow

On a machine with an NVIDIA GPU, Linux wheels of TensorFlow 2.21 bundle the CUDA libraries (a multi-GB install); on Windows, recent TensorFlow is CPU-only natively — GPU support runs via WSL2. Output below verified with TensorFlow 2.21 on a CPU-only machine, and that is the honest part: this code runs safely everywhere, and its output tells you which world you are in.

The two lines, defensively written

gpu_config.py
import tensorflow as tf

gpus = tf.config.list_physical_devices("GPU")
print("GPUs visible:", gpus)

if not gpus:
    print("no GPU on this machine - nothing to configure")

for gpu in gpus:
    # must happen BEFORE any tensor touches the GPU
    tf.config.experimental.set_memory_growth(gpu, True)
    print("memory growth enabled on:", gpu.name)
Output
GPUs visible: []
no GPU on this machine - nothing to configure

On a machine with one NVIDIA card, the same script prints instead (values per your hardware):

Output
GPUs visible: [PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')]
memory growth enabled on: /physical_device:GPU:0

The walkthrough

list_physical_devices("GPU") is the first diagnostic to run on any new machine — before wondering why training is slow. An empty list means TensorFlow cannot see a GPU: none exists, or the driver/CUDA setup is broken. The loop-over-gpus form works unchanged on zero, one, or eight cards.

set_memory_growth(gpu, True) flips the allocator to on-demand growth. Placement matters more than anything else on this page: call it after any operation has touched the GPU and you get RuntimeError: Physical devices cannot be modified after being initialized. "Touched" includes innocent-looking lines — creating a tensor, building a model. Top of the script, before everything.

The environment-variable alternative needs no code changes:

bash
export TF_FORCE_GPU_ALLOW_GROWTH=true

Handy when the script is not yours to edit. Same effect, same timing rule (set before launch).

A hard cap instead of growth — give TensorFlow a fixed slice and never more:

python
tf.config.set_logical_device_configuration(
    gpus[0],
    [tf.config.LogicalDeviceConfiguration(memory_limit=4096)])  # MB

The right tool when two jobs must share one card predictably: each gets a budget, neither can starve the other. Growth is polite but unbounded; a limit is a contract.

Common mistakes

Configuring too late. The RuntimeError above, or — worse — no error because the import-time code of some library already touched the GPU. Keep the configuration in the first lines of the entrypoint, before importing modules that build models at import time.

Diagnosing the wrong OOM. "Out of memory" with memory growth ON means the model genuinely needs more than the card has — lower the batch size. "Out of memory" the moment a second process starts means someone else's default reservation is the cause. nvidia-smi shows who holds what; a TensorFlow process at near-100% memory with 0% utilisation is the greedy-default signature.

Expecting nvidia-smi numbers to mean usage. Reserved is not used. With the default allocator, nvidia-smi reports the reservation — nearly the whole card — telling you nothing about what the model needs. With growth on, the number becomes meaningful.

Assuming the fix transfers to PyTorch. PyTorch's caching allocator grows on demand already; its OOM stories differ — see moving tensors between CPU and GPU.

Try it yourself

On a GPU machine, run a script that builds a small model and sleeps, and watch nvidia-smi in a second terminal — once with growth enabled, once without. The two memory numbers you see are this entire lesson in one comparison. On a CPU machine, run the script above and confirm the defensive form exits cleanly.

What to learn next

Researcher — Mathematics and papers.

Why pre-allocation, formally

GPU memory allocation (cudaMalloc) synchronises the device and costs orders of magnitude more than pointer-bump allocation from a held pool. TensorFlow's BFC allocator ("best-fit with coalescing") therefore reserves a large arena once and sub-allocates from it: allocation becomes a free-list search, deallocation a coalescing merge of adjacent free chunks. Default arena size is a large fraction of device memory (historically ~95%, via per_process_gpu_memory_fraction), claimed at device initialisation.

The trade-off is fragmentation control versus multi-tenancy. Within one arena, BFC bounds external fragmentation by coalescing; across processes, the arena is the fragmentation — an opaque reservation invisible to other allocators. Memory growth changes the policy to incremental arena extension: cudaMalloc on demand in chunks, never returned until process exit (release-on-free would reintroduce the syscall cost and fragmentation churn). Formally: reserved memory $R(t)$ becomes a non-decreasing step function tracking the high-water mark of live allocations $L(t)$, instead of the constant $R(t) \approx M_{\text{total}}$. Symbols: $R$ — bytes reserved from the device; $L$ — bytes in live tensors; $M_{\text{total}}$ — device capacity.

What consumes $L$ during training: parameters, gradients, optimizer state (2 extra copies under Adam), and — dominating in deep nets — activations retained for backprop, $O(B \sum_l d_l)$ for batch size $B$ and layer widths $d_l$. Remedies ordered by cost: smaller $B$; mixed precision (halves activation bytes; Micikevicius et al. 2018, Mixed precision training); gradient checkpointing (recompute activations, $O(\sqrt{n})$ memory for $n$ layers; Chen et al. 2016, Training deep nets with sublinear memory cost); gradient accumulation to fake large batches.

Multi-tenancy beyond politeness

Logical device configuration implements static partitioning — the only fairness mechanism TensorFlow offers natively. Real cluster sharing moved the problem out of the framework: NVIDIA MPS multiplexes kernels from multiple processes; MIG (A100 onward) partitions the card at hardware level into isolated instances with dedicated memory slices — the hotel finally hires a manager. Scheduler-level treatments (Xiao et al. 2018, Gandiva: introspective cluster scheduling for deep learning) time-slice whole GPUs across jobs.

References

  • Chen et al. (2016), Training deep nets with sublinear memory cost.
  • Micikevicius et al. (2018), Mixed precision training.
  • Xiao et al. (2018), Gandiva: introspective cluster scheduling for deep learning, OSDI.

What to learn next