GPUs: Memory, Scheduling and Cost

Sharing one GPU between models

A GPU with spare memory and spare compute can often serve several smaller models at once, instead of sitting mostly idle behind just one.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A GPU with room to spare can often serve more than one model at the same time. Otherwise one model gets the whole GPU to itself, while most of it sits idle.

The analogy you have already lived

A large kitchen with four stovetop burners does not need four separate kitchens. One family can cook two dishes on two burners at the same time. Same kitchen, same gas connection.

Giving each dish its own entire kitchen would work. But it would waste three empty kitchens' worth of space and equipment, for no real benefit. A GPU with spare room works the same way. Several smaller models can share it, the same way several dishes can share one stove.

Why it exists

Choosing a GPU for inference showed that a GPU's memory is often larger than one small model actually needs. Reading GPU utilisation honestly showed something else too: even a "busy" GPU can be using only a small fraction of its real capacity.

Put those two facts together, and a lot of GPU capacity across a real company can sit unused. Not because nothing needs it — because each model was given its own dedicated GPU, regardless of how much of it that model actually needs.

How it works

   ONE GPU, with spare memory and spare compute

   +--------------------------------------------------+
   |   model A's weights   |   model B's weights       |
   |   (using its share)   |   (using its share)       |
   |                                                    |
   |   request for A  -->  A answers using its slice    |
   |   request for B  -->  B answers using its slice    |
   +--------------------------------------------------+

   both models are ready, at the same time, on the same card

A real example you have seen

A company might offer several different AI features — a translator, a spam filter, a small recommendation model. It rarely gives each one a whole GPU to itself, if each one is small. Several of them often live together on shared hardware, each answering its own requests. One expensive resource gets split sensibly, instead of each demanding a full one.

The honest part

Sharing a GPU is not free. Two models sharing one card compete for the same memory and the same compute. One of them running unusually heavy work can slow the other one down — a real trade-off, not a pure win. Sharing makes sense when models are genuinely small relative to the GPU. It makes little sense when they are not.

Remember this

  • A GPU with spare memory and spare compute can often serve more than one model.
  • This works well for small models on a big GPU, not for models that already need most of it.
  • Sharing means models compete for the same resource. It saves money; it is not free.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Two models, one GPU, one process

shared_gpu.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

torch.cuda.empty_cache()
torch.cuda.reset_peak_memory_stats()

tok = AutoTokenizer.from_pretrained("distilgpt2")
tok.pad_token = tok.eos_token

mem_before = torch.cuda.memory_allocated() / 1024**2
model_a = AutoModelForCausalLM.from_pretrained("distilgpt2", dtype=torch.float16).to("cuda")
mem_after_a = torch.cuda.memory_allocated() / 1024**2
model_b = AutoModelForCausalLM.from_pretrained("distilgpt2", dtype=torch.float16).to("cuda")
mem_after_b = torch.cuda.memory_allocated() / 1024**2

print(f"this process's GPU memory before loading any model: {mem_before:8.1f} MB")
print(f"after loading model A:                               {mem_after_a:8.1f} MB  (+{mem_after_a-mem_before:.1f} MB)")
print(f"after ALSO loading model B:                          {mem_after_b:8.1f} MB  (+{mem_after_b-mem_after_a:.1f} MB)")

prompt = "The weather today is"
enc = tok(prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    out_a = model_a.generate(**enc, max_new_tokens=8, do_sample=False, pad_token_id=tok.eos_token_id)
    out_b = model_b.generate(**enc, max_new_tokens=8, do_sample=False, pad_token_id=tok.eos_token_id)

print("\nmodel A answer:", tok.decode(out_a[0], skip_special_tokens=True))
print("model B answer:", tok.decode(out_b[0], skip_special_tokens=True))
print("\nboth models answered from the SAME GPU, in the SAME process, at the same time.")
Output
this process's GPU memory before loading any model:      0.0 MB
after loading model A:                                  158.9 MB  (+158.9 MB)
after ALSO loading model B:                             318.6 MB  (+159.7 MB)

model A answer: The weather today is a bit of a shock, but it
model B answer: The weather today is a bit of a shock, but it

both models answered from the SAME GPU, in the SAME process, at the same time.

Two full copies of a real model, loaded and generating text from the same physical GPU. Together they use well under half a gigabyte. On a GPU with tens of gigabytes of memory, this is a small fraction of what is available. Real, measured spare room.

Line-by-line walkthrough

torch.cuda.memory_allocated(). This reports memory allocated by this process specifically, not the whole GPU. A tool like nvidia-smi reports total GPU memory used by every process combined. That is a genuinely different number — know which one you are looking at.

Loading model_b right after model_a, both .to("cuda"). Nothing special happens here. This is the entire mechanism. Two models both exist in the same GPU's memory at once, because there was room for both.

Both models answer with the identical text here. That is expected. They are two separate copies of the exact same model, given the exact same prompt with no randomness (do_sample=False). In a real system, model A and model B would usually be genuinely different models doing different jobs.

Common mistakes

Sharing without checking there is genuine spare capacity. Loading a second large model onto a GPU that is already nearly full crashes with an out-of-memory error. Not a graceful slowdown. Confirm real headroom first, using the measurement techniques from choosing a GPU for inference.

Assuming sharing has no performance cost. Two models on one GPU genuinely compete for the same compute resources. A traffic spike on one can slow the other down. Test under realistic concurrent load, not one request at a time.

No isolation between models that need it. In this simple, one-process approach, one model crashing can take the whole process down with it. Every model in it goes too. Real multi-tenant GPU sharing (see the researcher section) uses stronger isolation for exactly this reason.

Sharing models that are each already large. This pattern earns its keep when models are small relative to the GPU. Two models that each already use most of the GPU's memory or compute have very little left to actually share.

Try it yourself

Load a third copy of the model onto the same GPU and measure the memory increase again. Keep going until you can estimate, from real measurements, roughly how many copies this specific GPU could hold.

What to learn next

Researcher — Mathematics and papers.

Isolation mechanisms, from weakest to strongest

The demo shares a GPU the simplest possible way: multiple models in one process, one CUDA context. Production systems have several stronger options, trading flexibility for isolation:

  • Time-slicing — the GPU scheduler rapidly switches between processes' work, similar to how a CPU shares time between programs. Every process can use the full GPU when it gets its turn, but gets no guaranteed share and no memory isolation from other processes.
  • NVIDIA MPS (Multi-Process Service) allows multiple processes to submit work to the GPU concurrently, rather than strictly time-sliced. It improves throughput when workloads are individually too small to saturate the GPU alone. The cost is weaker fault isolation — a crash in one client process can, in some failure modes, affect others sharing the same MPS server.
  • NVIDIA MIG (Multi-Instance GPU), available on datacenter GPUs from the Ampere generation onward, partitions a single physical GPU into several fully isolated instances at the hardware level. Each gets its own dedicated memory and compute slice, invisible to and unaffected by the others. This is the strongest isolation available without using physically separate GPUs, at the cost of a fixed partition size chosen upfront rather than dynamic sharing.
  • Separate GPUs entirely — the strongest isolation, and the baseline this whole lesson is arguing against using by default for genuinely small models.

The packing problem

Deciding how to assign $m$ models to $n$ GPUs, given each model's memory and compute requirements and each GPU's capacity, is a bin-packing problem — NP-hard in general, but well approximated in practice by greedy heuristics (largest-model-first, best-fit) that are what most GPU schedulers (see scheduling GPUs on Kubernetes) implement under the hood. The quality of the packing directly determines how much of the "spare capacity" argument in this lesson is actually realised versus left stranded in unusable fragments.

Interference and tail latency under sharing

Sharing degrades tail latency disproportionately compared to average latency. A request landing at the same instant as another model's heavy burst experiences queuing delay that a request landing during a quiet moment does not. This asymmetry means sharing decisions should be evaluated against p95/p99 latency under realistic concurrent load (see tail latency and percentiles). Average latency measured in isolation systematically understates the real cost of sharing.

Reading

  • NVIDIA, Multi-Process Service (MPS) documentation — the mechanism and its isolation guarantees in detail
  • NVIDIA, Multi-Instance GPU (MIG) User Guide — partitioning semantics and supported hardware

What to learn next