MLOps

Docker for ML

Docker packs your model together with the exact Python and the exact library versions it needs, so a different computer cannot quietly change the answer.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. What it actually fixes, and what it does not
  6. Where you have already seen it
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Docker packs your code, your model and the exact versions of everything they need into one sealed box that runs the same on any computer.

The analogy you have already lived

You have packed for a trip somewhere you did not know well. You took your own charger, your own medicine, your own tea. Not because the destination has nothing — because you did not know what it had.

That is Docker. You pack the exact Python, the exact libraries and your model into one box. Then it does not matter what the destination machine has.

The name comes from shipping. Before steel containers, cargo was loaded loose, and every port unloaded it differently. A container is the same sealed box in every port.

Why it exists

Here is the failure, in one line: your model works on your laptop and gives different answers on the server.

Not an error message. Different answers.

Look at what changes when a model file moves to another machine:

  • The Python version. Your laptop has one. The server has another.
  • The library versions. Your scikit-learn wrote the file. Theirs reads it.
  • The NumPy version underneath both of them.
  • System pieces you never think about: the maths library that does the multiplication, the compression library, the C library the whole thing sits on.

A saved model is not self-describing. It is a frozen object that assumes the exact software that created it is still there.

How it works

Two words you need, and they are constantly mixed up.

An image is the packed box. It is built once and does not change.

A container is one running copy of that box. You can run twenty at the same time from one image, and closing one does not affect the others.

  Dockerfile        the packing list you write
      |
      |  docker build
      v
   image            the sealed box: OS + Python + libraries + your model
      |
      |  docker run
      v
  container         one running copy, doing the work

The packing list is a plain text file called Dockerfile. It reads top to bottom: start from this base, install these libraries, copy these files, run this command.

What it actually fixes, and what it does not

Fixes: wrong Python version, wrong library version, a missing system package, a step someone did by hand on one machine and forgot on another.

Does not fix: a model trained on the wrong data, a model that has gone stale, or a bug in your code. Docker makes your mistakes perfectly repeatable. That is genuinely useful, and it is not the same as making them go away.

Where you have already seen it

Almost every website and app you use runs inside containers. When a company deploys an update at 2 a.m. and nothing breaks, this is a large part of why.

For a learner, the everyday win is smaller and more immediate. You can hand a colleague one file and one command, and it works on their machine on the first try.

The honest part

Docker is a real second thing to learn. It has its own vocabulary, its own errors, and it is genuinely confusing for the first week. Almost everyone finds it confusing at first. Read that sentence again if you need it.

It is also not small. Images for machine learning run to hundreds of megabytes because scientific Python is large. On a metered mobile connection, that matters, and it is worth knowing before you start rather than halfway through a download.

Remember this

  • An image is the sealed box; a container is one running copy of it.
  • Docker fixes environment differences, not data or logic mistakes.
  • The whole packing list is one plain text file you can read in thirty seconds.

What to learn next

Developer — Code and libraries.

First, watch the problem happen

Before the Dockerfile, here is the failure it prevents. Train a model on one version of scikit-learn and load it on another.

bash
pip install scikit-learn pandas joblib
train.py
import joblib
import numpy as np
import pandas as pd
import sklearn
from sklearn.linear_model import LogisticRegression

rng = np.random.RandomState(0)
cols = ["income", "years", "age"]
X = pd.DataFrame(rng.normal(size=(200, 3)), columns=cols)
y = (X["income"] + 0.8 * X["years"] - 0.5 * X["age"]
     + rng.normal(0, 0.7, 200) > 0).astype(int)

model = LogisticRegression(max_iter=1000).fit(X, y)
joblib.dump(model, "model.joblib")
print("saved model.joblib, trained with scikit-learn", sklearn.__version__)
Output
saved model.joblib, trained with scikit-learn 1.7.2

Your version number will differ if you installed a different scikit-learn. That difference is the whole point of this section.

predict.py
import sys

import joblib
import pandas as pd
import sklearn

model = joblib.load("model.joblib")

print("python       ", ".".join(map(str, sys.version_info[:2])))
print("scikit-learn ", sklearn.__version__)
print("expects      ", list(model.feature_names_in_))

applicants = pd.DataFrame([
    {"income": 1.5, "years": 0.9, "age": -0.4},
    {"income": -1.2, "years": -0.5, "age": 1.1},
])
print("predictions  ", model.predict(applicants))

Run python train.py and then python predict.py, and everything is calm:

Output
python        3.10
scikit-learn  1.7.2
expects       ['income', 'years', 'age']
predictions   [1 0]

Now load the same file in an environment with an older scikit-learn:

bash
python -m venv oldsk
oldsk/bin/pip install "scikit-learn==1.5.2" pandas joblib   # Windows: oldsk\Scripts\pip
oldsk/bin/python predict.py
Output
.../site-packages/sklearn/base.py:376: InconsistentVersionWarning: Trying to unpickle estimator LogisticRegression from version 1.7.2 when using version 1.5.2. This might lead to breaking code or invalid results. Use at your own risk. For more info please refer to:
https://scikit-learn.org/stable/model_persistence.html#security-maintainability-limitations
  warnings.warn(
python        3.10
scikit-learn  1.5.2
expects       ['income', 'years', 'age']
predictions   [1 0]

Look closely at what happened. It warned. It did not stop. And this time the predictions were identical.

That is the worst possible outcome, and it is why this lesson exists. A warning that is usually harmless trains you to ignore it.

For a LogisticRegression, which stores two small arrays, a version gap is usually survivable. For a tree ensemble it is not. If the internal layout changed between releases, the same warning precedes wrong numbers or a crash. By then it is in production.

Now the Dockerfile that removes the question

requirements.txt
joblib==1.5.3
numpy==2.2.6
pandas==2.2.3
scikit-learn==1.7.2

Use the versions from your own pip list, not these. The point is that they are pinned, not which numbers they are.

Dockerfile
FROM python:3.11-slim

# Print appears immediately instead of sitting in a buffer. Matters the first
# time you are staring at an empty log wondering if anything is running.
ENV PYTHONUNBUFFERED=1

WORKDIR /app

# requirements.txt is copied on its own, before the code. Docker caches each
# step, so editing predict.py later does not reinstall every library.
COPY requirements.txt ./
RUN pip install --no-cache-dir -r requirements.txt

COPY model.joblib predict.py ./

# Containers run as root unless you say otherwise. Two lines to fix that.
RUN useradd --create-home appuser
USER appuser

CMD ["python", "predict.py"]
bash
python train.py               # produces model.joblib on your machine
docker build -t loan-model .
docker run --rm loan-model
Output
python        3.11
scikit-learn  1.7.2
expects       ['income', 'years', 'age']
predictions   [1 0]

Note the first line. The container is running Python 3.11 even though the model was trained on 3.10. The libraries are pinned, so the load is clean and there is no warning.

No output block for docker build. Its log is a stream of layer ids, download sizes and timings. Those differ on every machine and every day. A made-up one would teach you to expect something you will never see.

What is true: the first build downloads a few hundred megabytes. The finished image lands somewhere around half a gigabyte. Every later build reuses the cache and finishes in seconds. Check the real size for yourself with docker images.

The four lines that are doing real work

FROM python:3.11-slim — the base. -slim drops compilers and documentation, saving several hundred megabytes. Use -slim unless a package needs to compile from source, in which case you want a multi-stage build rather than the fat base.

COPY requirements.txt ./ before COPY predict.py — this is the layer-caching trick, and it is the difference between a two-second rebuild and a four-minute one. Docker caches each instruction. Any change invalidates that step and everything after it. Your code changes fifty times a day; your requirements change monthly. Put the slow, stable step first.

pip install --no-cache-dir — pip keeps downloaded wheels in a cache directory. Inside an image that cache is dead weight baked into a layer forever.

USER appuser — a container that runs as root, mounted onto a host directory, can write to that directory as root. This line takes ten seconds and closes that door.

Common mistakes

Unpinned requirements. scikit-learn with no version means your image is a different image every time it builds. You then have a Dockerfile in git that cannot rebuild what is running in production. Pin every direct dependency, and for real deployments pin transitive ones too with pip freeze > requirements.lock.txt.

Copying the whole directory. COPY . . sweeps in your .git folder, your virtual environment, your notebook checkpoints and your dataset. Write a .dockerignore on day one:

.dockerignore
.git
.venv
__pycache__
*.ipynb_checkpoints
data/
mlruns/
mlflow.db

Training inside the image. The build then takes hours, cannot use a GPU on most build systems, and produces a different model every rebuild. Train outside, save the artefact, copy it in.

Baking secrets in with ENV. Anything set in the Dockerfile is readable by anyone who pulls the image, including in deleted layers. Pass credentials at run time with -e or a mounted secret.

Assuming your laptop's architecture matches the server's. An image built on an Apple Silicon Mac is arm64; most cloud servers are amd64. Use docker build --platform linux/amd64 when they differ, and expect it to be slower.

Expecting Docker to fix a stale model. It freezes the environment, not the world. Freshness is monitoring, not packaging.

Try it yourself

Change one digit in requirements.txt — set scikit-learn==1.5.2 — then rebuild and run.

The InconsistentVersionWarning from the top of this page appears inside the container. That is the useful lesson: Docker did not make the version problem disappear. It made it a line in a file you can read, review and change on purpose.

What to learn next

Researcher — Mathematics and papers.

What an image actually is

An OCI image is a JSON configuration plus an ordered list of layers, each a tar archive of filesystem changes, addressed by the SHA-256 digest of its contents. The runtime stacks them with a union filesystem (overlayfs on Linux) and adds one thin writable layer per container.

Three consequences follow directly from that structure, and all three surprise people:

  • Deleting a file in a later layer does not reclaim its bytes. The upper layer records a whiteout marker; the original data remains in the lower layer and ships with the image. RUN pip install x && rm -rf /root/.cache in one instruction works. Splitting it across two does not.
  • Instruction order determines rebuild cost. The cache key for an instruction includes the digests of everything before it, so an early change invalidates the whole tail.
  • A tag is mutable; a digest is not. python:3.11-slim resolves to different content over time. python@sha256:... does not. Pin digests wherever a rebuild must be byte-identical.

Reproducibility, honestly graded

Docker gives you a fixed userspace. It does not give you determinism. What still varies:

Fixed by the imageStill varies
Python and library versionsHost kernel version
System libraries in the imageCPU instruction set (AVX-512 changes reduction order)
File layout and entrypointGPU driver, exposed through the host
Installed binariesThread count, and therefore BLAS reduction order

The CPU point is the one that bites numerically. Numerical libraries dispatch at runtime to the widest available vector instruction set, changing the order of floating-point accumulation. Floating-point addition is not associative, so results differ in the last bits. Usually irrelevant; occasionally decisive at a decision boundary.

For GPU work, the container carries the CUDA runtime; the driver comes from the host through the NVIDIA container toolkit. Compatibility is one-directional — a newer driver runs an older runtime, not the reverse.

Achieving bit-identical builds additionally requires hash-pinned dependencies (pip install --require-hashes, or a uv.lock), a pinned base digest, and SOURCE_DATE_EPOCH handling for timestamps. Tools like Nix and Bazel's rules_docker target this properly; a plain Dockerfile does not.

Size, and why it matters more than it looks

A typical scientific Python image decomposes roughly as: base python:3.11-slim around 130 MB on disk, then SciPy and NumPy in the tens of megabytes each, scikit-learn around 30 MB, and a CUDA-enabled PyTorch wheel that alone exceeds 2 GB.

Size is a latency problem, not a storage problem. Autoscaling pulls the image on every cold start of a new node. A 5 GB image on a 1 Gbit/s link is roughly a minute of pull before a single request is served, and that minute lands exactly when traffic is spiking.

Levers, in decreasing order of effect: use the CPU-only wheel index when you do not need CUDA; multi-stage builds that leave compilers behind; --no-cache-dir; distroless or Alpine bases, with the caveat that Alpine uses musl rather than glibc, which breaks manylinux wheels and forces slow source builds.

Isolation is not a security boundary

Containers share the host kernel. Namespaces and cgroups isolate the view and the resources, not the attack surface. A kernel vulnerability reachable from a container is a host compromise.

The baseline hardening set: run as a non-root UID, --read-only root filesystem with explicit tmpfs mounts, drop all capabilities and add back only what is needed, --security-opt no-new-privileges, and a seccomp profile. Where the workload is genuinely untrusted, use a VM-backed runtime such as gVisor or Kata Containers rather than hardening a shared kernel.

For ML specifically, note that joblib and pickle execute arbitrary code on load. Loading an untrusted model file is remote code execution, container or not. Formats such as safetensors exist precisely to remove that property.

Reading

What to learn next