Registries, Artifacts and Environments

Pinning ML dependencies

Pinning writes down the exact library versions a model was built with, so the same code cannot quietly start behaving differently.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Pinning means writing down the exact version of every library your code depends on. A fresh install then cannot quietly get a different one.

The analogy you have already lived

You have followed a recipe that says "a pinch of salt." The dish tasted different each time you made it. A recipe that says "half a teaspoon of salt" gives you the same dish every time. There is no room left to guess.

"Install scikit-learn" is the pinch-of-salt version. "Install scikit-learn version 1.7.2" is the half-teaspoon version.

Why it exists

Software libraries change over time — new features, bug fixes, and sometimes a changed default that quietly alters behaviour. A line like pip install scikit-learn, with no version attached, installs whatever the newest release happens to be that day.

Run that same line today and again in six months, and you can get two different versions, silently. A model that trained perfectly well can be loaded by a newer library that reads the file slightly differently. A retraining script can produce a slightly different model than it would have last month — for no reason the code itself explains.

How it works

   unpinned:  requirements.txt says "scikit-learn"
        |
        v
   today:      installs 1.7.2
   in 6 months: installs 1.9.0   <- different code is now running,
                                     with no change to your own files

   pinned:    requirements.txt says "scikit-learn==1.7.2"
        |
        v
   today:      installs 1.7.2
   in 6 months: installs 1.7.2   <- guaranteed, every time

Pinning does not stop libraries from changing. It stops your project from being dragged along by a change you never asked for and never reviewed.

A real example you have seen

An app that suddenly looks or behaves differently after "just an update" is a familiar version of this. Nobody changed a setting; an underlying piece changed instead. Pinning is the discipline that keeps that kind of surprise out of a machine learning project's dependencies specifically.

The honest part

Pinning trades one problem for another, and it is worth naming honestly. Pinned versions eventually fall behind, missing security fixes and real improvements. The answer is not to avoid pinning. Update pins deliberately, on your own schedule, and re-test when you do — rather than being updated by accident every time someone runs an install.

Remember this

  • An unpinned dependency can install a different version tomorrow than it installed today, with no change to your own code.
  • Pinning writes the exact version down, so an install is reproducible instead of a guess.
  • Pinning is not "set once and forget" — update pins on purpose, and re-test when you do.

What to learn next

Developer — Code and libraries.

Setup

No installs needed beyond what is already on your machine — this inspects versions that are already there.

Checking what is actually installed, and catching a real mismatch

pinning_deps.py
import sys

import numpy
import sklearn

# What is actually installed, right now, on this machine.
installed = {
    "python": sys.version.split()[0],
    "numpy": numpy.__version__,
    "scikit-learn": sklearn.__version__,
}
print("installed right now:", installed)


def check_environment(recorded, installed):
    """Compares the versions an artifact was trained with against what is
    installed now. Loading a model in the wrong environment is how
    'it worked on my machine' becomes a production incident."""
    problems = [
        f"{lib}: trained with {recorded[lib]}, running {installed[lib]}"
        for lib in recorded
        if recorded[lib] != installed.get(lib)
    ]
    return problems


# an artifact trained on the current machine: everything matches
same_env = check_environment(installed, installed)
print("\nartifact trained here, loaded here:", same_env or "no mismatches")

# an artifact whose metadata says it was trained on an older scikit-learn
old_artifact_metadata = {**installed, "scikit-learn": "1.2.0"}
mismatches = check_environment(old_artifact_metadata, installed)
print("\nartifact trained on scikit-learn 1.2.0, loaded here:")
for m in mismatches:
    print(" ", m)

# what requirements.txt should contain: name==exact-version, not a bare name
print("\nwhat requirements.txt should contain, exactly:")
for lib, version in installed.items():
    if lib != "python":
        print(f"  {lib}=={version}")
Output
installed right now: {'python': '3.10.11', 'numpy': '1.26.4', 'scikit-learn': '1.7.2'}

artifact trained here, loaded here: no mismatches

artifact trained on scikit-learn 1.2.0, loaded here:
  scikit-learn: trained with 1.2.0, running 1.7.2

what requirements.txt should contain, exactly:
  numpy==1.26.4
  scikit-learn==1.7.2

The specific version numbers reflect what is installed on the machine that ran this — yours will very likely differ, and that is expected. The comparison logic and its output shape are exact regardless of which versions you actually have.

Walking through it

installed is read from the libraries themselves, via __version__, not typed in by hand. This is the same information pip show <package> reports, read programmatically instead of by eye. It is the honest way to know what is really running, rather than what a file claims should be running.

check_environment compares two dictionaries and reports only the differences. This is the same check that belongs inside a real model loader. Before trusting a model file, confirm the environment reading it back matches, at least closely, the one that wrote it.

The mismatch example is deliberately artificial. old_artifact_metadata is constructed by hand here. It stands in for what a real packaged artifact's metadata would report, if it had genuinely been trained under an older library version.

Common mistakes

Pinning only the libraries you directly import. Your direct dependencies have their own dependencies, which also change over time. A full pin — often generated by a lockfile tool — covers the whole tree, not only the packages your code names in an import line.

Never re-pinning at all. A requirements.txt frozen three years ago and never revisited misses every security fix released since. Schedule pin updates — quarterly is a reasonable default — rather than either updating constantly or never.

Pinning in development, but not in the container that actually ships. A Dockerfile that runs pip install -r requirements.txt only guarantees reproducibility if requirements.txt itself is fully pinned. It is the same discipline from Docker for ML, applied specifically to the dependency list rather than the whole image.

Assuming a pinned version pins everything about behaviour. Some libraries depend on system-level pieces — a BLAS implementation, a compiler, a CUDA version — that a Python-level pin does not touch at all. Training and serving environment parity covers this gap in more depth.

Try it yourself

Change old_artifact_metadata to instead simulate a numpy mismatch instead of scikit-learn, and rerun. Notice the function needed no changes at all — check_environment was written generically over whatever keys the two dictionaries happen to share.

What to learn next

Researcher — Mathematics and papers.

Pinning versus lockfiles

A requirements.txt with == pins fixes the direct dependencies you name. It does not, on its own, fix the transitive dependency tree — the dependencies of your dependencies — unless every single one is also pinned by hand, which does not scale.

Lockfile tools — pip-tools' pip-compile, poetry.lock, uv.lock — resolve the full dependency graph once, and record every resulting version, transitive dependencies included. A later pip install from the lockfile reproduces the exact same environment down to the last package, well beyond the packages you named directly.

Hash pinning, for a stronger guarantee

pip-compile --generate-hashes adds a cryptographic hash of each package's exact file contents alongside its version. A version number alone assumes the package registry never serves different bytes under the same version string — rare, but it has happened, whether through a compromised registry account or a republished release. Hash pinning makes that assumption unnecessary: pip refuses to install anything whose downloaded bytes do not match the recorded hash, which is a supply-chain security property, not only a reproducibility one.

Reproducibility beyond the Python layer

Full bit-for-bit reproducibility of a training run depends on more than that. The BLAS/LAPACK implementation NumPy links against matters — OpenBLAS versus Intel MKL can produce different floating-point rounding in the same computation. So do the CUDA and cuDNN versions underneath a GPU-accelerated framework, and even CPU instruction set differences (AVX2 versus AVX-512) affecting floating-point summation order. Pineau et al., Improving Reproducibility in Machine Learning Research (JMLR, 2021, the report behind the NeurIPS reproducibility checklist), documents how often published results fail to reproduce for reasons that trace back to exactly this class of under-specified environment detail.

Environment reproducibility as a spectrum

LayerToolWhat it fixes
Language dependenciespip + lockfileExact Python package versions
System librariesDocker / NixBLAS, CUDA, compiler, OS packages
Hardware behaviourDeterminism flags, fixed seedsFloating-point rounding, GPU kernel non-determinism

Each layer down this table is progressively more expensive to guarantee, and most teams reasonably stop partway down it — pinning Python dependencies and containerising the rest, rather than chasing bit-exact reproducibility across every layer.

What to learn next