Saving a scikit-learn model safely
Saving a model seals everything it learned into a file for production to reuse — and doing it safely means whole pipelines, matching versions, and never loading strangers' files.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Saving a model writes everything it learned into one file, so another computer can load it and predict without retraining.
Think of making achaar. You spend a day cooking mangoes with spices, then seal the result in a jar. For months afterwards, anyone can open the jar and eat in seconds — the day's work is preserved inside.
Fittingly, Python's oldest jar format is literally called pickle. A trained model goes in, gets sealed as bytes in a file, and comes back out ready to predict.
Why it exists
Training happens once, on one machine, sometimes over hours. Predictions must happen forever after, on other machines, in milliseconds. The saved file is the handover between those two worlds. In real deployments the file is the product — that is what gets versioned, tested and shipped.
How it works
train once seal open anywhere
model + its ───→ model.joblib ───→ loaded copy predicts
learnings (the jar) without retrainingThree rules keep the jar safe:
- Seal the whole kitchen, not one spice. The model and every data-preparation step must go in together.
- Same recipe on both ends. The library versions that saved the file should match the ones loading it.
- Never eat from an unlabelled jar. A model file from an untrusted source can contain hidden instructions that run the moment you open it. Treat strange model files like strange food.
A real example you have seen
Your phone keyboard predicts your next word instantly, offline. Nobody trained that model on your phone. It was trained elsewhere, sealed into a file, and shipped inside the app — a jar of learning, opened millions of times a day.
Remember this
- A saved model file lets any machine predict without retraining.
- Save the whole pipeline — preparation steps and model, as one object.
- Opening a model file runs code. Only load files you trust.
What to learn next
- Model serving — putting the loaded pipeline behind an API.
- MLflow — registries that manage these artifacts and their versions.
- Docker for ML — pinning the whole environment, not only the model file.
Developer — Code and libraries.
Setup
pip install scikit-learn joblibTested against scikit-learn 1.7 and joblib 1.5.
Seal the pipeline, reopen it, use it
import os
import joblib
import numpy as np
import sklearn
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X = np.array([[1., 8.], [2., 7.], [3., 6.],
[6., 6.], [7., 5.], [8., 4.]])
y = np.array([0, 0, 0, 1, 1, 1])
pipe = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
# Save the WHOLE pipeline - scaler and model together, never separately
joblib.dump(pipe, "pass_model.joblib")
print("file size:", os.path.getsize("pass_model.joblib"), "bytes")
loaded = joblib.load("pass_model.joblib")
print("prediction:", loaded.predict([[7., 6.]]))
print("trained with sklearn", sklearn.__version__)file size: 1409 bytes prediction: [1] trained with sklearn 1.7.2
The byte count varies a little across sklearn and joblib versions; a tiny pipeline lands somewhere around a kilobyte or two.
The walkthrough
Why joblib and not plain pickle. joblib.dump is pickle underneath, with efficient handling of the large numpy arrays that live inside fitted estimators. For models, it is the community default. Everything true of pickle's security applies to joblib identically.
The loaded object needs no imports of your training code — until it does. Built-in estimators unpickle from scikit-learn itself. But a custom transformer unpickles by importing its class from wherever it lived at save time. If it lived in your training script, loading from a different script fails:
AttributeError: Can't get attribute 'ClipOutliers' on <module '__main__' (built-in)>
The tail of that message names wherever your loading process's main module lives, so it varies by how you ran Python.
The fix: define custom components in a real module, say mymodels/transforms.py, importable by both the trainer and the server.
Version skew warns — listen to it. Loading a file saved under a different scikit-learn version raises InconsistentVersionWarning. Sometimes it works anyway; sometimes attributes moved and predictions differ or crash. The discipline: record versions at save time, and pin them at load time.
import json, sklearn, joblib
meta = {"sklearn": sklearn.__version__, "joblib": joblib.__version__,
"trained_on_rows": int(len(X)), "features": ["study_hours", "sleep_hours"]}
with open("pass_model.meta.json", "w") as f:
json.dump(meta, f, indent=2)
print(open("pass_model.meta.json").read()){
"sklearn": "1.7.2",
"joblib": "1.5.3",
"trained_on_rows": 6,
"features": [
"study_hours",
"sleep_hours"
]
}Your version lines will match whatever you have installed. A sidecar file like this turns "which versions did we train this with?" from archaeology into a file read.
Common mistakes
Saving the model but not the scaler. The server then feeds raw numbers to a model trained on scaled ones. No error, wrong answers. Fitting and saving one pipeline object makes this mistake unwritable.
Loading pickles you did not create. Unpickling executes code chosen by whoever made the file — a model file can shell out, read your keys, anything. This is not theoretical; it is the standard attack on ML artifact stores. Load only files from your own systems, verify checksums, and prefer safer formats for sharing.
Letting versions drift. Training on a laptop with 1.7 and serving on an old image with 1.2 works right up until it does not. Put the exact versions in your server's requirements, sourced from the metadata file.
Retraining "the same" model instead of persisting it. Refitting from the same CSV on the server sounds equivalent — then a pandas upgrade reorders NaN handling, or sampling differs, and the deployed model is quietly not the audited one. Ship the artifact, not the recipe.
Try it yourself
Save the pipeline, then in a fresh Python session load it and predict — confirm nothing from the training script is needed. Then open pass_model.joblib in a hex viewer and find the string StandardScaler in the bytes: models are readable objects, not magic.
What to learn next
- Model serving — putting the loaded pipeline behind an API.
- MLflow — registries that manage these artifacts and their versions.
- Docker for ML — pinning the whole environment, not only the model file.
Researcher — Mathematics and papers.
Why pickle is code execution
Pickle is a stack-based virtual machine. Deserialisation replays opcodes, and the REDUCE opcode calls an arbitrary callable with arbitrary arguments — that is how objects reconstruct themselves via __reduce__. Nothing restricts that callable: os.system serialises fine. Loading is therefore exactly as trusting as executing the file. CVE databases and ML-supply-chain reports document in-the-wild malicious model files; scanning tools (e.g. picklescan) inspect opcodes but are bypassable, so scanning is a mitigation, not a guarantee.
Safer interchange formats
- skops (
skops.io.dump/load): serialises the estimator object graph but reconstructs only from an allowlist of trusted types; unknown types require explicit opt-in (trusted=[...]). Designed for sharing scikit-learn models publicly, e.g. on model hubs. - ONNX via
skl2onnx: exports the prediction function to a framework-neutral computation graph. Removes the Python and scikit-learn dependency entirely and freezes semantics — no version-skew ambiguity — at the cost of covering inference only, for supported estimator types. See ONNX. - PMML (via
sklearn2pmml): the older XML interchange standard, still common in banking and telecom scoring engines.
The decision rule: joblib inside one trusted, version-pinned system; skops when humans exchange sklearn objects; ONNX when serving crosses language or platform boundaries.
Versioning semantics
scikit-learn explicitly does not guarantee pickle compatibility across versions; InconsistentVersionWarning (added in 1.2) formalised the check by embedding __sklearn_version__ in the pickle. Fitted-attribute layouts are internal API. Consequences worth engineering for:
- The artifact's contract is (model bytes, exact dependency set). Content-address it: store a SHA-256 of the file alongside the metadata, which also gives tamper evidence.
- Migration across versions = load under the old version, re-fit or re-export under the new one, re-validate metrics. Registries like MLflow automate the bookkeeping; see also experiment tracking.
Artifact size
Size is dominated by fitted arrays: linear models are O(d); a random forest stores every node of every tree, commonly hundreds of megabytes at production scale; k-NN pickles its entire training set. Size therefore leaks modelling choices — and occasionally training data, which matters under privacy constraints: a pickled k-NN or SVM (support vectors are training rows) ships raw samples to whoever holds the file.
What to learn next
- Model serving — putting the loaded pipeline behind an API.
- MLflow — registries that manage these artifacts and their versions.
- Docker for ML — pinning the whole environment, not only the model file.