What is MLOps?
MLOps is the set of habits and tools that keep a trained model working after it leaves your laptop, where nobody is watching it.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
MLOps is everything you do to keep a trained model working after it leaves your laptop.
MLOps is short for machine learning operations — operations meaning the daily running of a system, not the building of it.
The analogy you have already lived
You have cooked something at home that turned out perfect. Someone at the table said you should open a restaurant.
A restaurant is a completely different problem. The same dish, two hundred times, every day. A different cook on Tuesday. A supplier who sends smaller onions this month. An inspector who wants to see the kitchen.
The recipe was never the hard part. Running it reliably is.
A trained model is the recipe. MLOps is the restaurant.
The failure this whole field is named after
Here is the story that turned MLOps into a job title.
You train a model on your laptop. It scores well on your test data. You are pleased. You hand the file to a teammate, and it stops working.
It does not crash. Crashing would be kind. It returns answers, and the answers are wrong.
Some real reasons this happens:
- Their pandas version differs from yours, and a function you relied on changed its behaviour.
- The live data arrives with the columns in a different order.
- A column that was a number in your file arrives as text from the database.
- You cleaned the training data by hand, once, in a step nobody wrote down.
- The model learned last year's customers, and this year's customers behave differently.
Not one of those is a machine learning problem. Every one of them will take your model down.
How it works
MLOps is six repeating steps. Each word is a place where things break.
data -> train -> test -> package -> serve -> watch
^ |
|______________ retrain when it goes stale _____________|Data — where the examples come from, and who guarantees they keep arriving in the same shape.
Train — running the learning, in a way someone else can repeat next month.
Test — checking the model, and also checking the data that feeds it.
Package — wrapping the model with the exact library versions it needs, so a different computer cannot change the answer.
Serve — putting it somewhere an app can send a question and get an answer back in milliseconds.
Watch — noticing when the answers quietly start getting worse.
The arrow back to the start is the part beginners miss. A model is not a thing you finish. It is a thing you keep.
Where you have already seen this working
- Your bank app flags a strange payment within a second, day and night, with nobody watching a screen.
- Your email spam filter keeps working even though spammers deliberately change their wording every week.
- A ride app predicts your arrival time using traffic that changed twenty minutes ago.
- A shopping app re-ranks its home page as new products arrive.
Every one of those is a model plus a great deal of unglamorous machinery around it.
The honest part
Most of MLOps is not machine learning.
It is packaging, testing, logging and watching. Many people find that disappointing when they first hear it. It is also the reason most trained models never reach a single real user.
The hardest failure to catch is the silent one. A crashed service pages somebody at 3 a.m. A model that has quietly become wrong pages nobody, for months.
Remember this
- Training a model is the small part. Keeping it working is the rest.
- Most production failures are boring: versions, column order, data types, undocumented steps.
- The dangerous failure is silent — a confident wrong answer with no error message.
What to learn next
- Experiment tracking — how to make a good result repeatable next month.
- Docker for ML — freezing the library versions so a different machine cannot change the answer.
- Monitoring and model drift — catching the silent failure.
Developer — Code and libraries.
The failure, in twelve lines of output
Talking about training-serving skew is less useful than watching it happen. Training-serving skew means the data your model sees when it is being used differs from the data it learned on.
Here is the cheapest version of it. It is also the one that has cost the most money: two feature columns in the wrong order.
Setup
pip install scikit-learn pandas numpyRuns on any laptop in about a second. No GPU, no download, no cloud account.
import warnings
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
# Ten loan applications. income = monthly income in thousands of rupees,
# years = years of credit history. repaid: 1 = paid the loan back, 0 = did not.
train = pd.DataFrame({
"income": [9.0, 12.0, 15.0, 18.0, 22.0, 30.0, 38.0, 45.0, 52.0, 60.0],
"years": [0.5, 1.0, 4.0, 2.0, 1.5, 3.0, 2.5, 4.5, 3.5, 5.0],
})
repaid = np.array([0, 0, 0, 0, 0, 1, 1, 1, 1, 1])
# C=0.02 is heavy regularisation. On ten rows it stops the model from
# reporting 99.99% confidence, which would hide the point being made here.
model = LogisticRegression(C=0.02).fit(train, repaid)
print("columns the model was fitted on:", list(model.feature_names_in_))
# One applicant: earns 50k a month, has 4 years of credit history.
row = pd.DataFrame([{"income": 50.0, "years": 4.0}])
print("correct order -> repaid =", model.predict(row)[0],
" confidence =", round(model.predict_proba(row)[0, 1], 3))
# The bug: a teammate builds the same applicant as a bare list of numbers,
# in the order the incoming JSON happened to use — years first, income second.
swapped = np.array([[4.0, 50.0]])
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
guess, conf = model.predict(swapped)[0], model.predict_proba(swapped)[0, 1]
print("swapped, numpy -> repaid =", guess, " confidence =", round(conf, 3))
print(" sklearn said:", caught[0].message if caught else "nothing")
# The identical mistake, but with the column names still attached.
try:
model.predict(pd.DataFrame([{"years": 4.0, "income": 50.0}]))
except ValueError as err:
print("swapped, named ->", str(err).splitlines()[0])columns the model was fitted on: ['income', 'years'] correct order -> repaid = 1 confidence = 0.983 swapped, numpy -> repaid = 0 confidence = 0.025 sklearn said: X does not have valid feature names, but LogisticRegression was fitted with feature names swapped, named -> The feature names should match those that were passed during fit.
Read those four lines slowly
The same applicant. The same model. Two answers that are opposites.
Line 2 is the truth. A person earning 50k with 4 years of history gets approved, at 98.3% confidence.
Line 3 is the bug. The two numbers arrived in the wrong order, so the model read the applicant as earning 4k a month with 50 years of history. It rejected them, at 97.5% confidence in the rejection.
There was no exception. There was no error code. A rejection letter goes out and nobody ever finds out why.
Line 4 is the only thing that saved you, and it is a warning, not an error. Warnings scroll past in production logs. Nobody reads them.
Line 5 is the fix, and it costs nothing. Pass a DataFrame with named columns and scikit-learn refuses to run. A loud failure at 10 a.m. beats a silent one for six months.
Why this is an MLOps lesson and not a pandas lesson
Notice what would have caught this bug at each stage of the pipeline:
| Stage | What would have caught it |
|---|---|
| Package | Storing feature_names_in_ next to the model file |
| Serve | Validating the incoming request against that stored list |
| Test | One test that sends a known applicant and asserts a known answer |
| Watch | An alert on "share of applications approved" dropping overnight |
Four independent chances. A team with none of them ships this bug. A team with any one of them catches it. That redundancy is MLOps.
Common mistakes
Passing bare NumPy arrays to a model at serving time. It is faster and it removes your only safety net. Build a DataFrame from a dictionary, keyed by name, and let the library check you.
Saving the model but not the recipe. A .joblib file alone is not reproducible. Save the library versions, the feature names, the training data hash and the test scores alongside it, in a small JSON file you write yourself.
Cleaning training data in a notebook cell. Anything you do to the training data must also happen to live data, byte for byte. If that step lives in a cell you ran once, it does not exist. Put every transformation inside a scikit-learn Pipeline so it travels with the model.
Treating the test score as a promise. A score on held-out data tells you how the model does on data from that week. It says nothing about next quarter. See model evaluation for what the number does and does not mean.
Try it yourself
Change C=0.02 to the default C=1.0 and run it again. Both confidences will move very close to 1.000 and 0.000.
Then ask yourself the uncomfortable question. If your dashboard only plots average confidence, would either version of this bug appear on it?
What to learn next
- Experiment tracking — how to make a good result repeatable next month.
- Docker for ML — freezing the library versions so a different machine cannot change the answer.
- Monitoring and model drift — catching the silent failure.
Researcher — Mathematics and papers.
The framing paper
Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems, is the origin document for this field and still the sharpest thing written about it. Its central observation is a figure: the box labelled "ML code" occupies a few percent of a real system's area. Configuration, data collection, feature extraction, serving infrastructure, process management and monitoring occupy the rest.
Its most-cited contribution is the CACE principle — Changing Anything Changes Everything. In a system with $n$ input features, no feature has an isolated effect. Removing feature $x_i$ changes the learned weights on all remaining features, because the optimiser re-solves a coupled problem. Consequently:
- Model components cannot be unit-tested in the way software components can.
- There is no meaningful notion of an isolated "interface" between a feature and a model.
- Ablation cost is $O(n)$ retrains, not $O(1)$ inspections.
The paper's taxonomy of ML-specific debt is worth memorising: entanglement (CACE), correction cascades, undeclared consumers, unstable data dependencies, underutilised data dependencies, feedback loops (direct and hidden), pipeline jungles, dead experimental codepaths, and configuration debt.
Training-serving skew, formally
Let $p_{\text{tr}}(x, y)$ be the joint distribution the model was fitted on and $p_{\text{sv}}(x, y)$ the distribution encountered at serving time. Skew is any $p_{\text{tr}} \neq p_{\text{sv}}$, but the causes decompose usefully:
- Implementation skew — the feature transformation code differs between the training path and the serving path. This is a software defect, not a distributional one, and it is by far the most common. The demonstration in the Developer tab is this case.
- Covariate shift — $p_{\text{tr}}(x) \neq p_{\text{sv}}(x)$ while $p(y \mid x)$ is stable (Shimodaira, 2000).
- Label shift — $p_{\text{tr}}(y) \neq p_{\text{sv}}(y)$ while $p(x \mid y)$ is stable.
- Concept drift — $p(y \mid x)$ itself changes over time.
Only the last three are statistical problems. The first is a plumbing problem, and it dominates real incident counts. Feature stores exist mainly to make implementation skew structurally impossible by serving both paths from one definition.
Measurable maturity
Breck et al. (2017), The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, converts "is this system production-ready" into 28 scored tests across four axes: features and data, model development, ML infrastructure, and monitoring. The rubric is deliberately blunt — each test scores 0, 0.5 or 1 — and the paper reports that surveyed Google teams cluster at low scores on the monitoring axis specifically.
The highest-value tests in the rubric, by the authors' own reporting, are unglamorous: test that every feature is beneficial, test that the model is not stale, test for training-serving skew, and test that the serving path handles a rollback.
Underspecification
D'Amour et al. (2020), Underspecification Presents Challenges for Credibility in Modern Machine Learning, gives the sharpest statistical account of why an in-distribution test score fails to predict deployment behaviour.
The argument: a training pipeline typically admits a large set of predictors with near-identical held-out risk. The pipeline does not constrain which member of that set you get — random seed, data ordering and hardware non-determinism do. Members of the set differ arbitrarily in behaviour under distribution shift and on subgroups.
The practical consequence is that seed variance is not noise to be averaged away in a paper. It is a lower bound on the deployment variance you should expect.
Data quality upstream
Sambasivan et al. (2021), "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI, documents 92% of interviewed practitioners experiencing compounding downstream failures originating in data collection and annotation. Cascades are described as opaque, diagnosed late, and typically triggered by events far upstream of the model.
Cost accounting
Three numbers that are routinely underestimated when planning ML systems:
- Inference dominates lifetime compute. A model trained once and served for a year at even modest traffic spends far more FLOPs on inference than on training. Optimise training cost only after inference cost is understood.
- Retraining is neither free nor reliably beneficial. Retraining cadence should be derived from measured performance decay, not from a calendar.
- Human labelling is often the binding constraint. Label latency, not model latency, sets how fast you can detect and correct a regression.
Papers
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — papers.nips.cc/paper/5656
- Breck et al., The ML Test Score, IEEE Big Data 2017 — research.google/pubs/pub46555
- D'Amour et al., Underspecification Presents Challenges for Credibility in Modern Machine Learning, 2020 — arxiv.org/abs/2011.03395
- Sambasivan et al., Data Cascades in High-Stakes AI, CHI 2021 — research.google/pubs/pub49953
- Shimodaira, Improving predictive inference under covariate shift, JSPI 2000 — the origin of importance-weighted correction.
- Paleyes, Urma and Lawrence, Challenges in Deploying Machine Learning: a Survey of Case Studies, 2020 — arxiv.org/abs/2011.09926
What to learn next
- Experiment tracking — how to make a good result repeatable next month.
- Docker for ML — freezing the library versions so a different machine cannot change the answer.
- Monitoring and model drift — catching the silent failure.