MLflow
MLflow is a free tool that records every training run for you, stores the trained model with its input schema, and shows the whole history in a web page.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
MLflow is a free tool that keeps the record of your training runs for you, so you do not have to remember to keep it.
The analogy you have already lived
You once decided to keep a diary of your walks. Distance, time, how you felt. You wrote in it for two days and then stopped.
Then a phone app started counting your steps without being asked. Months later you could see every single day. You never wrote one line.
The diary failed because it needed you. The app worked because it did not.
MLflow is the app. Experiment tracking is the diary you were supposed to keep.
Why it exists
In the previous lesson you built a small tracker yourself. It worked, and it showed exactly what a tracker has to store.
Then the real questions start arriving:
- Can two people on the same team see each other's runs?
- Can I compare thirty runs on a chart instead of reading numbers?
- Where did the actual trained model file go?
- Which run produced the model that is live right now?
- Will this model still load in eight months?
You can build all of that. It is several months of work. MLflow is that work, already done, free and open source.
What is inside it
MLflow is four tools sharing one home.
Tracking — records every run: the settings, the scores, and files you attach.
Models — a standard folder layout for a saved model. It stores the model, the library versions, and the shape of the input it expects.
Model Registry — a named list of models with stages, so "the model in production" is a thing you can point at, not a filename someone remembers.
Evaluation — running a stored model against a dataset and recording the results.
Most people start with Tracking, and that alone earns its place on day one.
How it works
your training script
|
| mlflow.log_param("depth", 5)
| mlflow.log_metric("accuracy", 0.87)
| mlflow.sklearn.log_model(model, name="model")
v
[ MLflow ] --> one database file: every run, sortable
|
+------> a folder per run: the model, the versions, the schema
|
v
mlflow ui --> a web page in your browser, on your own machineNothing leaves your laptop. There is no account, no upload, no bill.
The one habit that pays for itself
MLflow can store the signature of a model — the names and types of the columns it expects.
That sounds like paperwork. It is what stops the failure from the first lesson. There, data arrived with the columns in the wrong order and the model answered anyway. With a signature stored, a missing column stops the model instead of confusing it.
The honest part
MLflow is not magic and it does not make you organised.
It will happily record four hundred runs named run_1, with no notes. You are then no better off than with a folder of files.
The tool solves storage. Naming things, writing down why you tried something, and deleting dead branches are still yours to do.
It is also a big install — several hundred megabytes with its dependencies. On a slow connection, plan for that.
Remember this
- MLflow records runs automatically, so the habit does not depend on your memory.
- It stores the model plus its schema, not only a file.
- It runs entirely on your own machine, free, with no account.
What to learn next
- Docker for ML — closing the reproducibility gap MLflow leaves open.
- Model serving — turning a logged model into something an app can call.
- CI/CD for machine learning — gating a promotion on tests that actually ran.
Developer — Code and libraries.
Setup
pip install mlflow scikit-learn pandasThis is a large install — MLflow pulls in a web server, a database layer and plotting libraries. Expect a few hundred megabytes and a few minutes on a normal connection. Everything after that is local and offline.
A complete tracked sweep
The example below writes to a single SQLite file in your working directory. It never contacts a server.
import mlflow
import mlflow.sklearn
import numpy as np
import pandas as pd
from mlflow.models import infer_signature
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
# Where the records go: one SQLite file in this folder. No server, no account.
mlflow.set_tracking_uri("sqlite:///mlflow.db")
mlflow.set_experiment("loan-default")
rng = np.random.RandomState(0)
cols = ["income", "years", "age", "cards", "enquiries", "utilisation"]
X = pd.DataFrame(rng.normal(size=(300, 6)), columns=cols)
y = (X["income"] + 0.8 * X["years"] - 0.5 * X["age"]
+ rng.normal(0, 0.7, 300) > 0).astype(int)
for C in [0.01, 0.1, 1.0, 10.0]:
with mlflow.start_run(run_name=f"logreg-C{C}"):
model = LogisticRegression(C=C, max_iter=1000)
accuracy = float(cross_val_score(model, X, y, cv=5).mean())
model.fit(X, y)
mlflow.log_param("C", C)
mlflow.log_param("model", "logistic_regression")
mlflow.log_metric("cv_accuracy", accuracy)
# The signature records the column names and types. Without it the
# model will happily accept columns in the wrong order.
mlflow.sklearn.log_model(
model, name="model",
signature=infer_signature(X, model.predict(X)))
print(f"C={C:<6} cv_accuracy={accuracy:.4f}")
runs = mlflow.search_runs(order_by=["metrics.cv_accuracy DESC"])
best = runs.iloc[0]
print("\nbest run")
print(" C ", best["params.C"])
print(" cv_accuracy", round(best["metrics.cv_accuracy"], 4))
print(" runs logged", len(runs))
loaded = mlflow.pyfunc.load_model(f"runs:/{best['run_id']}/model")
sample = X.iloc[6:12]
print("\nreloaded model")
print(" normal order ->", loaded.predict(sample))
# Columns arriving in a scrambled order: the signature puts them back by name.
print(" reversed cols ->", loaded.predict(sample[cols[::-1]]))
# A column missing entirely: the signature refuses to run at all.
try:
loaded.predict(sample.drop(columns=["utilisation"]))
except Exception as err:
print(" missing col ->", type(err).__name__)2026/08/20 14:08:08 INFO mlflow.store.db.utils: Creating initial MLflow database tables... 2026/08/20 14:08:08 INFO mlflow.store.db.utils: Updating database tables 2026/08/20 14:08:09 INFO mlflow.tracking.fluent: Experiment with name 'loan-default' does not exist. Creating a new experiment. C=0.01 cv_accuracy=0.8767 C=0.1 cv_accuracy=0.8833 C=1.0 cv_accuracy=0.8867 C=10.0 cv_accuracy=0.8900 best run C 10.0 cv_accuracy 0.89 runs logged 4 reloaded model normal order -> [1 1 0 1 0 0] reversed cols -> [1 1 0 1 0 0] missing col -> MlflowException
The three timestamped INFO lines appear once, on the first run, while the database is created. Your timestamps will differ, and the four accuracy figures should match. Run it twice and runs logged becomes 8 — MLflow appends, it does not replace.
Read the last three lines again
normal order and reversed cols produce identical predictions. The second call passed all six columns backwards. The stored signature matched them by name and put them back in place.
Compare that with the first lesson in this section, where the same mistake silently reversed a loan decision. The signature is a dozen extra characters and it removes an entire class of production bug.
missing col raises MlflowException. The full message is long: it prints the data it received and the schema it expected, side by side. It is loud, ugly and exactly right — the model refused rather than guessing.
See it in a browser
mlflow ui --backend-store-uri sqlite:///mlflow.dbOpen http://127.0.0.1:5000. You get a sortable table of runs, checkboxes to compare any two, and the stored files for each run. Stop it with Ctrl+C.
No output block for this one. The terminal prints a server banner whose contents vary by version, and the interesting result is a web page.
The bits of the API worth knowing
mlflow.set_tracking_uri("sqlite:///mlflow.db") — the storage backend. Recent MLflow versions have retired the old ./mlruns folder store and will raise an error if you point at one, so start with SQLite. For a team, swap in a PostgreSQL URI and nothing else in your script changes.
with mlflow.start_run(): — the context manager marks the run finished even if your training crashes. A crashed run recorded as FAILED is far more useful than a missing row.
mlflow.autolog() — one line before training, and MLflow logs the parameters, metrics and model for supported libraries without any log_param calls. Convenient for exploring, and worth turning off for production pipelines where you want to control precisely what is recorded.
infer_signature(X, y_pred) — reads the column names and types out of your training frame. Pass a DataFrame, not a NumPy array, or you get an unnamed tensor schema and lose the name checking entirely.
runs:/<run_id>/model — a URI that resolves through the tracking store. It survives files being moved, which raw paths do not.
A note for older installs. The name= argument to log_model arrived in MLflow 3. On MLflow 2.x the same call is mlflow.sklearn.log_model(model, artifact_path="model", signature=...). Check with python -c "import mlflow; print(mlflow.__version__)" if the call raises a TypeError.
Common mistakes
Logging the model without a signature. MLflow warns and carries on. You lose schema enforcement, which is most of the reason to use log_model instead of joblib.dump.
Committing mlflow.db and mlruns/ to git. They are outputs, and they grow fast. Add both to .gitignore and keep the tracking store somewhere shared instead.
Assuming a logged model is reproducible. MLflow records your Python package versions. It does not record your system libraries, your CUDA driver or your operating system. That gap is what Docker exists to close.
Using mlflow ui as a production server. It is a development server bound to localhost. A shared deployment needs mlflow server with a real database, an artifact store, and authentication in front of it.
One experiment for everything. Runs from three unrelated projects in one list is the same problem as one folder of files. Call set_experiment with a meaningful name per project.
Try it yourself
Add mlflow.log_param("data_hash", ...) using the fingerprint function from the experiment tracking lesson. Then change one value in X, re-run, and sort by accuracy in the UI.
You will see two runs with the same settings and different scores. The fingerprint column is the only thing on the page that explains why.
What to learn next
- Docker for ML — closing the reproducibility gap MLflow leaves open.
- Model serving — turning a logged model into something an app can call.
- CI/CD for machine learning — gating a promotion on tests that actually ran.
Researcher — Mathematics and papers.
What the Model format actually is
An MLflow model is a directory containing an MLmodel YAML descriptor plus artefacts. The descriptor lists one or more flavors — alternative interfaces to the same serialised object.
A scikit-learn model is typically written with two: sklearn (native, restores the exact estimator) and python_function (a generic predict(DataFrame) -> DataFrame|ndarray contract). The second is what makes serving infrastructure model-agnostic: a deployment target implements pyfunc once and can then serve PyTorch, XGBoost, ONNX and scikit-learn identically.
The descriptor also carries signature (input and output schema), saved_input_example_info, and an environment specification — conda.yaml, python_env.yaml and requirements.txt are all emitted.
Signature enforcement semantics
Worth knowing precisely, because the rules are asymmetric:
- Column matching for a
DataFrameinput is by name, not position. Extra columns are dropped; missing required columns raise. - Type checking is safe-cast only.
inttodoublepasses;doubletointdoes not, because it would lose information. - Optional columns are supported through
Schemawithrequired=False. - A
TensorSpecsignature (what you get from a NumPy input) checks shape and dtype only. Passing a NumPy array therefore gives you strictly weaker guarantees than aDataFrame.
The design point is that enforcement happens in pyfunc, above the framework, so the same guarantee holds whichever flavor is underneath.
Reproducibility ceiling
MLflow captures the Python dependency set. It does not capture:
- system shared libraries (
glibc,libgomp, BLAS implementation and version), - the CUDA driver and runtime, or the GPU architecture,
- environment variables that change numerical behaviour (
OMP_NUM_THREADS,MKL_NUM_THREADS,CUBLAS_WORKSPACE_CONFIG), - the ordering of the data as it was fed.
A requirements.txt with unpinned transitive dependencies is not a reproducible environment either. mlflow.models.build_docker and hash-pinned lockfiles close part of this gap; see Docker for ML for the rest.
Registry semantics, and their limits
The registry adds named models, integer versions, aliases and tags. Two properties are worth stating plainly because they are commonly misread:
- A registry entry is a pointer plus metadata. Promoting a version does not copy, revalidate or re-test anything.
- Stage transitions carry no built-in gating. "Only promote if the evaluation job passed" is a policy you implement in CI, not something the registry enforces.
Named aliases (@champion, @challenger) replaced the older fixed stages, precisely because fixed stages encoded a workflow that did not match most teams.
Scaling characteristics
The tracking store is a normal relational database with a metrics table holding one row per logged value per step. This is the usual bottleneck: a training job logging five metrics every step for 100k steps writes half a million rows.
Practical mitigations, in the order they usually become necessary: batch with log_metrics rather than repeated log_metric calls, reduce logging frequency for step-level metrics, move from SQLite to PostgreSQL well before a second concurrent writer appears, and move artifacts to object storage with the server proxying access.
Where it sits among alternatives
| Tool | Model | Strongest at | Main cost |
|---|---|---|---|
| MLflow | Open source, self-hosted | Model packaging and the pyfunc contract | You operate the server |
| Weights & Biases | Hosted, free tier | Live training dashboards, sweeps, collaboration | Data leaves your network |
| DVC | Open source, git-based | Versioning data and pipelines alongside code | No UI for comparing runs |
| Aim / Neptune | Open source / hosted | Fast UI over very many runs | Smaller packaging story |
These compose more often than they compete. DVC for data versioning plus MLflow for run tracking is a common and coherent pairing.
Reading
- Zaharia et al., Accelerating the Machine Learning Lifecycle with MLflow, IEEE Data Eng. Bull. 2018 — the design paper, and still the clearest statement of why the pyfunc abstraction exists.
- MLflow Models documentation — mlflow.org/docs/latest/ml/model
- Chen et al., Developments in MLflow: A System to Accelerate the Machine Learning Lifecycle, DEEM@SIGMOD 2020 — what changed once real users arrived.
What to learn next
- Docker for ML — closing the reproducibility gap MLflow leaves open.
- Model serving — turning a logged model into something an app can call.
- CI/CD for machine learning — gating a promotion on tests that actually ran.