AI glossary

MLOps

In one sentence MLOps is the discipline of running machine learning in production — versioning, deploying, monitoring and retraining models as a routine, not a heroic event.

By Updated

MLOps is the set of practices for taking models from notebook to production and keeping them healthy there — the ML equivalent of DevOps.

Winning a cooking competition once is talent. Running a restaurant that serves the same quality every day — through staff changes, supplier changes, and the Diwali rush — is systems. A model that scored well in a notebook is the competition dish; MLOps is the restaurant: everything required so that predictions ship reliably today, next month, and after the fifth retrain.

The reason it exists as its own discipline: ML systems break in ways ordinary software does not. Code is versioned, but here data and models need versioning too — "which training-data and which checkpoint produced the model now serving?" must have an answer during an incident. Software fails loudly with exceptions; models fail silently, staying up while accuracy rots from data-drift. And a result that cannot be reproduced — same data, same code, same metrics — cannot be debugged.

The practical toolkit, by concern:

experiment tracking    MLflow, W&B          which run produced what, with which settings
data/model versioning  DVC, model registry  reproducibility and rollback
pipelines              Airflow, Kubeflow    retraining as an automated, tested path
serving                vLLM, Triton, cloud  see model-serving
monitoring             Evidently, custom    drift, latency, live accuracy

Adopt in that order as you grow; a solo project needs tracking and versioning long before it needs pipelines. The mindset shift is the real content: treat models as perishable products with a lifecycle — deployed, monitored, retrained, retired — not as artefacts finished on the day the metric looked good.

Where to go next