AI glossary

Model serving

In one sentence Model serving is running a trained model as a live service that applications call — the engineering between a saved file and real users.

By Updated

Model serving is making a trained model available as a running service — typically an API — that applications query and get predictions back from, reliably, at scale.

A brilliant chef with no restaurant feeds nobody. Between the skill and the customers stands unglamorous infrastructure: a kitchen sized for the crowd, order-taking, quality control, and a plan for the Sunday rush. A trained checkpoint on disk is the chef; serving is the restaurant. The gap between "works in my notebook" and "handles ten thousand users" is where most production ML effort actually goes.

The core loop is an HTTP endpoint: request in → preprocess → model inference → postprocess → response out. Around that loop, serving adds what notebooks never need: batching requests to keep the GPU fed, autoscaling for traffic spikes, versioning (run v2 alongside v1, shift traffic gradually, roll back in seconds), monitoring of latency percentiles and error rates, and — over time — watching for data-drift.

The tooling landscape by weight class: FastAPI wrapping a model covers small services; dedicated servers (NVIDIA Triton, TorchServe, BentoML) add batching and multi-model management; LLMs get their own specialised stack (vLLM, TGI, SGLang) built around KV-cache management and continuous batching; and managed endpoints (SageMaker, Vertex, Hugging Face) trade money for operations. One structural decision shapes products: online serving answers one request now (chatbots, fraud checks), while batch scoring processes millions overnight (recommendations precomputed for tomorrow) — the second is cheaper and simpler wherever freshness allows.

Where to go next