Model deployment
Deployment means putting your model somewhere with a public address that stays awake, so other people can use it when your laptop is shut.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Deployment means putting your model somewhere with a public address that stays awake when your laptop is shut.
The analogy you have already lived
Cooking at home and running a food stall are not the same job.
At home you cook when you feel like it. If you burn something, you start again and nobody notices. When you go to sleep, the kitchen sleeps too.
A stall needs a fixed spot people can find. It needs opening hours somebody keeps. It needs to handle a rush without collapsing, and it needs to survive one customer ordering something strange.
Your model on your laptop is home cooking. Deployment is opening the stall.
Why it exists
A model in a notebook has three problems the moment somebody else wants to use it.
Nobody can reach it. It lives on your machine, behind your Wi-Fi router. There is no address to send a request to.
It sleeps when you do. Close the lid and your project is offline. Your demo is at 9 a.m. and you are asleep at 8.
It handles one person. A notebook runs one thing at a time. Thirty classmates clicking at once is a different problem.
Deployment fixes all three. It moves the model to a computer that is always on, always reachable, and ready for many requests.
How it works
YOUR LAPTOP A DEPLOYED SERVICE
------------------- ---------------------------
model.py https://your-app.example/predict
runs when you press play runs all day, every day
dies when you close the lid has an address anyone can call
one person at a time many people at onceThe thing you build is an API — a fixed address that answers requests. It takes a question in a known format and sends an answer back in a known format. Your website, your friend's app and a script can all talk to that one address.
The money conversation, first
This section is here first because it decides everything else, and because most tutorials skip it.
Renting a computer with a graphics card costs roughly 25 to 60 rupees an hour. A powerful one costs more. Run it day and night for a month and the bill reaches tens of thousands of rupees. For a student project with four users, that is money set on fire.
So start with the honest ranking.
Cheapest: do not host a language model at all. Call somebody else's. You pay a tiny amount per request, and nothing at all when nobody is using it. Several providers have a free daily allowance that comfortably covers a college project.
Next cheapest: host only the small thing. A spam classifier, a sentiment model or an embedding model runs fine on an ordinary processor. That fits in a free tier. The big language model stays behind somebody else's API.
Only then: rent a graphics card. Do this when three things are true. Real users, a real reason the model must be yours, and a real budget.
Most projects never need step three. Reaching for it first is the most common and most expensive mistake in this whole section.
Free places to put a project today
- Hugging Face Spaces — a free processor tier that hosts a small model with a web page in front of it. It sleeps when idle and wakes on the next visit. The best free option for a demo you want to share.
- Render, Railway, Fly.io — free or very cheap small tiers for a plain web service. Limits are tight; read them.
- Cloudflare and similar edge platforms — a free daily allowance for running small models close to users.
- A free virtual machine — some cloud providers offer a permanently free small server. Terms and availability change often, so check before you build a plan around one.
Free tiers all share one behaviour worth knowing: they sleep. Nobody visits for fifteen minutes, the service shuts down, and the next visitor waits twenty seconds while it wakes. That is called a cold start. It is fine for a demo and wrong for anything a person is waiting on.
Where you have already seen this
Every AI feature you use is somebody's deployment. The chat box on a bank site, the caption suggestion in a photo app, the spam filter in your inbox.
You never see the model. You see an address that answers.
The honest part
Deployment is where student projects most often stall, and the reason is rarely the model. It is the boring parts.
Somebody has to keep it running. Something has to be watched. The bill has to be paid every month. When it breaks at 2 a.m., it stays broken until a person wakes up.
That is not a warning to avoid it. It is the reason one line on a CV beats another. "I deployed it and it stayed up for six months" is worth far more than "I trained a model in a notebook".
Remember this
- Deployment gives your model an address that stays awake.
- Start with the cheapest option that works, and usually that means calling an API instead of hosting one.
- Free tiers sleep. Plan for the first visitor after a quiet hour to wait.
What to learn next
- Docker for ML — the packaging step, in full.
- Model serving — throughput, batching and scaling in depth.
- Monitoring and model drift — noticing when a deployed model goes quietly wrong.
Developer — Code and libraries.
Setup
pip install fastapi httpx uvicornFastAPI turns Python functions into HTTP endpoints. httpx powers the in-process test client used below, so you can verify everything without opening a port.
A service you can test without running a server
TestClient calls your app directly in the same process. No port, no browser, no waiting. This runs anywhere and prints the same thing every time.
from fastapi import FastAPI
from fastapi.testclient import TestClient
from pydantic import BaseModel
app = FastAPI()
BANNED = {"free money", "click here"}
class Message(BaseModel):
text: str
@app.get("/health")
def health():
# Every hosting platform pings a path like this to decide if your app is alive.
return {"status": "ok"}
@app.post("/classify")
def classify(message: Message):
lowered = message.text.lower()
spam = any(phrase in lowered for phrase in BANNED)
return {"label": "spam" if spam else "ham", "length": len(message.text)}
client = TestClient(app)
print(client.get("/health").status_code, client.get("/health").json())
print(client.post("/classify", json={"text": "Click here for FREE MONEY"}).json())
print(client.post("/classify", json={"text": "Your parcel is out for delivery"}).json())
print(client.post("/classify", json={"wrong_field": 1}).status_code)200 {'status': 'ok'}
{'label': 'spam', 'length': 25}
{'label': 'ham', 'length': 31}
422The keyword rule stands in for a real model so the example runs anywhere. Swap it for pipeline(...) from Hugging Face and nothing else about this file changes.
Read the last line
422 is FastAPI rejecting a request that did not match Message. You wrote no validation code. The Pydantic class was the validation.
That is the whole argument for declaring your input shape. A malformed request is refused at the door with a clear error, instead of reaching your model and crashing it.
The one design decision that matters most
Load the model once, at startup. Not inside the request handler.
from contextlib import asynccontextmanager
from transformers import pipeline
MODELS = {}
@asynccontextmanager
async def lifespan(app: FastAPI):
# Runs once when the process starts, before any request is served.
MODELS["clf"] = pipeline("sentiment-analysis",
model="distilbert-base-uncased-finetuned-sst-2-english")
yield
MODELS.clear()
app = FastAPI(lifespan=lifespan)
@app.post("/sentiment")
def sentiment(message: Message):
return MODELS["clf"](message.text)[0]Loading inside the handler re-reads hundreds of megabytes from disk on every single request. Response times go from 40 milliseconds to several seconds, and memory use climbs until the process is killed. This one mistake accounts for more "my API is slow" questions than everything else combined.
Running it for real
uvicorn app:app --host 0.0.0.0 --port 8000--host 0.0.0.0 means "accept connections from outside this machine". Without it, only your own computer can reach the service, which is exactly the wrong behaviour on a server and exactly the right behaviour on your laptop.
Open http://localhost:8000/docs and FastAPI gives you a clickable page for every endpoint, generated from your type hints. It is the fastest way to hand a working demo to someone.
Packaging it with Docker
A container is your code, your Python version and your libraries in one image, so the server runs what you tested.
FROM python:3.12-slim
WORKDIR /app
# Copy requirements first: this layer is cached until the file changes.
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
# Hosting platforms supply the port in an environment variable.
ENV PORT=8000
EXPOSE 8000
CMD ["sh", "-c", "uvicorn app:app --host 0.0.0.0 --port ${PORT}"]docker build -t my-model .
docker run -p 8000:8000 my-modelTwo details save real time. Copying requirements.txt before the rest means a code change does not reinstall every library. And reading ${PORT} from the environment is required by most hosting platforms, which assign the port themselves.
Free deployment, step by step
Hugging Face Spaces is the shortest path from working code to a public URL.
- Create a Space, choose the free CPU hardware and the Docker or Gradio template.
- Push your repository, including the
Dockerfile. - Wait for the build. You get a public address.
A Gradio Space is even shorter — no Dockerfile, and you get a web interface without writing HTML:
import gradio as gr
from transformers import pipeline
clf = pipeline("sentiment-analysis",
model="distilbert-base-uncased-finetuned-sst-2-english")
def predict(text):
result = clf(text)[0]
return f"{result['label']} ({result['score']:.2f})"
gr.Interface(fn=predict, inputs="text", outputs="text").launch()Put gradio and transformers in requirements.txt and push. That is a deployed model.
Render, Railway and Fly.io take a Dockerfile and give you a URL on a free or near-free tier. They sleep when idle, so expect a slow first request.
Serverless GPU — Modal, RunPod and Replicate charge per second of execution and nothing while idle. This is the honest middle ground when you genuinely need a GPU but not for 24 hours a day. Cold starts are measured in tens of seconds because the model has to load.
Common mistakes
Loading the model per request. Covered above. Check this first, every time, before optimising anything else.
No health endpoint. Platforms restart containers that fail a health check. Without one, a hung process is left serving errors forever.
Secrets in the repository. An API key in app.py pushed to a public Space is a key you must now rotate. Read keys from environment variables, always.
Ignoring memory limits. Free tiers give you a fixed amount of RAM. A model that loads fine on your 16 GB laptop is killed instantly on a 512 MB tier, and the log shows nothing useful. Check the tier's limit against your model size before you build.
No request timeout. A single slow request holds a worker forever, and enough of them take the whole service down. Set a timeout, and return 503 when overloaded rather than queueing without bound.
Testing only the happy path. Send an empty string, a 100,000-character string, emoji, and a wrong field. Fix what breaks now, not after somebody finds it.
Serving a large language model yourself by default. If you find yourself renting a GPU for a project with ten users, stop and price an API call instead. It is usually a hundred times cheaper.
Try it yourself
Deploy app.py to a free Space. Then add a /stats endpoint returning how many requests each label has produced since startup, and confirm the count resets when the Space sleeps and wakes.
That reset is the lesson. It is why anything you need to keep must go in a database, not in a Python variable.
What to learn next
- Docker for ML — the packaging step, in full.
- Model serving — throughput, batching and scaling in depth.
- Monitoring and model drift — noticing when a deployed model goes quietly wrong.
Researcher — Mathematics and papers.
Deciding what to host
The question is not technical. It is a cost model with three terms: fixed cost of idle capacity, marginal cost per request, and the cost of latency violations.
monthly_cost = hours_up × hourly_rate + requests × marginal_costFor self-hosting on a dedicated GPU the second term is near zero and the first dominates. For a per-token API the reverse holds. The break-even request volume is:
break_even_requests = (hours_up × hourly_rate) / api_cost_per_requestA cloud GPU at roughly $0.50/hour run continuously is about $360/month. Against a hosted small model at a fraction of a cent per request, break-even sits in the millions of requests per month. Below that, self-hosting an LLM is a decision about data residency, latency floor or model customisation — not about cost. State which one, or do not do it.
Distinguish this sharply from hosting a small discriminative model. A quantised classifier or embedding model on CPU costs single-digit dollars per month and has no comparable API alternative at low volume.
Latency budgets
Report distributions, never means. Response time is right-skewed and the mean hides the tail that users actually experience.
Decompose the budget:
total = network_RTT + queue_wait + preprocess + model_forward + postprocessFor generative models the relevant metrics are different again:
- TTFT (time to first token) — dominated by prefill, scales with prompt length.
- TPOT (time per output token) — dominated by memory bandwidth, roughly constant per token.
- End-to-end = TTFT + TPOT × output_length.
Streaming changes perceived latency without changing end-to-end latency. Optimise TTFT for interactive interfaces and total throughput for batch pipelines; these pull in opposite directions, which is why one deployment rarely serves both well.
Queue wait is where p99 goes wrong. From Kingman's approximation, waiting time grows as ρ/(1−ρ) with utilisation ρ, so a service run at 90% utilisation has roughly nine times the queueing delay of one at 50%. Provisioning on mean utilisation guarantees tail-latency failures.
CPU inference is a real option, and underused
For models under roughly 500M parameters, CPU serving with the right runtime is frequently sufficient and an order of magnitude cheaper.
- ONNX Runtime — graph-level fusion and constant folding, plus INT8 dynamic quantisation. Typical result on transformer encoders: 2–4× throughput versus eager PyTorch, with under 1% accuracy loss.
- OpenVINO — similar, tuned for Intel hardware.
- Distillation — a 6-layer distilled encoder retains most task accuracy at a fraction of the cost (Sanh et al., 2019).
Measure before assuming a GPU is needed. Many production text classifiers should never have been on one.
Scale-to-zero and cold starts
Serverless GPU platforms eliminate idle cost, and the price is cold-start latency. The dominant term is weight loading, not container start:
cold_start ≈ container_pull + runtime_init + weight_load_from_storageWeight load is bounded by storage bandwidth. Mitigations in order of effect: keep weights in a memory-mapped format like safetensors so loading is page faults rather than deserialisation; cache the image and weights on local NVMe rather than object storage; keep one warm instance if the SLO cannot absorb the tail.
The economics are straightforward. Scale-to-zero wins when duty cycle is below roughly 20% and the p99 latency budget can absorb the cold start. Otherwise keep a floor of one instance.
Release strategy
Model updates are not code updates. A model change alters the output distribution without changing any interface, so tests pass and behaviour shifts.
- Shadow deployment — send production traffic to the new model, discard its output, compare offline. The only zero-risk way to measure a new model on real traffic.
- Canary — a small traffic percentage, with automatic rollback on metric regression.
- Blue-green — full switch with instant rollback. Doubles capacity cost during the window.
Pin the model version alongside the code version and log both with every prediction. Without that pairing, an incident investigation cannot determine which model produced a given output.
What to monitor
Split into three tiers, because they fail on different timescales.
- Infrastructure — latency percentiles, error rate by class, saturation, GPU utilisation and memory. Minutes.
- Model — input-distribution drift (population stability index, or KS on continuous features), prediction-distribution drift, confidence histograms, abstention rate. Days.
- Business — the outcome the model exists to change, and human override rate. Weeks.
Ground-truth labels usually arrive with delay, and often never for the cases where the model was wrong and nobody complained. That feedback asymmetry is why input drift is monitored as a leading indicator: it is available immediately, whereas accuracy is not. See monitoring and drift.
References
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — the canonical account of why the model is the small part.
- Sanh et al., DistilBERT, 2019 — arxiv.org/abs/1910.01108
- Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention, SOSP 2023 — arxiv.org/abs/2309.06180
- Beyer et al., Site Reliability Engineering, O'Reilly 2016 — chapters on SLOs and error budgets.
- Breck et al., The ML Test Score, IEEE Big Data 2017 — a concrete production-readiness rubric.
What to learn next
- Docker for ML — the packaging step, in full.
- Model serving — throughput, batching and scaling in depth.
- Monitoring and model drift — noticing when a deployed model goes quietly wrong.