Health checks and readiness probes
Liveness and readiness are two different questions a server must answer separately, so traffic-routing systems never send real requests to an instance that is alive but not yet able to serve.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A server needs two different health signals. Is the process alive at all? And is it actually able to serve a real request right now?
The analogy you have already lived
A shop can have its shutter open and lights on — visibly occupied, not shut down. Yet the staff can still be restocking shelves, not yet ready to serve a customer. "Open" and "ready to serve you" are two different facts, true at different times.
A server is the same. It can be running, and still not ready.
Why it exists
Model serving introduced a single health check: "yes, I am alive and my model is loaded". That is a reasonable simplification for a small system. At real production scale, a server needs to separate two distinct questions.
Liveness — is the process itself alive, or has it hung, deadlocked, or crashed and needs restarting? Readiness — can this instance actually handle a real request right now, or is it still starting up, or temporarily overloaded?
A system that only asks the first question sends real traffic to an instance that is technically alive but not yet ready. During a cold start, for example, those requests then fail.
How it works
/healthz (liveness) -> "is the process itself alive?"
answers "ok" almost immediately, even mid-startup
/readyz (readiness) -> "can I actually serve a real prediction right now?"
answers "not ready" until the model is fully loadedThe system routing traffic — a load balancer, Kubernetes — uses readiness to decide who gets real requests. It uses liveness to decide who needs restarting.
A real example you have seen
A food delivery app showing a restaurant as "open" the moment its shift starts. That is true even while the kitchen is still setting up and cannot yet accept new orders. "Open" and "accepting orders" are tracked and shown as separate states.
Remember this
- Liveness asks "is the process alive"; readiness asks "can it serve a real request right now" — these are different questions.
- An instance can be alive but not ready, especially during startup.
- Traffic should only be routed to instances that report ready, not only alive.
What to learn next
- Warm-up and cold starts — exactly the window a readiness probe exists to hide from real traffic.
- Serving a model with FastAPI — the concurrency model these probe endpoints have to share the server with, without being starved out under load.
- Model serving — the single combined health check this lesson splits into two, and why that split matters at production scale.
Developer — Code and libraries.
Setup
pip install fastapi "uvicorn[standard]" httpxTwo endpoints, answering two different questions
import threading
import time
from fastapi import FastAPI
from fastapi.responses import JSONResponse
from fastapi.testclient import TestClient
app = FastAPI()
STATE = {"model_loaded": False}
def load_model_slowly():
time.sleep(0.3) # standing in for reading a large file from disk
STATE["model_loaded"] = True
threading.Thread(target=load_model_slowly, daemon=True).start()
@app.get("/healthz")
def healthz():
# Liveness: is the process itself alive and able to answer at all?
# This should stay "ok" even while the model is still loading.
return {"status": "ok"}
@app.get("/readyz")
def readyz():
# Readiness: is this instance actually able to serve a real prediction?
if not STATE["model_loaded"]:
return JSONResponse(status_code=503, content={"status": "loading"})
return {"status": "ready"}
client = TestClient(app)
print("immediately after startup:")
print(" /healthz ->", client.get("/healthz").status_code, client.get("/healthz").json())
print(" /readyz ->", client.get("/readyz").status_code, client.get("/readyz").json())
time.sleep(0.4)
print("\nafter the model finishes loading:")
print(" /healthz ->", client.get("/healthz").status_code, client.get("/healthz").json())
print(" /readyz ->", client.get("/readyz").status_code, client.get("/readyz").json())immediately after startup:
/healthz -> 200 {'status': 'ok'}
/readyz -> 503 {'status': 'loading'}
after the model finishes loading:
/healthz -> 200 {'status': 'ok'}
/readyz -> 200 {'status': 'ready'}Line-by-line walkthrough
/healthz never checks STATE["model_loaded"] — it answers ok the instant the process can respond at all, which is exactly what a liveness check should mean: "do not restart me, I am not hung".
/readyz checks the one thing that matters for accepting real traffic, and returns 503 (service unavailable) until it is true. A 503 here is not a failure to be alarmed about — it is the correct, expected answer during the small window every fresh instance passes through.
Common mistakes
Using one endpoint for both purposes. If your only health check also verifies the model is loaded, an orchestrator using it for liveness will keep restarting a perfectly healthy instance that is still starting up — a restart loop caused by the health check itself, not any real problem.
Making /readyz too expensive. A readiness check that runs a full prediction on every poll adds real load, multiplied by how often it is polled by every instance. Check the cheap thing that implies readiness — "is the model object loaded" — not a full round-trip through it.
No readiness check for downstream dependencies. If your server depends on a low-latency feature lookup service being reachable, readiness should reflect that too — a server with a loaded model but no reachable feature store is not genuinely ready to answer real requests correctly.
Treating "ready once" as "ready forever". A healthy instance can become unready later — a downstream dependency going down, memory pressure. /readyz should be checked continuously by the routing system, not only once at startup.
Try it yourself
Add a third state: after the model loads, simulate it becoming temporarily unready again (a downstream dependency failing) by flipping STATE["model_loaded"] back to False for two seconds, then true again. Confirm /readyz reflects the change immediately while /healthz stays ok throughout.
What to learn next
- Warm-up and cold starts — exactly the window a readiness probe exists to hide from real traffic.
- Serving a model with FastAPI — the concurrency model these probe endpoints have to share the server with, without being starved out under load.
- Model serving — the single combined health check this lesson splits into two, and why that split matters at production scale.
Researcher — Mathematics and papers.
The Kubernetes probe model, in full
Kubernetes formalises three probe types, though only two are covered above:
- Liveness probe — failure triggers a container restart. Should reflect only "is this process fundamentally broken (hung, deadlocked)", since a false positive here causes unnecessary, potentially cascading restarts.
- Readiness probe — failure removes the instance from the service's load-balanced endpoint set, without restarting it. Should reflect genuine capacity to serve, and can legitimately flap between ready and not-ready under normal operation (temporary overload, a dependency blip).
- Startup probe — specifically for slow-starting containers, delaying when liveness and readiness probes begin at all, preventing a legitimately slow (but healthy) startup from being killed by an impatient liveness probe before it ever gets the chance to finish.
Conflating liveness and readiness into one check is a common production misconfiguration, and its specific failure mode is a restart loop: an instance takes a while to become ready, a liveness probe not distinguished from readiness kills it for "failing", the replacement instance faces the identical startup time, and repeats — the instance may never successfully start under real load.
Probe configuration as a latency-versus-stability trade-off
Probe interval, timeout, failure threshold and success threshold jointly determine how quickly the system reacts to a real state change versus how much it protects against a noisy, transient failure being over-interpreted. A short interval with a low failure threshold reacts fast but risks flapping an instance in and out of service over brief blips; a long interval with a high threshold is stable but slow to detect a genuine failure, extending the window where traffic keeps hitting a broken instance.
Graceful shutdown as readiness's counterpart
Readiness governs traffic entering an instance; graceful shutdown governs traffic already in flight when an instance is told to stop. The correct sequence on shutdown: mark /readyz unready first (stopping new traffic from routing there), wait for the load balancer to notice and stop sending new requests, allow in-flight requests to complete, then terminate the process. Skipping the readiness-first step and terminating immediately drops in-flight requests that had already been routed there, an entirely avoidable source of errors during ordinary, planned deploys.
Papers and systems
- The Kubernetes documentation on liveness, readiness and startup probes is the authoritative and most current reference for exact probe semantics: kubernetes.io/docs/concepts/configuration/liveness-readiness-startup-probes
- Google's SRE book (Beyer et al., Site Reliability Engineering, 2016), Chapter 22, covers graceful degradation and the broader health-signalling design this lesson's two endpoints are a minimal instance of.
- Fowler, HealthCheck, martinfowler.com — an accessible treatment of health-check design patterns predating and generalising beyond Kubernetes specifically.
What to learn next
- Warm-up and cold starts — exactly the window a readiness probe exists to hide from real traffic.
- Serving a model with FastAPI — the concurrency model these probe endpoints have to share the server with, without being starved out under load.
- Model serving — the single combined health check this lesson splits into two, and why that split matters at production scale.