KServe and Seldon
KServe and Seldon are Kubernetes platforms that turn everything built by hand earlier in this section — canaries, versioned endpoints, rollback — into a few lines of YAML, at the cost of needing a real cluster to run on.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
KServe and Seldon are ready-made platforms that run and route traffic between model versions on Kubernetes. They exist so you do not have to build that routing logic yourself.
The analogy you have already lived
You have used a building's lift instead of climbing a rope up the lift shaft yourself. The rope-climbing approach works — you would eventually get to the fifth floor. But the lift makes the same journey, engineered once, by people who thought hard about pulleys and safety brakes. Every resident afterward only presses a button.
Every earlier lesson in this section — shadow deployment, canary releases, blue-green deploys — was climbing the rope by hand. That was on purpose, so the mechanism stays fully visible. KServe and Seldon are the lift: the same ideas, engineered once by a team that runs them for thousands of companies. You configure them with a short YAML file, instead of a custom router you maintain yourself.
Why it exists
The routers built earlier in this section were small enough to write in fifty lines. That was only possible because they ran on one machine, handling one process. A real production model usually runs across many machines. It needs its containers scheduled and restarted automatically, and its traffic load-balanced across replicas. All of that has to keep working while a canary rollout, a rollback, or a version switch is happening on top of it.
Building that infrastructure yourself is a large, ongoing engineering project having nothing to do with machine learning. KServe and Seldon are two of the more established open-source answers. Both run on Kubernetes — a system for running and managing containers across many machines. Both let you describe what you want (route 10% of traffic to a new model version) in a short configuration file, while the platform handles how.
How it works
you write a short YAML file describing the model and the traffic split
|
v
kubectl apply -f my-model.yaml
|
v
KServe / Seldon creates and manages the actual running containers,
load balancing, health checks, and traffic-splitting for you
|
v
you change ONE number in that file to shift traffic --
the platform re-applies it, live, without you writing any routing codeA real example you have seen
A large online store running dozens of recommendation and fraud models at once is a realistic case for a platform like this. Nobody at that scale is hand-writing a router class per model, the way this section did for teaching purposes. A platform like KServe or Seldon exists specifically so a team can add the fortieth model without adding fortieth version of the same routing logic.
Remember this
- KServe and Seldon do on Kubernetes, as YAML, the same things this section built earlier by hand, as Python.
- They exist because real production serving needs infrastructure — scheduling, scaling, load balancing — that is a large engineering effort on its own.
- This lesson's code runs entirely on your own machine and illustrates the concept the YAML configures — actually running the YAML itself needs a real Kubernetes cluster.
What to learn next
- Scheduling GPUs on Kubernetes — the layer underneath either platform when your model actually needs a GPU, not only CPU replicas.
- Switching between model providers — a related but distinct migration, where what changes is not your own model's version but the company serving it.
- Health checks and readiness probes — the mechanism Kubernetes itself uses to decide whether a pod deployed by either platform is safe to send traffic to.
Developer — Code and libraries.
The honest constraint on this lesson
KServe and Seldon are Kubernetes-native systems — the YAML below only does anything on a real cluster (a managed one, or a local one like kind or minikube). Nothing in this lesson can be run as a plain Python script the way earlier lessons in this section were. What follows instead: real, verified YAML showing exactly what each platform's configuration looks like, and a small, genuinely runnable Python simulation of the underlying routing idea — kept visibly separate from each other.
KServe: a canary rollout in three edits
KServe wraps each model in an InferenceService. A canary rollout is one extra field.
apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
name: "loan-scorer"
spec:
predictor:
model:
modelFormat:
name: sklearn
storageUri: "gs://my-bucket/models/loan-scorer/v1"apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
name: "loan-scorer"
spec:
predictor:
model:
modelFormat:
name: sklearn
storageUri: "gs://my-bucket/models/loan-scorer/v2"
canaryTrafficPercent: 10kubectl apply -f 2-canary.yaml
kubectl get isvc loan-scorerstorageUri now points at v2, and canaryTrafficPercent: 10 tells KServe to send 10% of traffic to it while the other 90% keeps going to the last known-good revision — the exact ramp from canary releases for models, expressed as one field instead of a hand-written hash router.
apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
name: "loan-scorer"
spec:
predictor:
model:
modelFormat:
name: sklearn
storageUri: "gs://my-bucket/models/loan-scorer/v2"
# canaryTrafficPercent removed -- v2 is now the only version canaryTrafficPercent: 0Promotion and rollback, in this scheme, are both the same one-line edit to the same file — remove the field to finish the rollout, or set it to 0 to abandon it. This is the KServe equivalent of the atomic switch built by hand in rolling back a model.
Seldon: the same idea, expressed as a list of predictors
Seldon Core's SeldonDeployment names each version as its own predictor, each with its own traffic weight.
apiVersion: machinelearning.seldon.io/v1alpha2
kind: SeldonDeployment
metadata:
name: loan-scorer
spec:
name: loan-scorer-deployment
predictors:
- name: model-v1
replicas: 3
traffic: 90
componentSpecs:
- spec:
containers:
- name: classifier-v1
image: myregistry/loan-scorer:v1
graph:
name: classifier-v1
type: MODEL
endpoint:
type: REST
- name: model-v2
replicas: 1
traffic: 10
componentSpecs:
- spec:
containers:
- name: classifier-v2
image: myregistry/loan-scorer:v2
graph:
name: classifier-v2
type: MODEL
endpoint:
type: RESTThe traffic: 90 and traffic: 10 fields are required to add up to 100 across all predictors in the list — Seldon rejects a SeldonDeployment where they do not. Each predictor also declares its own replicas, so a canary can genuinely run on fewer machines than the stable version, matching its smaller share of traffic.
A real, runnable piece: how the traffic split actually behaves
Both platforms implement their percentage split using the underlying Kubernetes networking layer (Istio or Knative), which routes per request, at random, weighted by the percentage — not by hashing a stable user identifier, the way the hand-built canary router in canary releases for models did. This is a real and easy-to-miss difference, and it is worth seeing concretely.
"""Simulates what a service-mesh weighted route (the mechanism underneath
KServe's canaryTrafficPercent and Seldon's per-predictor `traffic` field)
actually does: split PER REQUEST, at random, by weight -- not sticky per
user. This illustrates the routing CONCEPT the YAML above configures;
running the YAML itself needs a real cluster.
"""
import random
def weighted_route(canary_percent: float, rng: random.Random) -> str:
"""One call = one independent coin flip, weighted by canary_percent."""
return "canary" if rng.uniform(0, 100) < canary_percent else "stable"
rng = random.Random(0)
routes = [weighted_route(10, rng) for _ in range(20_000)]
canary_count = routes.count("canary")
print(f"requested canary: 10% measured canary: {canary_count/len(routes)*100:.2f}% ({canary_count}/{len(routes)})")
# the same "user" issuing 10 separate requests -- NOT sticky
rng2 = random.Random(1)
same_user_routes = [weighted_route(10, rng2) for _ in range(10)]
print("one user, 10 separate requests, at 10% canary:", same_user_routes)requested canary: 10% measured canary: 10.10% (2019/20000) one user, 10 separate requests, at 10% canary: ['stable', 'stable', 'stable', 'stable', 'stable', 'stable', 'stable', 'stable', 'canary', 'canary']
The measured split lands close to the requested 10%, exactly as the hand-built hash router did — but the same simulated user landed on stable for eight calls and canary for two, in no particular pattern. A real user making several requests through a default KServe or Seldon canary could genuinely see both versions across those requests. If your use case needs the same user to consistently see the same version — which the earlier canary lesson argued for directly — you need to explicitly configure session affinity (a consistent-hash load-balancing policy) on top of the platform's default weighted routing; it is not the default behaviour.
import random
from weighted_routing import weighted_route
def test_zero_percent_never_routes_to_canary():
rng = random.Random(0)
assert all(weighted_route(0, rng) == "stable" for _ in range(500))
def test_measured_split_is_close_to_the_requested_weight():
rng = random.Random(0)
routes = [weighted_route(25, rng) for _ in range(20_000)]
measured = routes.count("canary") / len(routes) * 100
assert 23.0 <= measured <= 27.0pytest test_weighted_routing.py -q.. [100%] 2 passed in 0.02s
What these platforms add beyond routing
Both KServe and Seldon are considerably more than a traffic splitter. KServe integrates with Knative for scale-to-zero and autoscaling on request concurrency, supports many model formats (scikit-learn, TensorFlow, PyTorch, ONNX, and custom containers) behind an identical API shape, and includes built-in explainability and drift-detection components. Seldon similarly supports arbitrary inference graphs — chaining a preprocessor, a model, and an explainer as separate steps — not only a flat list of model versions.
Common mistakes
Treating percentage traffic splitting as sticky by default. As demonstrated above, it is not, on either platform, without extra configuration. Building a user-facing feature that assumes consistency across a session on the default setup is a real, subtle bug.
Deploying without resource requests and limits set. The resources block omitted from the shortened examples above for readability is not optional in practice — without it, Kubernetes cannot schedule your model sensibly alongside everything else on the cluster, and a runaway model can starve its neighbours.
Skipping local testing entirely because "it needs a cluster". The routing and contract logic — the exact things this whole section built and tested by hand in plain Python — can and should still be tested exactly as shown throughout this section, before anything is wrapped in a SeldonDeployment or InferenceService at all. The platform changes how a tested idea gets deployed, not whether it needs testing.
Assuming Seldon's traffic percentages are validated for you at write time. They are checked when the resource is applied to the cluster, not by any local tool — a typo that makes the weights sum to 90 instead of 100 is only caught once you actually kubectl apply it.
Try it yourself
If you have access to minikube or kind, install KServe's quickstart environment and apply the three YAML files above in order, watching kubectl get isvc loan-scorer after each one. Seeing the PREV and LATEST traffic columns change in real output, on a real (if local) cluster, is worth doing once — it is the same concept this lesson's Python simulation illustrates, running for real.
What to learn next
- Scheduling GPUs on Kubernetes — the layer underneath either platform when your model actually needs a GPU, not only CPU replicas.
- Switching between model providers — a related but distinct migration, where what changes is not your own model's version but the company serving it.
- Health checks and readiness probes — the mechanism Kubernetes itself uses to decide whether a pod deployed by either platform is safe to send traffic to.
Researcher — Mathematics and papers.
What Kubernetes actually provides underneath both platforms
Both InferenceService and SeldonDeployment are Custom Resource Definitions (CRDs) — Kubernetes' mechanism for extending its API with new kinds of objects, reconciled by a controller that continuously compares the desired state (your YAML) against the actual state (running pods, service mesh routing rules) and takes action to close any gap. This reconciliation loop is the general Kubernetes operator pattern; KServe's controller and Seldon's operator are both instances of it, specialised for model serving.
Where the traffic-splitting mechanism actually lives
KServe's percentage-based canary is implemented via Knative Serving's revision-based routing, itself backed by Istio's (or a compatible mesh's) VirtualService weighted destination rules. Weighted routing in Istio operates at the level of independent request-routing decisions by default; consistent-hash-based session affinity is available as an explicit trafficPolicy.loadBalancer.consistentHash configuration, not the default, which is the origin of the non-sticky behaviour demonstrated in the developer section. Seldon Core's traffic splitting, when deployed with an Istio integration, follows the same underlying mechanism; its non-Istio ambassador/default routing similarly performs independent per-request weighted selection.
Explainability and drift detection as first-class citizens
A meaningful difference from the routers built earlier in this section: both platforms treat model explainability (e.g., KServe's integration with Alibi Explain, Seldon's own explainer components) and outlier/drift detection as deployable components in the same graph as the model itself, rather than a separate monitoring system bolted on afterward. This connects the deployment mechanics in this lesson directly to measuring data drift — on these platforms, a drift detector is a peer service in the same inference graph as the model, not an external batch job.
Choosing between a hand-rolled router and a platform
The trade-off is not purely technical. A hand-rolled router (as built throughout this section) is transparent, dependency-free, and easy to reason about precisely because it is small — appropriate for a single service, a small team, or a case where the exact routing semantics (like sticky assignment) matter more than convenience. A platform like KServe or Seldon amortises real engineering effort (autoscaling, multi-format model support, explainability, drift detection) across every model a team deploys, at the cost of a genuine new dependency: a Kubernetes cluster, its own operational burden, and — as shown above — default behaviour (non-sticky routing) that may not match what a hand-rolled version was built to guarantee.
Further reading
- KServe documentation, Canary Rollout Strategy — kserve.github.io/website/docs/model-serving/predictive-inference/rollout-strategies/canary
- Seldon Core documentation, Inference Graphs — docs.seldon.ai/seldon-core-1/v1.19/configuration/routing/inference-graph
- Burns et al., Kubernetes: Up and Running, O'Reilly — the operator and CRD pattern both platforms are built on.
What to learn next
- Scheduling GPUs on Kubernetes — the layer underneath either platform when your model actually needs a GPU, not only CPU replicas.
- Switching between model providers — a related but distinct migration, where what changes is not your own model's version but the company serving it.
- Health checks and readiness probes — the mechanism Kubernetes itself uses to decide whether a pod deployed by either platform is safe to send traffic to.