Ray Serve
Ray Serve runs several copies of a model across many processes or machines behind one address, scaling a service by adding replicas instead of rewriting it.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Ray Serve runs several copies of your model across many processes or machines, behind one single address.
The analogy you have already lived
Think of a busy dosa stall with one cook. However fast that one cook works, there is a ceiling on how many dosas can come out per minute. A second cook, working the same recipe at a second tawa, roughly doubles that ceiling. Customers still queue at one counter — they never need to know or care which cook made theirs.
Ray Serve is the system that manages a whole row of cooks working the same recipe. It directs each new order to whichever one is free.
Why it exists
Serving a model with FastAPI runs as one process. One process has a ceiling on how many requests it can genuinely handle at once. That is true however well the concurrency inside it is written.
Ray Serve is built on Ray, a system for running Python code across many processes and machines. It lets you say "run 3 copies of this model". It then routes incoming requests across them, adding or removing copies as load changes.
How it works
requests arrive
|
v
Ray Serve's router
|
+----> replica 1 (a copy of your model)
+----> replica 2 (a copy of your model)
| (busy replicas get fewer new requests)
v
answer sent backA replica is one running copy of your model, doing the same job as every other replica. Adding replicas raises the ceiling on how many requests can be handled at the same time.
A real example you have seen
A large video platform's recommendation service needs to answer requests from millions of users at once. No single machine could do that alone. Many identical copies run behind one system that spreads the load across them instead.
Remember this
- Ray Serve runs multiple replicas of a model, spreading load across them automatically.
- Scaling means adding replicas, not rewriting the service.
- It is built on Ray, a broader system for running Python across many machines.
What to learn next
- Serving many models on one machine — the single-process version of the scaling problem Ray Serve solves across many.
- BentoML — a lighter-weight alternative when a full cluster is more than a service needs.
- Text Generation Inference (TGI) — a server specialised for one specific, very demanding workload: LLM token generation.
Developer — Code and libraries.
Setup
pip install "ray[serve]" scikit-learn numpy requestsTwo replicas, one address
import numpy as np
from sklearn.linear_model import LogisticRegression
from ray import serve
import requests
rng = np.random.RandomState(0)
X = rng.uniform(0, 10, (200, 2))
y = (X[:, 0] + X[:, 1] > 10).astype(int)
clf = LogisticRegression().fit(X, y)
@serve.deployment(num_replicas=2)
class RiskScorer:
def __init__(self, model):
self.model = model
async def __call__(self, request):
data = await request.json()
row = np.array([[data["income"], data["years"]]])
prob = float(self.model.predict_proba(row)[0, 1])
return {"approved": prob >= 0.5, "probability": round(prob, 4)}
app = RiskScorer.bind(clf)
serve.run(app)
r1 = requests.post("http://127.0.0.1:8000/", json={"income": 6.0, "years": 5.0})
r2 = requests.post("http://127.0.0.1:8000/", json={"income": 1.0, "years": 1.0})
print("strong applicant:", r1.json())
print("weak applicant :", r2.json())
serve.shutdown()strong applicant: {'approved': True, 'probability': 0.9253}
weak applicant : {'approved': False, 'probability': 0.0}Ray also prints a substantial block of its own startup and shutdown logs (worker process IDs, dashboard URL, per-replica status) to the terminal — that log content varies by Ray version and is not reproduced here; only the two print results above are the program's own output.
Line-by-line walkthrough
@serve.deployment(num_replicas=2) is the entire scaling decision — two running copies of RiskScorer, both loaded with the same trained model, both able to answer a request independently.
RiskScorer.bind(clf) builds the deployment graph without running it yet; serve.run(app) actually starts the replicas and the router in front of them, all reachable at the one address http://127.0.0.1:8000/, exactly as promised in the beginner section — the caller never chooses which replica answers.
Common mistakes
Setting num_replicas far above what the model or the machine can actually support. More replicas than CPU cores available buys contention, not more real throughput. Match replica count to real capacity, and measure.
Loading the model inside __call__ instead of __init__. The same "load once" rule from model serving applies per replica here — __init__ runs once per replica at startup, __call__ runs on every request.
Treating num_replicas as a fixed number forever. Ray Serve supports autoscaling replica count based on load; a fixed number is the simplest starting point, not the production end state for traffic that varies through the day.
Try it yourself
Change num_replicas to 1, re-run, and time 50 requests sent one after another versus 50 sent concurrently (using a small thread pool, as in gRPC vs REST for inference's timing style). Then set it back to a higher number and compare.
What to learn next
- Serving many models on one machine — the single-process version of the scaling problem Ray Serve solves across many.
- BentoML — a lighter-weight alternative when a full cluster is more than a service needs.
- Text Generation Inference (TGI) — a server specialised for one specific, very demanding workload: LLM token generation.
Researcher — Mathematics and papers.
Ray's actor model underneath Serve
Ray Serve deployments are built on Ray's actor abstraction: each replica is a long-lived process holding state (the loaded model) between calls, scheduled and supervised by Ray's own cluster scheduler. This differs from a stateless function-as-a-service model in the same way model serving's "load once" server differs from a cold function invocation — the state (the model) survives across many calls to the same replica.
Ray's scheduler places replicas across available CPU or GPU resources in a cluster, which is what lets num_replicas scale from a laptop's core count to a many-machine cluster with the same deployment code, unlike a single-process framework where scaling requires an external orchestrator (Kubernetes) managing separate containers.
Autoscaling and request routing
Ray Serve supports an autoscaling replica count, configured with minimum and maximum bounds and a target ongoing-requests-per-replica metric, adjusting the running replica count as measured load changes. Routing between replicas defaults to a power-of-two-choices load-balancing policy: sample two replicas at random, route to whichever currently has fewer in-flight requests — a well-studied middle ground between pure round-robin (ignores actual load) and checking every replica's load on every request (expensive at scale).
Composing multiple models as a graph
Beyond scaling one model, Ray Serve supports deployment graphs: multiple deployments (potentially different models, or pre/post-processing steps) composed together, each independently scaled, with Ray handling the data flow between them. This directly extends the single-model case here toward the multi-stage pipelines common in real recommendation and ranking systems, where a candidate-generation model, a ranking model and a business-rules step often run as separate, independently-scaled services.
Comparison with the other servers in this section
Ray Serve's differentiator relative to Triton and BentoML is native multi-machine scaling and arbitrary Python composition, at the cost of a heavier runtime dependency (a full Ray cluster) than either a single-process FastAPI server or a Triton container needs. Triton remains the stronger choice for pure GPU-batched inference throughput on fixed hardware; Ray Serve for services that need to compose several models or scale elastically across a cluster.
Papers and systems
- Moritz et al., Ray: A Distributed Framework for Emerging AI Applications, OSDI 2018 — the underlying actor and task scheduling system Ray Serve is built on.
- Mitzenmacher, The Power of Two Choices in Randomized Load Balancing, 2001 — the theoretical basis for the routing policy described above.
- Ray Serve's own autoscaling documentation is the current, authoritative source for its scaling configuration: docs.ray.io/en/latest/serve
What to learn next
- Serving many models on one machine — the single-process version of the scaling problem Ray Serve solves across many.
- BentoML — a lighter-weight alternative when a full cluster is more than a service needs.
- Text Generation Inference (TGI) — a server specialised for one specific, very demanding workload: LLM token generation.