BentoML
BentoML is a Python framework that turns a trained model into a packaged, servable API with far less boilerplate than writing every route and Dockerfile by hand.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
BentoML is a Python tool that packages a trained model into a ready-to-serve API, handling most of the boilerplate for you.
The analogy you have already lived
Serving a model with FastAPI is like building your own tiffin box from scratch. You cut the metal, weld the compartments, and it works exactly as you designed it. BentoML is closer to buying a well-designed tiffin box instead. The compartments, the latch and the carry handle already exist, sized for the standard case. You still choose what goes in it.
Why it exists
Writing a FastAPI server by hand means deciding, every time, how to load the model and structure the response. It also means building a Docker image and writing the deployment config. Most of those decisions are the same for most projects.
BentoML standardises them. Define a small "service" class around your model, and it handles the API layer, request validation, and building a deployable package. A team is then not reinventing the same server shape on every project.
How it works
your trained model (any framework: sklearn, PyTorch, XGBoost...)
|
v
save it into BentoML's model store (like git, but for model files)
|
v
write a small Service class: what does "predict" do?
|
v
bentoml serve -> a running API, with auto-generated docs
bentoml build -> a portable, deployable packageThe service class is the only genuinely new thing to write. Everything else — the web server, the package format — is BentoML's job.
A real example you have seen
Many small and mid-sized companies use frameworks exactly like this. They go from "a data scientist has a trained model" to "there is a working API" in an afternoon. No separate infrastructure engineer needs to write a server by hand for every model.
Remember this
- BentoML packages a model into a servable API with far less boilerplate than a hand-built server.
- You write one small service class; BentoML builds the API and the deployable package around it.
- It works with most popular ML frameworks, not only one.
What to learn next
- Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
- Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
- Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.
Developer — Code and libraries.
Setup
pip install bentoml scikit-learn numpySaving a model and serving it
import numpy as np
from sklearn.linear_model import LogisticRegression
import bentoml
from bentoml.models import BentoModel
rng = np.random.RandomState(0)
X = rng.uniform(0, 10, (200, 2))
y = (X[:, 0] + X[:, 1] > 10).astype(int)
clf = LogisticRegression().fit(X, y)
# Saves the trained model into BentoML's local model store, tagged and
# versioned automatically -- the same idea as the registry in
# [versioning and retiring features](/learn/production-feature-pipelines/feature-versioning-and-deprecation), for model artifacts instead of features.
saved = bentoml.sklearn.save_model("risk_model", clf)
print("saved tag:", saved.tag)
@bentoml.service(resources={"cpu": "1"})
class RiskScorer:
model_ref = BentoModel("risk_model:latest")
def __init__(self):
self.model = bentoml.sklearn.load_model(self.model_ref)
@bentoml.api
def predict(self, income: float, years: float) -> dict:
row = np.array([[income, years]])
prob = float(self.model.predict_proba(row)[0, 1])
return {"approved": prob >= 0.5, "probability": round(prob, 4)}
# During development, instantiate the service directly and call its method --
# no server, no port, the same pattern used for FastAPI's TestClient in
# [model serving](/learn/mlops/model-serving).
svc = RiskScorer()
print(svc.predict(income=6.0, years=5.0))
print(svc.predict(income=1.0, years=1.0))saved tag: risk_model:3brrjxngtckp74aj
{'approved': True, 'probability': 0.9253}
{'approved': False, 'probability': 0.0}The exact tag after risk_model: is a random identifier BentoML generates per save — yours will differ every run. That is deliberate: it is how BentoML tells two versions of the same model apart.
Line-by-line walkthrough
bentoml.sklearn.save_model writes the model to a local store, distinct from a plain joblib.dump — it also records metadata (framework, save time, a unique tag) that the rest of BentoML relies on.
@bentoml.service turns a plain Python class into something BentoML can serve. @bentoml.api marks which method is a callable endpoint. Instantiating RiskScorer() directly, as done here, runs your model logic with zero networking involved — useful for fast local iteration before running a real server.
To actually run it as a live HTTP API, on the command line:
bentoml serve bento_service:RiskScorerNo output block for that command — like uvicorn in model serving, it prints a live startup banner that varies by version and is not meaningful as static text.
Common mistakes
Loading the model inside predict instead of __init__. This repeats the "load every request" mistake from model serving, inside a different framework. Load once, in __init__, always.
Referencing BentoModel("risk_model:latest") in a class attribute, then never re-saving after retraining. :latest resolves at class-definition time in some setups — pin to the exact tag you validated in anything you deploy, and treat :latest as a development convenience only.
Skipping BentoML's own validation (Pydantic-style type hints on predict) and re-implementing manual checks. The type hints on predict's parameters are not decoration — BentoML uses them to validate incoming requests automatically, the same job model serving's Applicant model does for FastAPI.
Try it yourself
Add a second method to RiskScorer, predict_batch(self, rows: list[dict]) -> list[dict], that scores several applicants in one call. Compare calling it once with 50 rows against calling predict 50 times, the same batching argument made in model serving's exercise.
What to learn next
- Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
- Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
- Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.
Researcher — Mathematics and papers.
What a "framework" actually standardises
BentoML's value is not the inference call itself — model.predict_proba(row) is identical whether wrapped in Flask, FastAPI, or BentoML. What it standardises is everything around that call: a model store with content-addressed versioning, automatic request-schema generation from Python type hints, a build system that produces a runnable, versioned artifact (a "Bento"), and adapters targeting several deployment backends (Docker, Kubernetes via Yatai, or several managed cloud runtimes) from that one artifact.
This is the same trade every application framework makes relative to a hand-rolled server: less flexibility at the edges, in exchange for a large amount of correctly-implemented, rarely-differentiated infrastructure code nobody has to write or review again.
Runners and separated scaling
BentoML's Runner abstraction — the API used before the 1.2 @bentoml.service style shown above — allows the model-execution step to scale independently of the HTTP-handling step, running model inference in separate worker processes coordinated by the API server. This matters specifically for CPU-bound models under Python's GIL, and mirrors the process-scaling argument in serving a model with FastAPI: the API layer and the compute layer have different scaling bottlenecks, and conflating them into one process limits both.
Framework choice as an organisational decision
At the scale of a single model with light traffic, BentoML, a hand-rolled FastAPI server, and Ray Serve differ mainly in developer ergonomics. The decision matters more at the scale of many models maintained by many teams: a shared packaging convention (what BentoML, and Triton's model repository, both provide) reduces the variance in how each model gets deployed, which is often the larger operational cost than any single framework's raw throughput.
Papers and systems
- BentoML's own documentation on the Bento build format and deployment targets is the primary source for its packaging model: docs.bentoml.com
- Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — early academic articulation of separating the "container" serving layer from framework-specific model code, the same separation BentoML's Runner design formalises.
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2014 — the "glue code" and "pipeline jungles" patterns this class of framework specifically targets.
What to learn next
- Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
- Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
- Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.