Serving Models in Production

BentoML

BentoML is a Python framework that turns a trained model into a packaged, servable API with far less boilerplate than writing every route and Dockerfile by hand.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

BentoML is a Python tool that packages a trained model into a ready-to-serve API, handling most of the boilerplate for you.

The analogy you have already lived

Serving a model with FastAPI is like building your own tiffin box from scratch. You cut the metal, weld the compartments, and it works exactly as you designed it. BentoML is closer to buying a well-designed tiffin box instead. The compartments, the latch and the carry handle already exist, sized for the standard case. You still choose what goes in it.

Why it exists

Writing a FastAPI server by hand means deciding, every time, how to load the model and structure the response. It also means building a Docker image and writing the deployment config. Most of those decisions are the same for most projects.

BentoML standardises them. Define a small "service" class around your model, and it handles the API layer, request validation, and building a deployable package. A team is then not reinventing the same server shape on every project.

How it works

your trained model (any framework: sklearn, PyTorch, XGBoost...)
      |
      v
  save it into BentoML's model store  (like git, but for model files)
      |
      v
  write a small Service class: what does "predict" do?
      |
      v
  bentoml serve   -> a running API, with auto-generated docs
  bentoml build   -> a portable, deployable package

The service class is the only genuinely new thing to write. Everything else — the web server, the package format — is BentoML's job.

A real example you have seen

Many small and mid-sized companies use frameworks exactly like this. They go from "a data scientist has a trained model" to "there is a working API" in an afternoon. No separate infrastructure engineer needs to write a server by hand for every model.

Remember this

  • BentoML packages a model into a servable API with far less boilerplate than a hand-built server.
  • You write one small service class; BentoML builds the API and the deployable package around it.
  • It works with most popular ML frameworks, not only one.

What to learn next

  • Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
  • Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
  • Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.

Developer — Code and libraries.

Setup

bash
pip install bentoml scikit-learn numpy

Saving a model and serving it

bento_service.py
import numpy as np
from sklearn.linear_model import LogisticRegression

import bentoml
from bentoml.models import BentoModel

rng = np.random.RandomState(0)
X = rng.uniform(0, 10, (200, 2))
y = (X[:, 0] + X[:, 1] > 10).astype(int)
clf = LogisticRegression().fit(X, y)

# Saves the trained model into BentoML's local model store, tagged and
# versioned automatically -- the same idea as the registry in
# [versioning and retiring features](/learn/production-feature-pipelines/feature-versioning-and-deprecation), for model artifacts instead of features.
saved = bentoml.sklearn.save_model("risk_model", clf)
print("saved tag:", saved.tag)


@bentoml.service(resources={"cpu": "1"})
class RiskScorer:
    model_ref = BentoModel("risk_model:latest")

    def __init__(self):
        self.model = bentoml.sklearn.load_model(self.model_ref)

    @bentoml.api
    def predict(self, income: float, years: float) -> dict:
        row = np.array([[income, years]])
        prob = float(self.model.predict_proba(row)[0, 1])
        return {"approved": prob >= 0.5, "probability": round(prob, 4)}


# During development, instantiate the service directly and call its method --
# no server, no port, the same pattern used for FastAPI's TestClient in
# [model serving](/learn/mlops/model-serving).
svc = RiskScorer()
print(svc.predict(income=6.0, years=5.0))
print(svc.predict(income=1.0, years=1.0))
Output
saved tag: risk_model:3brrjxngtckp74aj
{'approved': True, 'probability': 0.9253}
{'approved': False, 'probability': 0.0}

The exact tag after risk_model: is a random identifier BentoML generates per save — yours will differ every run. That is deliberate: it is how BentoML tells two versions of the same model apart.

Line-by-line walkthrough

bentoml.sklearn.save_model writes the model to a local store, distinct from a plain joblib.dump — it also records metadata (framework, save time, a unique tag) that the rest of BentoML relies on.

@bentoml.service turns a plain Python class into something BentoML can serve. @bentoml.api marks which method is a callable endpoint. Instantiating RiskScorer() directly, as done here, runs your model logic with zero networking involved — useful for fast local iteration before running a real server.

To actually run it as a live HTTP API, on the command line:

bash
bentoml serve bento_service:RiskScorer

No output block for that command — like uvicorn in model serving, it prints a live startup banner that varies by version and is not meaningful as static text.

Common mistakes

Loading the model inside predict instead of __init__. This repeats the "load every request" mistake from model serving, inside a different framework. Load once, in __init__, always.

Referencing BentoModel("risk_model:latest") in a class attribute, then never re-saving after retraining. :latest resolves at class-definition time in some setups — pin to the exact tag you validated in anything you deploy, and treat :latest as a development convenience only.

Skipping BentoML's own validation (Pydantic-style type hints on predict) and re-implementing manual checks. The type hints on predict's parameters are not decoration — BentoML uses them to validate incoming requests automatically, the same job model serving's Applicant model does for FastAPI.

Try it yourself

Add a second method to RiskScorer, predict_batch(self, rows: list[dict]) -> list[dict], that scores several applicants in one call. Compare calling it once with 50 rows against calling predict 50 times, the same batching argument made in model serving's exercise.

What to learn next

  • Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
  • Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
  • Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.

Researcher — Mathematics and papers.

What a "framework" actually standardises

BentoML's value is not the inference call itself — model.predict_proba(row) is identical whether wrapped in Flask, FastAPI, or BentoML. What it standardises is everything around that call: a model store with content-addressed versioning, automatic request-schema generation from Python type hints, a build system that produces a runnable, versioned artifact (a "Bento"), and adapters targeting several deployment backends (Docker, Kubernetes via Yatai, or several managed cloud runtimes) from that one artifact.

This is the same trade every application framework makes relative to a hand-rolled server: less flexibility at the edges, in exchange for a large amount of correctly-implemented, rarely-differentiated infrastructure code nobody has to write or review again.

Runners and separated scaling

BentoML's Runner abstraction — the API used before the 1.2 @bentoml.service style shown above — allows the model-execution step to scale independently of the HTTP-handling step, running model inference in separate worker processes coordinated by the API server. This matters specifically for CPU-bound models under Python's GIL, and mirrors the process-scaling argument in serving a model with FastAPI: the API layer and the compute layer have different scaling bottlenecks, and conflating them into one process limits both.

Framework choice as an organisational decision

At the scale of a single model with light traffic, BentoML, a hand-rolled FastAPI server, and Ray Serve differ mainly in developer ergonomics. The decision matters more at the scale of many models maintained by many teams: a shared packaging convention (what BentoML, and Triton's model repository, both provide) reduces the variance in how each model gets deployed, which is often the larger operational cost than any single framework's raw throughput.

Papers and systems

  • BentoML's own documentation on the Bento build format and deployment targets is the primary source for its packaging model: docs.bentoml.com
  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — early academic articulation of separating the "container" serving layer from framework-specific model code, the same separation BentoML's Runner design formalises.
  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2014 — the "glue code" and "pipeline jungles" patterns this class of framework specifically targets.

What to learn next

  • Ray Serve — a Python-native alternative built for scaling across many machines, not only one.
  • Serving many models on one machine — the problem a shared model store and packaging format make much easier to manage.
  • Model serving — the fundamentals BentoML automates once you understand what it is actually doing underneath.

What to learn next

These follow on from what you just read.

  • Serving Models in Production

    Ray Serve

    Ray Serve runs several copies of a model across many processes or machines behind one address, scaling a service by adding replicas instead of rewriting it.

  • Serving Models in Production

    Text Generation Inference (TGI)

    TGI is a server built specifically for generating text from large language models fast, using continuous batching to keep a GPU busy across requests of very different lengths.

  • Serving Models in Production

    gRPC vs REST for inference

    gRPC and REST are two different ways for a caller to talk to a model server, trading gRPC's speed and strict structure against REST's simplicity and universal support.