Caching and Cost Control

Model cascades

A cascade tries a cheap, fast model first, and only calls in the expensive model for the cases the cheap one is not confident about.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model cascade tries a cheap, fast model first. It calls in an expensive model only for questions the cheap one cannot confidently answer.

The analogy you have already lived

A hospital does not send every patient straight to a senior specialist. A triage nurse sees everyone first. Most cases — a sprained ankle, a common cold — get handled right there, quickly and cheaply.

The nurse sends on only the cases that genuinely need it. Unclear symptoms. Something serious. Something they are not confident about. The specialist's expensive time gets spent where it actually matters.

A model cascade runs the same system on your requests.

Why it exists

Cutting LLM API costs mentioned using the smallest model that can do the job. The problem is that "the job" is rarely one difficulty level. Some questions are trivial. Some are genuinely hard. A single model choice forces a compromise. Pick the cheap model, and get some hard questions wrong. Pick the expensive model, and overpay for every easy one.

A cascade removes the compromise. It lets each question find its own right-sized model, automatically.

How it works

   question arrives
          |
          v
   ask the CHEAP, fast model
          |
          v
   is it confident in its answer?
        |                    |
       yes                   no
        |                    |
   use that answer      send it up to the
   (fast, cheap)         EXPENSIVE model
                          (slower, costs more,
                           but gets the hard
                           ones right)

The key idea is confidence — the cheap model does not just answer, it also reports how sure it is. A low-confidence answer is a signal to escalate, not a signal to trust.

A real example you have seen

Spam filters instantly block obvious spam, and instantly allow obvious real mail. They hold only a small number of uncertain messages for closer checking. Most email never needs "closer checking." The system spends its effort only where the easy rules genuinely are not enough.

The honest part

A cascade is only as good as the cheap model's ability to know what it does not know. Say the cheap model is confidently wrong — sure of itself on a question it actually gets wrong. The cascade never escalates that case, and the mistake goes uncaught. Measuring how often that happens matters as much as measuring the cost saved.

Remember this

  • A cascade tries the cheap model first, and escalates only what it is unsure about.
  • Confidence, not just the answer, is what decides whether to escalate.
  • It only works well if the cheap model's confidence is honest — genuinely lower on cases it actually gets wrong.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

A cascade, measured against both extremes

cascade.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier

rng = np.random.RandomState(0)
n = 3000
X = rng.uniform(-3, 3, (n, 5))
# A genuinely nonlinear boundary, so a linear model gets a real subset of
# cases wrong that a nonlinear model handles correctly.
score = X[:, 0] + 0.5 * X[:, 1] - 0.4 * X[:, 2] + 1.2 * np.sin(X[:, 0] * X[:, 1])
y = (score + rng.normal(0, 0.4, n) > 0).astype(int)

split = 2200
X_train, y_train = X[:split], y[:split]
X_test, y_test = X[split:], y[split:]

# The "cheap model": fast, small, linear -- decent but structurally limited.
cheap = LogisticRegression().fit(X_train, y_train)
# The "expensive model": stronger, standing in for a big, costly one.
expensive = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_train, y_train)

CHEAP_COST, EXPENSIVE_COST = 1, 20  # illustrative relative cost units, not real prices
CONFIDENCE_THRESHOLD = 0.75

cheap_proba = cheap.predict_proba(X_test)
cheap_pred = cheap_proba.argmax(axis=1)
cheap_confidence = cheap_proba.max(axis=1)

escalated = cheap_confidence < CONFIDENCE_THRESHOLD
cascade_pred = cheap_pred.copy()
cascade_pred[escalated] = expensive.predict(X_test[escalated])
cascade_cost = (~escalated).sum() * CHEAP_COST + escalated.sum() * (CHEAP_COST + EXPENSIVE_COST)

def accuracy(pred):
    return (pred == y_test).mean()

print(f"test set size: {len(X_test)}")
print(f"escalated to expensive model: {escalated.sum()} ({escalated.mean()*100:.1f}%)\n")
print(f"{'strategy':16} {'accuracy':>9} {'cost units':>11}")
print(f"{'always cheap':16} {accuracy(cheap_pred):9.3f} {len(X_test)*CHEAP_COST:11d}")
print(f"{'always expensive':16} {accuracy(expensive.predict(X_test)):9.3f} {len(X_test)*EXPENSIVE_COST:11d}")
print(f"{'cascade':16} {accuracy(cascade_pred):9.3f} {cascade_cost:11d}")
Output
test set size: 800
escalated to expensive model: 143 (17.9%)

strategy          accuracy  cost units
always cheap         0.873         800
always expensive     0.909       16000
cascade              0.904        3660

A real, reproducible run. The cascade recovers almost all of the expensive model's accuracy gain (0.904 versus 0.909). It spends under a quarter of what always using the expensive model would cost (3,660 versus 16,000 cost units). The exact numbers depend on this synthetic dataset and this threshold. The shape of the result does not: most of the accuracy, a fraction of the cost, is the general finding a well-tuned cascade aims for.

Line-by-line walkthrough

cheap_proba.max(axis=1). For a binary classifier, this is how confident the model was in whichever class it picked. Close to 1.0 means very sure; close to 0.5 means it is essentially guessing.

CONFIDENCE_THRESHOLD = 0.75. The line that decides how much escalates. Lower it, and fewer cases escalate — cheaper, but some of the cheap model's genuine mistakes go uncaught. Raise it, and more cases escalate — closer to the expensive model's accuracy, but closer to its cost too.

Cost accounting includes the cheap call even when escalating. (escalated.sum()) * (CHEAP_COST + EXPENSIVE_COST). A real cascade always pays for the first, cheap attempt, whether or not it ends up escalating. That cost is small but real, and should not be left out of the total.

Common mistakes

Trusting confidence without checking it is honest. A model can be confidently wrong. Before deploying a cascade, check the cheap model's accuracy specifically among the cases it was confident about. If that is meaningfully worse than its overall accuracy, its confidence is not a safe signal to route on.

Picking a threshold once and never revisiting it. As the questions your system sees change over time, the right threshold can shift. Recheck it against fresh, real examples periodically, the same way any other tuned constant needs revisiting.

Using a cascade where mistakes are expensive and rare escalation is not enough. For high-stakes decisions, a small number of confidently-wrong cheap answers slipping through can matter more than the money saved. Match the threshold's strictness to how costly a wrong answer actually is.

Measuring accuracy but not measuring cost per case. A cascade's whole value proposition is a joint claim about accuracy and cost. Report both together, as the table above does, or the comparison is incomplete.

Try it yourself

Sweep CONFIDENCE_THRESHOLD from 0.5 to 0.95 in steps of 0.05, and plot escalation rate against accuracy. Find the threshold where accuracy stops improving meaningfully for each extra percent escalated. That knee in the curve is usually the right operating point.

What to learn next

Researcher — Mathematics and papers.

Cascades as a constrained optimisation problem

Formally, a cascade with $K$ models ordered from cheapest ($M_1$) to most capable ($M_K$) chooses, for each input $x$, a stopping stage $k(x)$ based on a confidence estimate $\hat{c}_k(x)$ from model $M_k$, escalating while $\hat{c}_k(x) < \tau_k$. The design problem is choosing thresholds $\tau_1, \ldots, \tau_{K-1}$ to minimise expected cost

$$\mathbb{E}\left[\sum_{j=1}^{k(x)} \text{cost}(M_j)\right]$$

subject to a constraint on expected accuracy loss relative to always using $M_K$. This is a genuine Pareto frontier. For any target accuracy, there is a cost-minimal set of thresholds. Cascades that ignore this, using fixed, untuned thresholds, leave savings on the table.

Calibration is the load-bearing assumption

Everything above assumes $\hat{c}_k(x)$ is a calibrated confidence estimate — that among cases where the model reports confidence 0.8, roughly 80% are actually correct. Modern neural networks, particularly large ones trained with cross-entropy loss, are well documented to be systematically overconfident (Guo et al., 2017), which biases a naive cascade toward under-escalating. Production cascades typically apply post-hoc calibration (temperature scaling, or isotonic regression) to the confidence signal before using it for routing decisions, not the raw softmax output.

Cascades versus mixture-of-experts and routing models

A model cascade routes sequentially and pays for the cheap model on every request, even ones that escalate. A learned router is a separate, typically small model trained specifically to predict which downstream model will answer a given input correctly. It can route directly to the right model in one step, at the cost of needing labelled routing data to train it. FrugalGPT (Chen, Zaharia and Zou, 2023) formalises and compares both strategies for LLM APIs specifically. It reports that cascades can match or exceed the accuracy of the best single model, at a substantial fraction of its cost, across several public benchmarks — the general result the demo's synthetic experiment illustrates in miniature.

The mixture-of-experts architecture (see sparse models) is a related but distinct idea. It operates inside one model at the layer level, routing each token to a subset of parameters — rather than between separately deployed models at the request level.

Papers

  • Guo, Pleiss, Sun and Weinberger, On Calibration of Modern Neural Networks, ICML 2017 — arxiv.org/abs/1706.04599
  • Chen, Zaharia and Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, 2023 — arxiv.org/abs/2305.05176

What to learn next