Write down the constraints before choosing a model
Latency, memory, cost, privacy, and explainability rules eliminate most model families before accuracy is ever discussed — so list them first.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
List the hard limits — speed, memory, money, rules — before comparing models, because the limits eliminate most candidates for you.
Nobody buys a sofa without measuring the door. The most beautiful sofa in the shop is worthless if it cannot enter the house. So you measure the door first, and suddenly the shop has only five sofas instead of a hundred.
Model choice works the same way. Measure the doorway first: how fast, how big, how expensive, how explainable the model must be. Most of the menu then disappears, before any accuracy contest begins.
Why it exists
Teams routinely spend a month building the most accurate model, then discover it cannot ship. It needs a GPU the budget does not cover. It answers in 900 milliseconds when the checkout page allows 100. It cannot explain a loan rejection, which the regulator requires.
Every one of those failures was knowable on day one, for the price of asking.
How it works
Walk through five doors, and write a number or a rule for each:
SPEED how fast must one answer come? "under 100 ms"
SIZE where does the model live? "on the phone, under 50 MB"
MONEY what may each prediction cost? "under 0.01 rupees"
REASONS must decisions be explained? "yes — banking regulation"
PRIVACY may the data leave the device/country? "no"Then check the menu. A giant neural network fails the phone-size door. Calling a paid API per prediction fails the money door at high volume. A black-box model — one whose individual decisions cannot be traced to reasons — fails the regulator door.
What remains is the honest shortlist, and it is often short: a small tree model, or a linear model with good features.
A real example you have seen
Your phone keyboard predicts the next word without internet. That model was chosen under brutal constraints: a few megabytes, instant answers, battery life, and your messages never leaving the device. A far more accurate server-side model exists — and was rejected by the doors, not by accuracy.
Remember this
- Constraints eliminate model families before accuracy gets a vote. Collect them first.
- The five doors: speed, size, money, reasons, privacy.
- A model that fails a hard constraint has accuracy zero, whatever the leaderboard says.
What to learn next
- Check the data exists before promising the model — the other pre-modelling audit.
- Buying accuracy with size, and when to stop — the trade this lesson's doors constrain.
- Why run models on device — the constraint set at its most extreme.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26. Timings below are from one laptop CPU — yours will differ, and that is fine: the ratio is the lesson.
Measuring two models against the doors
Accuracy contests get all the attention. This script runs the other contest — artifact size and single-prediction latency, the numbers a serving budget cares about.
import pickle
import time
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
X, y = make_classification(n_samples=5000, n_features=20, random_state=0)
one_row = X[:1]
for model in [LogisticRegression(max_iter=1000),
RandomForestClassifier(n_estimators=300, random_state=0)]:
model.fit(X, y)
size_kb = len(pickle.dumps(model)) / 1024
start = time.perf_counter()
for _ in range(200):
model.predict(one_row) # one row at a time, like a live API
ms = (time.perf_counter() - start) / 200 * 1000
name = type(model).__name__
print(f"{name:24s} artifact {size_kb:8.1f} KB one prediction {ms:5.2f} ms")LogisticRegression artifact 0.8 KB one prediction 0.04 ms RandomForestClassifier artifact 6695.1 KB one prediction 5.12 ms
The walkthrough
Three orders of magnitude, both columns. The forest is roughly 8,000 times larger and 100 times slower per prediction on this machine. If the accuracy gap between them is one point, the constraint table — not the accuracy table — should pick the winner.
One row at a time is deliberate. Batch prediction amortises overhead and flatters every model. A live API receives requests one by one; measure the shape you will serve. The time.perf_counter loop with 200 repetitions smooths timer noise.
Exact numbers are environment-dependent. Different CPU, different BLAS library, different results — sometimes by 5x. Re-run on hardware shaped like production before believing any absolute number. Ratios travel better than milliseconds.
Pickle size is a lower bound on the door. Serving adds the Python runtime, libraries, and memory overhead at inference. For a phone or an edge device, measure the whole packaged artifact.
Common mistakes
Benchmarking accuracy first and constraints last. The month spent tuning the losing family is unrecoverable. Filter by constraints, then tune inside the survivors.
Trusting batch throughput for a latency promise. "10,000 rows per second" and "one row in 10 ms" are unrelated claims. Percentile latency on single requests is the honest measure.
Forgetting the constraint on training, not serving. A model retrained nightly must finish training nightly. A 14-hour training run fails a 24-hour retrain cycle the first time data grows.
Treating explainability as a feature to add later. If regulation requires reasons per decision, the requirement shapes model family choice now. Bolting explanations onto a black box afterwards satisfies neither the regulator nor the customer — see explainability.
Try it yourself
Add HistGradientBoostingClassifier to the loop. Before running, write down your guess for both columns. Boosted trees usually land between the two extremes — check how close your intuition was, and which of the five doors it passes for your project.
What to learn next
- Check the data exists before promising the model — the other pre-modelling audit.
- Buying accuracy with size, and when to stop — the trade this lesson's doors constrain.
- Why run models on device — the constraint set at its most extreme.
Researcher — Mathematics and papers.
Constraints as the feasible set
Model selection under constraints is:
$$ \max_{f \in \mathcal{F}} \; \text{utility}(f) \quad \text{s.t.} \quad g_i(f) \leq b_i, \; i = 1, \dots, k $$
Where:
- $\mathcal{F}$ — the model family under consideration.
- $g_i$ — resource functions: p99 latency, peak memory, cost per query, energy per inference.
- $b_i$ — the budgets, fixed by product and regulation, not by the ML team.
The practical content: the feasible region ${f : g_i(f) \leq b_i}$ is often empty for entire families, which is why constraint collection precedes model search. Multi-objective formulations (accuracy-latency Pareto fronts) make the trade explicit rather than sequential.
Hardware-aware model design
The constraint-first view produced a research lineage of its own:
- MobileNets (Howard et al., 2017) parameterise the accuracy-compute trade with width and resolution multipliers — constraints as architecture inputs.
- EfficientNet (Tan and Le, 2019) treats the budget as the independent variable of compound scaling.
- Hardware-aware neural architecture search (Tan et al., 2019, MnasNet) puts measured latency directly into the search objective, replacing FLOP proxies — FLOPs correlate imperfectly with wall-clock latency across accelerators.
- Compression as constraint-satisfaction-after-the-fact: pruning, quantization, and knowledge distillation shrink a model into a budget the original violated.
Latency is a distribution, not a number
Serving latency is heavy-tailed; means mislead. Budgets should bind percentiles ($p95$, $p99$), because tail latency compounds across service fan-out — Dean and Barroso (2013), The Tail at Scale (CACM). For ML specifically, batching policy couples latency to throughput: dynamic batching trades p50 latency for GPU utilisation, so the constraint must state both the percentile and the traffic assumption under which it holds.
Regulatory constraints as hard requirements
Explainability requirements are jurisdiction- and domain-specific: adverse-action reasons in US credit (ECOA / Regulation B), documentation and transparency duties for high-risk systems in the EU AI Act (2024). Two design responses: inherently interpretable models for the constrained decision (Rudin, 2019, Nature Machine Intelligence), or post-hoc attribution methods (SHAP: Lundberg and Lee, 2017) with their known instabilities. Which response satisfies a given regulator is a legal determination — scope it with counsel, not with a library import.
What to learn next
- Check the data exists before promising the model — the other pre-modelling audit.
- Buying accuracy with size, and when to stop — the trade this lesson's doors constrain.
- Why run models on device — the constraint set at its most extreme.