Baselines and Choosing a Model

Buying accuracy with size, and when to stop

Each doubling of model size buys a smaller accuracy gain at the same doubling of cost, and the stopping point is a business calculation you can put in a table.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Bigger models score better at a shrinking rate while costing more at a constant rate — so somewhere there is a stopping point, and it is findable.

Tuition works this way. The first month of coaching lifts a student's maths marks from 55 to 75. The second month adds 8 more. The sixth month adds 1, and costs the same as the first. At some point the family stops paying — not because improvement ended, but because it stopped being worth the fee.

Model size is tuition for machines. The first increase buys a lot. Later increases buy decimal places.

Why it exists

"More accuracy" sounds unarguable, so teams keep buying it past the point of sense. The discipline comes from seeing both curves at once:

  • Accuracy climbs fast, then flattens — diminishing returns, each step buying less than the last.
  • Cost climbs steadily and forever: training time, serving hardware, memory, latency, electricity.

A flattening benefit against a straight-line cost guarantees a crossover. After it, growth makes the system worse in every way the business feels — slower, pricier — for gains nobody can detect.

How it works

Make the two curves visible in one table, then find where your project stops:

size        accuracy      gain      cost
  1x          89%          —         1x
  4x          93%         +4.0      4x     ← big gain, cheap
 16x          94%         +1.0     16x     ← gain shrinking
 64x        95.3%         +0.9     64x     ← paying 4x for less than before

stop where:  value of the next gain  <  cost of the next doubling

The stopping rule needs one business number: what a point of accuracy is worth. One point of fraud recall — one more real fraud caught in every hundred — might be worth lakhs, so buy the big model. One point on a movie recommender might be worth nothing a user notices — stop early.

A real example you have seen

Phone cameras stopped racing megapixels years ago. 12 well-processed megapixels beat 108 badly processed ones, and users could not see the difference that remained. The industry found its crossover and moved the race elsewhere. Model size races end the same way.

Remember this

  • Accuracy grows with size at a shrinking rate; cost grows at a constant one.
  • Put both curves in one table — the crossover appears by itself.
  • The stopping rule needs a business number: the value of one more point.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26. Timings are from one laptop CPU; treat them as shape, not gospel.

The diminishing-returns table, generated live

Model size here is the number of boosted trees. Each row quadruples the size of the previous one.

size_sweep.py
import pickle
import time
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=12000, n_features=20, n_informative=12,
                           flip_y=0.05, random_state=1)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, random_state=0)

# throwaway fit so first-call warm-up does not pollute the timings
HistGradientBoostingClassifier(max_iter=5).fit(X_tr[:500], y_tr[:500])

print("trees   accuracy   train seconds   model size")
prev = None
for max_iter in [10, 50, 200, 800]:
    model = HistGradientBoostingClassifier(max_iter=max_iter, random_state=0,
                                           early_stopping=False)
    start = time.perf_counter()
    model.fit(X_tr, y_tr)
    seconds = time.perf_counter() - start
    acc = model.score(X_te, y_te)
    size_kb = len(pickle.dumps(model)) / 1024
    gain = "" if prev is None else f"   (+{acc - prev:.3f})"
    print(f"{max_iter:5d}   {acc:8.3f}   {seconds:13.1f}   {size_kb:7.0f} KB{gain}")
    prev = acc
Output
trees   accuracy   train seconds   model size
   10      0.891             0.1        78 KB
   50      0.932             0.2       217 KB   (+0.041)
  200      0.944             0.6       738 KB   (+0.011)
  800      0.953             2.2      2822 KB   (+0.009)

The walkthrough

Read the gain column downward: +0.041, +0.011, +0.009. Each quadrupling buys roughly a quarter of the previous gain, while the size and time columns quadruple on schedule. That is the whole phenomenon, in four rows, in three seconds of compute.

The last row costs 4x and buys +0.009. Whether that trade is good is not a modelling question. At fraud-detection stakes, 0.9 points may fund the hardware many times over; on a like-button predictor it is waste. The table's job is making the trade visible; the brief's metric line supplies the verdict.

This data has a ceiling. flip_y=0.05 re-assigns 5% of labels at random, and in a two-class problem about half of those land on the wrong class — so roughly 2.5% of rows are mislabelled and no model can pass about 0.975. The curve flattens below that, at 0.953, because capacity runs out before the noise floor does. Both walls are real: see can this even be learned?. Past either one, additional size can only memorise noise.

Same sweep, other families: hidden units for an MLP, model variant for pretrained networks (resnet18 → 50 → 101). The curve keeps its shape; only the axis labels change.

Common mistakes

Reporting the best model without the sweep. A lone "0.953" hides that 0.944 was available at a quarter of the cost. Decision-makers deserve the table, not the peak.

Sweeping size while accuracy is starving for something else. If features are weak or labels noisy, size buys almost nothing from the first doubling. Diminishing returns from the start is a diagnosis: fix data, not size — error analysis tells you which.

Ignoring serving multiplication. Training cost is paid once; serving cost is per prediction, forever. A model 4x slower at a million predictions a day is a permanent bill. The constraints lesson has the doors this table must fit through.

Trusting one split for 0.009-sized differences. Split luck moves accuracy by more than the last row's gain. Small differences need the repeated-fold treatment from comparing models fairly before you pay for them.

Try it yourself

Extend the sweep to 3200 trees. Predict the gain before running. Then compute a "rupees per accuracy point" column: assume serving cost scales with model size and a point of accuracy earns ₹10,000 a month. Find where your table says stop.

What to learn next

Researcher — Mathematics and papers.

Scaling laws: the phenomenon has equations

Across domains, loss follows a power law in model parameters $N$, dataset size $D$, and compute $C$:

$$ L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N} + L_\infty $$

Where:

  • $L$ — test loss; $L_\infty$ — the irreducible term (entropy of the data).
  • $N_c, \alpha_N$ — fitted constants; $\alpha_N \approx 0.076$ for transformer language models.
  • Analogous exponents exist for $D$ and $C$.

Kaplan et al. (2020), Scaling Laws for Neural Language Models, established the form; Hoffmann et al. (2022), Training Compute-Optimal Large Language Models (Chinchilla), corrected the allocation: at fixed compute, parameters and tokens should scale together (~20 tokens per parameter), and many models were oversized for their data. The power-law form is diminishing returns: each loss decrement costs multiplicatively more of everything.

Compute-optimal is not deployment-optimal

Chinchilla optimises training compute, but serving cost scales with $N$ per token forever. For inference-heavy deployments the optimum shifts toward smaller models trained longer ("over-training"): LLaMA-class models (Touvron et al., 2023) train far past Chinchilla-optimal token counts precisely to buy a cheaper serving point. The lesson's table is this argument at sklearn scale: the stopping point depends on the integral of serving cost, not the training bill.

Getting off the curve

The trade is not immutable; several techniques dominate points on the naive curve:

  • Distillation (Hinton et al., 2015): train large, compress into small — the student often beats a same-size model trained directly. See knowledge distillation.
  • Quantization and pruning shift the size axis at near-zero accuracy cost within limits (quantization in practice, pruning; lottery-ticket results, Frankle & Carbin, 2019).
  • Cascades: a small model handles easy cases, deferring hard ones to the large model — the expected serving cost is a weighted mix while accuracy tracks the large model (Viola & Jones, 2001 for the classical form; modern LLM routing continues it).

Measurement discipline at small deltas

Late-curve gains (+0.009) approach the standard error of evaluation. The binomial standard error on 3,000 test points at accuracy 0.95 is $\sqrt{0.95 \cdot 0.05 / 3000} \approx 0.004$ — the last row's gain is ~2 standard errors, borderline. Sweeps intended to justify spending should therefore report per-size confidence intervals (repeated folds or seed ensembles), or the crossover analysis inherits noise as signal. Mechanics in noisy validation curves.

What to learn next