Validation and Hyperparameter Search

AIC and BIC

AIC and BIC score a model as fit minus a fee for complexity, so you can compare candidates from a single fit each — no refitting, no folds, and a different answer from each of the two.

On this page 5
  1. Why these exist
  2. How they work
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

AIC and BIC score a model by how well it fits the data, minus a fee charged for every extra knob it needed.

Think of hiring staff for a small shop. Each new worker sells a little more, so takings go up with every hire. But every worker also draws a salary. The owner does not ask "did takings rise?" — takings always rise. She asks whether the rise covered the salary.

A model with more parameters always fits its training data better, in the same way. AIC and BIC pay attention to the salary bill. A parameter is one number the model learns from data — one knob it gets to turn.

Why these exist

You already have an honest way to compare models: K-fold cross-validation. Split, refit, score, repeat.

Sometimes you cannot afford that. Fitting a single seasonal forecasting model can take minutes, and a search over fifty candidates means fifty of those, times five folds. Sometimes you must not do it: with 80 monthly observations, chopping the series into folds leaves training slices too short to say anything.

AIC and BIC give you a comparison number from one fit of each model. No folds, no refits. That is the entire appeal.

How they work

   how well the model fits          fee for each
   the data it was given            knob it used
   ────────────────────────    +    ─────────────    =    score
   (bigger fit → lower score)       (more knobs →
                                     higher score)

   lower score wins

The two differ in one place: the size of the fee.

  • AIC charges a flat fee per knob, whatever the dataset size.
  • BIC charges a fee that grows as the dataset grows.

So on a large dataset, BIC is the stricter of the two and tends to pick the smaller model. AIC is more relaxed and often keeps a knob or two extra. Neither is wrong — they were built to answer different questions, which the Researcher section separates properly.

A real example you have seen

Every automatic forecasting tool does this. A package picking the order of an ARIMA model fits a few dozen candidates. It keeps the one with the best information criterion. Weather services, demand planners and sales forecasts all lean on that search. Nobody sits there cross-validating 40 candidate models by hand.

Remember this

  • Score = how well it fits, plus a fee for every parameter. Lower wins.
  • One fit per model — no folds. That is why they exist.
  • BIC charges more on big datasets, so it favours smaller models than AIC does.

What to learn next

  • K-fold cross-validation — the assumption-light alternative these criteria approximate.
  • Learning curves — the other single-graph answer to "is a bigger model worth it?"
  • ARIMA — the setting where automatic selection by information criterion is standard practice.

Developer — Code and libraries.

Setup

bash
pip install scikit-learn statsmodels

Verified with statsmodels 0.14.6, scikit-learn 1.7.2, numpy 1.26.4 on CPU. Seeded, so the numbers below should reproduce.

Fitting the same data seven ways

We generate data from a cubic, then fit polynomials from degree 1 to degree 7. The true answer is degree 3. Watch what each criterion picks.

criteria_vs_cv.py
import numpy as np
import statsmodels.api as sm
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures

rng = np.random.default_rng(0)
n = 120
x = rng.uniform(-3, 3, n)
y = 0.5 * x**3 - 2 * x + 1 + rng.normal(0, 3, n)      # the truth is a cubic

print("degree  params      AIC      BIC   5-fold MSE")
for deg in range(1, 8):
    X = np.vander(x, deg + 1, increasing=True)        # column 0 is the intercept
    fit = sm.OLS(y, X).fit()
    pipe = make_pipeline(PolynomialFeatures(deg), LinearRegression())
    mse = -cross_val_score(pipe, x.reshape(-1, 1), y, cv=5,
                           scoring="neg_mean_squared_error").mean()
    print(f"{deg:6d}  {deg+1:6d}  {fit.aic:7.1f}  {fit.bic:7.1f}  {mse:11.2f}")
Output
degree  params      AIC      BIC   5-fold MSE
     1       2    674.6    680.1        16.25
     2       3    676.5    684.9        17.38
     3       4    613.9    625.0         9.74
     4       5    615.1    629.1         9.68
     5       6    612.8    629.6         9.32
     6       7    614.1    633.6         9.39
     7       8    616.1    638.4         9.59

Three different winners, and none of them is a mistake. BIC picks degree 3 — the true model. AIC picks degree 5. Cross-validation, which measures prediction error directly, also picks degree 5.

That agreement between AIC and cross-validation is not a coincidence, and the Researcher section explains where it comes from.

Read the differences, not the ranking

The raw numbers mean nothing on their own — only gaps between models fitted to the same data are interpretable.

akaike_weights.py
import numpy as np

aic = np.array([674.6, 676.5, 613.9, 615.1, 612.8, 614.1, 616.1])   # from above
delta = aic - aic.min()
weights = np.exp(-delta / 2) / np.exp(-delta / 2).sum()
for deg, d, w in zip(range(1, 8), delta, weights):
    print(f"degree {deg}   dAIC {d:5.1f}   weight {w:.3f}")
Output
degree 1   dAIC  61.8   weight 0.000
degree 2   dAIC  63.7   weight 0.000
degree 3   dAIC   1.1   weight 0.221
degree 4   dAIC   2.3   weight 0.121
degree 5   dAIC   0.0   weight 0.383
degree 6   dAIC   1.3   weight 0.200
degree 7   dAIC   3.3   weight 0.074

Degrees 1 and 2 are dead — 60 points behind. Degrees 3 through 6 are indistinguishable. The "winner" carries 38% of the evidence, and four models share the rest.

Reporting "AIC selected degree 5" hides all of that. Reporting the gaps does not.

The walkthrough

statsmodels counts the coefficients, not the noise level. Its aic uses the number of regression coefficients, with the residual variance left out of the count. Other packages include it and land two points higher, uniformly. Gaps between models are unaffected, which is the only thing you should read.

Cross-validation and AIC agreed here, and that is the normal case. Both are aiming at prediction error on new data. BIC is aiming at something else — the model that generated the data — so it goes its own way.

np.vander builds the design matrix directly so that statsmodels and the sklearn pipeline see the same seven models. Comparing criteria computed on one parameterisation against error computed on another is a common way to get nonsense.

The weights formula is standard practice, from Burnham and Anderson's model-selection text. It converts differences into a rough share of evidence. Treat it as a readability aid, not a probability.

Common mistakes

Comparing AIC across different datasets. The criterion contains the likelihood of these rows. Drop 12 rows with missing values in one model and not another, and the two numbers are no longer comparable. The fix: fit every candidate on identical rows.

Comparing a model of y against a model of log(y). These have different response variables, so their likelihoods live on different scales. The comparison is meaningless, and the log version usually "wins" by a mile. The fix: transform back with the Jacobian, or compare on a shared held-out error metric instead.

Using AIC on a neural network. The maths assumes the parameter count reflects model flexibility and that the sample is large relative to it. A network with 40 million parameters and 50,000 images breaks both assumptions. The fix: hold out data. There is no shortcut here.

Treating a 1-point AIC gap as a decision. As the weights above show, models within two points are tied. The fix: report the gap, keep the simpler model, and use domain judgement to break ties.

Try it yourself

Change n = 120 to n = 2000 and rerun the first script. Predict first: does BIC's preferred degree change, and does the gap between BIC's choice and AIC's choice grow or shrink? Then set the noise to rng.normal(0, 8, n) and watch both criteria retreat to simpler models as the signal drowns.

What to learn next

  • K-fold cross-validation — the assumption-light alternative these criteria approximate.
  • Learning curves — the other single-graph answer to "is a bigger model worth it?"
  • ARIMA — the setting where automatic selection by information criterion is standard practice.

Researcher — Mathematics and papers.

AIC: an estimate of out-of-sample divergence

For a model fitted by maximum likelihood,

$$ \mathrm{AIC} = -2 \ln \hat{L} + 2k $$

Where:

  • $\hat{L}$ — the maximised likelihood of the observed data under the fitted model.
  • $k$ — the number of estimated parameters, including the residual variance when the convention includes it.

The derivation (Akaike, 1973; 1974) is a bias correction. The maximised log-likelihood is an optimistically biased estimator of the expected log-likelihood on fresh data drawn from the same distribution, because the parameters were tuned on the sample. Under regularity conditions and with the true distribution in or near the candidate set, that bias is asymptotically $k$, giving the $+2k$ penalty on the $-2\ln$ scale. Minimising AIC therefore approximately minimises the expected Kullback–Leibler divergence between the fitted model and the data-generating distribution.

AICc corrects the small-sample failure of that asymptotic argument (Sugiura, 1978; Hurvich and Tsai, 1989):

$$ \mathrm{AICc} = \mathrm{AIC} + \frac{2k(k+1)}{n - k - 1} $$

with $n$ the sample size. The standing advice is to use AICc whenever $n/k < 40$; it converges to AIC as $n$ grows, so using it always costs nothing.

BIC: an approximation to a marginal likelihood

$$ \mathrm{BIC} = -2 \ln \hat{L} + k \ln n $$

Schwarz (1978) derived this from a Laplace approximation to the marginal likelihood $p(D \mid M) = \int p(D \mid \theta, M)\, p(\theta \mid M)\, d\theta$: for large $n$, $-2 \ln p(D \mid M) = \mathrm{BIC} + O(1)$. The $O(1)$ term carries the prior, which is why BIC is a crude Bayes factor and not a good one. Differences in BIC approximate twice the log Bayes factor between models.

The practical consequence is the penalty: $\ln n$ exceeds $2$ once $n > 7$, and keeps growing.

Efficiency versus consistency — the reason they disagree

The disagreement in the developer section is structural, not noise.

  • BIC is consistent. If the data-generating model is among the candidates, $P(\text{BIC selects it}) \to 1$ as $n \to \infty$.
  • AIC is efficient (asymptotically loss-optimal). When the true model is not in the candidate set — the usual situation — AIC selects a model whose predictive risk approaches the best achievable in the set at the optimal rate.

Yang (2005), Can the strengths of AIC and BIC be shared?, answers no: no selection rule can be both consistent and minimax-rate optimal. Choosing between them is choosing which question you are asking. "Which model made this data?" → BIC. "Which model predicts best?" → AIC, or cross-validation.

Stone (1977) proved that AIC and leave-one-out cross-validation are asymptotically equivalent for parametric models under regularity conditions — the same criterion arrived at from two directions. Shao (1997) develops the analogous correspondence for BIC, which lines up with leave-$v$-out schemes whose held-out fraction grows with $n$.

This gives the practical decision rule. When refits are cheap and the sample is reasonable, cross-validate: it makes fewer assumptions and survives model misspecification. When refits are expensive, when $n$ is small, or when the model class is genuinely parametric and maximum-likelihood fitted, information criteria buy nearly the same answer for the cost of one fit each — $O(1)$ against $K$ refits.

Where $k$ stops being countable

The penalty assumes you can count parameters. Regularisation, shrinkage, early stopping and hierarchical structure all break that. Extensions replace $k$ with an effective count:

  • Generalised degrees of freedom — for ridge, $\mathrm{tr}(H)$ with $H$ the hat matrix; for a wide class of estimators, $\sum_i \partial \hat{y}_i / \partial y_i$ (Efron, 2004).
  • DIC (Spiegelhalter et al., 2002) uses a posterior-based effective parameter count; it misbehaves for hierarchical and mixture models.
  • WAIC (Watanabe, 2010) is singularity-tolerant and defined via the posterior predictive density; PSIS-LOO (Vehtari, Gelman and Gabry, 2017) estimates leave-one-out predictive density with importance sampling and diagnostics that report their own failure. For Bayesian workflows these have effectively replaced AIC and BIC.

For deep networks none of this transfers. The parameter count vastly exceeds the sample, the likelihood surface has many equivalent optima, and the asymptotics behind both criteria do not hold. Held-out evaluation remains the only defensible instrument there.

What to learn next

  • K-fold cross-validation — the assumption-light alternative these criteria approximate.
  • Learning curves — the other single-graph answer to "is a bigger model worth it?"
  • ARIMA — the setting where automatic selection by information criterion is standard practice.