Validation and Hyperparameter Search
Random search
Drawing each setting at random from a range beats a neat grid at the same budget, because random points explore every knob fully instead of recycling a few values.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Random search picks each setting at random from a range you give, tries it, and repeats.
At the same budget it usually beats the grid.
Picture a farmer testing soil quality in a square field. Walking in four neat rows, taking four samples per row, sounds thorough — sixteen samples. But look at what got covered: only four distinct east–west positions, each measured four times. If soil quality mostly varies east to west, twelve of those sixteen samples repeated information.
Now scatter sixteen samples at random. Every sample lands at a new east–west position and a new north–south position. Sixteen distinct values of each, whichever direction turns out to matter.
Why this had to be invented
The grid's weakness is exactly the neat rows. Model settings are like the field: usually one or two of them matter a lot, and the rest barely matter at all. The catch is that nobody knows in advance which ones matter.
A grid hedges badly against that ignorance. It spends its budget re-testing the same few values of every setting, over and over, in different combinations. Random search hedges perfectly. Every trial tries a fresh value of every setting. Whichever setting turns out to be the important one, it got explored thoroughly.
There is also a happy freedom: with a grid, adding one more setting multiplies the cost. With random search, you decide the number of trials first — thirty, sixty — and it stays fixed no matter how many settings you search over.
How it works
grid, 16 tries random, 16 tries
x x x x . x x
x .
x x x x x x
x . x
x x x x x x
. x
x x x x x x
distinct values tried distinct values tried
per setting: 4 per setting: 16One more rule of thumb is worth carrying. Around sixty random trials give you a 95% chance of landing at least once inside the best 5% of the search space. That holds regardless of how many settings there are.
A real example you have seen
Tasting street food in a new city. Visiting stalls in strict street-by-street order covers a few streets deeply. Wandering at random samples the whole city's variety with the same number of meals. Your odds of finding a great stall are better. Greatness is concentrated in pockets you cannot predict.
Remember this
- Random search tries a fresh value of every setting on every trial.
- It wins over grids because usually only a few settings matter — and you don't know which.
- Budget is fixed up front; roughly 60 trials reliably finds a top-5% configuration.
What to learn next
- Bayesian optimisation — when trials are too expensive to spend blindly.
- Successive halving and Hyperband — random search that abandons losers early.
- Tuning gradient-boosted trees — the search spaces that matter for the models that win on tables.
Developer — Code and libraries.
Setup
pip install scikit-learn scipyOutputs verified with scikit-learn 1.7.2 and scipy 1.14.1 on CPU. Runs in under a minute.
Same data, same budget, better answer
This is the identical dataset and pipeline from the grid search lesson, which spent 80 fits to reach 0.916. Random search gets the same 80 fits.
from scipy.stats import loguniform
from sklearn.datasets import make_classification
from sklearn.model_selection import RandomizedSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
X, y = make_classification(n_samples=500, n_features=10, n_informative=5,
flip_y=0.05, random_state=4)
pipe = Pipeline([("scale", StandardScaler()), ("svm", SVC())])
space = {"svm__C": loguniform(1e-2, 1e3),
"svm__gamma": loguniform(1e-4, 1e1)}
search = RandomizedSearchCV(pipe, space, n_iter=16, cv=5, random_state=0)
search.fit(X, y)
print("models trained:", 16 * 5)
best = {k: round(v, 4) for k, v in search.best_params_.items()}
print("best settings:", best)
print("best CV accuracy:", round(search.best_score_, 3))models trained: 80
best settings: {'svm__C': 658.1332, 'svm__gamma': 0.0083}
best CV accuracy: 0.9260.926 versus the grid's 0.916, at identical cost. Notice where it won: C=658, gamma=0.008 — a corner of the space the grid never offered, because 658 was not on anyone's list. Random search is not smarter; it is less committed to a list.
The walkthrough
loguniform is not decoration. Settings like C and gamma act multiplicatively — the jump from 0.01 to 0.1 matters as much as 10 to 100. loguniform(1e-2, 1e3) samples evenly in orders of magnitude. Plain uniform(0.01, 1000) would put 99% of draws above 10, never exploring the small end. Choosing the right distribution per setting is most of the skill here.
Distributions, not lists. The space gives ranges; each trial draws fresh values. You can mix in lists too — {"svm__kernel": ["rbf", "poly"]} — and random search samples from them uniformly.
random_state=0 makes the drawn trials reproducible. Without it, every rerun searches different points, and yesterday's result cannot be reproduced or debugged.
n_iter is the budget dial. 16 here for comparability with the grid; 30–60 is a sound default for a few settings. The cost is n_iter × cv fits, flat in the number of settings searched.
Common mistakes
Uniform ranges for multiplicative settings. The most common and most damaging one — learning rates, regularisation strengths, gamma, all need loguniform. Symptom: every good trial has the setting near one end of your range.
Too few trials, then blaming the method. Eight random trials can lose to a well-placed 3×3 grid; that is variance, not refutation. The 60-trial rule exists because randomness needs sample size.
Rerunning until a nice number appears. Ten reruns with different seeds, reporting the best, is score-shopping — the selection bias that inflates every reported maximum. One seeded run; confirm the final winner on an untouched test set.
Ranges that beg the question. loguniform(1, 10) for C cannot discover that the answer is 658. Start ranges generously wide — three to five orders of magnitude — then narrow after seeing where winners cluster.
Try it yourself
Run the search five times with random_state 0 through 4 and collect the five best scores. The spread you see is the honest uncertainty of a 16-trial search. Then set n_iter=60 once — does the best score beat all five small runs? Check where the winning C and gamma land each time.
What to learn next
- Bayesian optimisation — when trials are too expensive to spend blindly.
- Successive halving and Hyperband — random search that abandons losers early.
- Tuning gradient-boosted trees — the search spaces that matter for the models that win on tables.
Researcher — Mathematics and papers.
The coverage argument
Bergstra and Bengio (2012), Random search for hyper-parameter optimization (JMLR 13), formalise the field-sampling intuition. Model the objective as having low effective dimensionality: $f(x_1, \dots, x_p) \approx g(x_S)$ for a small, unknown subset $S$ of hyperparameters. A grid with $m$ points per axis projects onto only $m$ distinct values along $x_S$; $n$ i.i.d. random points project onto $n$ distinct values along every axis. Since $S$ is unknown a priori, random sampling is the minimax-sensible allocation. Their empirical result: on neural-network tuning tasks, random search matched or beat grid search using a small fraction of the compute.
The success probability bound
Let the "good region" be the top $\varepsilon$ fraction of the search space by score. The probability that $n$ independent draws all miss it is $(1 - \varepsilon)^n$, so
$$ P(\text{at least one hit}) = 1 - (1 - \varepsilon)^n $$
With $\varepsilon = 0.05$, $n = 60$: $1 - 0.95^{60} \approx 0.954$. The bound is dimension-free — the celebrated and slightly misleading property. The caveat: $\varepsilon$ is a volume fraction. In high dimensions, a fixed-radius neighbourhood of the optimum can occupy exponentially small volume, so "top 5% of the space" and "near the optimum" quietly diverge. The bound guarantees hitting good regions as defined by measure, nothing more.
Sampling distributions as priors
The search distribution is an explicit prior over configurations. Log-uniform draws for scale parameters correspond to Jeffreys-type uninformative priors on positive scale quantities — the principled reason behind the loguniform advice. Conditional spaces (e.g. degree exists only when kernel="poly") require tree-structured sampling; this is the configuration-space formalism of Bergstra, Bardenet, Bengio and Kégl (2011), which their TPE algorithm — covered under Bayesian optimisation — inherits directly.
Beyond i.i.d.: low-discrepancy sampling
I.i.d. sampling clumps; quasi-random sequences (Sobol' 1967; Halton 1960) fill space with low discrepancy, guaranteeing more even per-axis coverage at small $n$. Empirical comparisons in hyperparameter tuning show consistent small gains over i.i.d. sampling (Bousquet et al., 2017). Latin hypercube sampling gives exact per-axis stratification. Most modern tuners (Optuna's QMC sampler, Ax) offer these as drop-in replacements for the i.i.d. stage.
Position in the modern stack
Random search is now the baseline layer rather than the final tool: Hyperband is random search plus adaptive resource allocation; TPE and GP-based Bayesian optimisation begin with random startup trials and fall back to it for exploration. Its virtues — parallelism that needs no coordination, no modelling assumptions, dimension-free budgeting — keep it the recommended default whenever trials are cheap and plentiful, and the honest benchmark any smarter method must beat to justify its complexity.
What to learn next
- Bayesian optimisation — when trials are too expensive to spend blindly.
- Successive halving and Hyperband — random search that abandons losers early.
- Tuning gradient-boosted trees — the search spaces that matter for the models that win on tables.