AI for Science and Engineering
Choosing the next experiment to run
Active learning for experiments picks the next expensive measurement by asking a model where it is least sure, so a limited budget of real experiments teaches the most per run.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Active learning for experiments picks the next real-world measurement by asking a model where it is least sure.
Think about a cook adjusting salt one pinch at a time, tasting after each pinch. That beats dumping in a random amount of every spice and hoping for the best. Each taste test tells the cook exactly where to focus next.
Active learning does the same thing for real experiments — a wet-lab measurement, a wind-tunnel run, a materials test. It chooses each next expensive trial based on what the last one taught.
Why it exists
Every lesson in this section so far has assumed you already have training data. Real science does not start that way. A researcher usually has a limited budget for how many real experiments they can run. That budget might be money, time, or a finite number of samples of a rare material.
Choosing those experiments randomly wastes budget on regions already well understood. Regions that actually matter — where the outcome is genuinely uncertain — might never get tested at all. Active learning flips that. After each experiment, it asks the current model where it is most unsure, and spends the next expensive measurement exactly there.
How it works
A few initial experiments, spread across the possibilities
|
v train a model that reports its OWN uncertainty, not only a guess
Model says: "I am most unsure about THIS specific untested option"
|
v
Run the real, expensive experiment exactly there
|
v add the new result to the training data, retrain
Repeat, spending each new experiment where it teaches the mostThe model used for this needs to report not only a prediction, but how confident it is in that prediction. An ordinary regression model that only outputs a single number cannot drive this loop on its own.
Where this shows up
Materials scientists use active learning to decide which of thousands of possible material compositions to actually synthesise and test next. Often, only a handful of real syntheses are affordable.
Drug discovery teams use the same idea to choose which candidate molecules deserve the next round of real laboratory testing. This builds on the scoring and generation covered earlier in this section.
An honest warning
Active learning helps you spend a limited experimental budget well. It does not remove the need for the real experiment. It also does not guarantee the model's notion of "uncertain" lines up with what actually matters scientifically. A model can be confidently wrong in a region it never thought to doubt — the same way the surrogate models could, earlier in this section.
Remember this
- Active learning chooses the next real experiment based on where the current model is least confident, instead of choosing randomly.
- It needs a model that reports its own uncertainty, not only a single predicted number.
- It stretches a limited experimental budget further — it does not remove the need for the real experiment.
What to learn next
- Surrogate models for simulation — the fast, learned model this lesson's uncertainty estimate is built on.
- Gaussian process regression — the model family used in the code below, chosen specifically because it reports uncertainty.
- Prediction intervals for regression — the general idea of a model reporting a range, not only a point guess.
Developer — Code and libraries.
This example compares active learning against random sampling on the same fixed budget of "expensive" measurements, so the benefit is directly visible in the final numbers.
Setup
pip install scikit-learn numpyMinimal runnable code
expensive_experiment stands in for a real, slow, noisy measurement. Both strategies get the same budget: 15 total measurements. One picks each new point by maximum model uncertainty; the other picks randomly.
import numpy as np
from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import RBF
rng = np.random.default_rng(0)
def expensive_experiment(x):
return np.sin(3 * x) + 0.3 * x + rng.normal(0, 0.05, size=np.shape(x))
candidate_pool = np.linspace(0, 5, 200).reshape(-1, 1)
true_curve = np.sin(3 * candidate_pool.ravel()) + 0.3 * candidate_pool.ravel()
kernel = RBF(length_scale=0.6)
def run_strategy(pick_next_fn, budget=15, seed=1):
r = np.random.default_rng(seed)
x_seen = [0.0, 2.5, 5.0]
y_seen = list(expensive_experiment(np.array(x_seen)))
for _ in range(budget - 3):
gp = GaussianProcessRegressor(kernel=kernel, alpha=0.05, optimizer=None)
gp.fit(np.array(x_seen).reshape(-1, 1), y_seen)
next_x = pick_next_fn(gp, r)
x_seen.append(next_x)
y_seen.append(expensive_experiment(np.array(next_x)))
gp.fit(np.array(x_seen).reshape(-1, 1), y_seen)
predicted = gp.predict(candidate_pool)
mae = np.mean(np.abs(predicted - true_curve))
return mae, sorted(x_seen)
def pick_by_uncertainty(gp, r):
# active learning: ask the model where it is LEAST sure, and measure there
_, std = gp.predict(candidate_pool, return_std=True)
return candidate_pool.ravel()[np.argmax(std)]
def pick_randomly(gp, r):
return r.choice(candidate_pool.ravel())
active_mae, active_points = run_strategy(pick_by_uncertainty)
# average random sampling over several seeds since a single random draw is noisy
random_maes = [run_strategy(pick_randomly, seed=s)[0] for s in range(10)]
print(f"budget: 15 measurements each")
print(f"active learning (measure where uncertain) -> curve error (MAE): {active_mae:.4f}")
print(f"random sampling (avg of 10 runs) -> curve error (MAE): {np.mean(random_maes):.4f}")
print()
print("x-values active learning chose to measure:")
print(" " + ", ".join(f"{x:.2f}" for x in active_points))budget: 15 measurements each active learning (measure where uncertain) -> curve error (MAE): 0.0448 random sampling (avg of 10 runs) -> curve error (MAE): 0.0846 x-values active learning chose to measure: 0.00, 0.33, 0.63, 0.95, 1.26, 1.58, 1.88, 2.50, 3.09, 3.42, 3.74, 4.05, 4.37, 4.67, 5.00
Walkthrough
Both strategies get the exact same budget of 15 measurements against the exact same unknown curve. Active learning ends up with roughly half the error of random sampling, averaged over ten random runs to make sure that comparison is not a lucky fluke.
gp.predict(candidate_pool, return_std=True) is the key line — it returns not only a prediction but a standard deviation std at every candidate point, and pick_by_uncertainty always chooses the single point where that standard deviation is highest. Early on, that tends to be wherever the pool has been measured least; later, it shifts toward wherever the underlying curve is genuinely hardest to predict, such as its more sharply curving regions.
optimizer=None and a fixed kernel keep the Gaussian process's internal settings stable across every refit in the loop. Letting the optimiser re-tune those settings from only two or three points at a time makes early iterations noisy and unstable — a real, practical wrinkle in getting this loop to behave well with very little starting data.
The printed x-values active learning chose show this playing out directly: after the three starting points, it spends its early measurements filling in the widest unmeasured gaps first, then settles into a fairly even spread — exactly what "measure where uncertain" produces once most of the domain has been covered at least once.
Common mistakes
Using a plain regressor that has no notion of uncertainty. A random forest or plain neural network gives a prediction, not a "how sure am I" signal, without extra work. Gaussian processes provide this natively, which is why they are the classic choice for this loop.
Judging active learning from a single random comparison. A single random run can occasionally beat active learning by chance. Averaging several random runs, as this example does, is necessary before concluding anything.
Letting the acquisition strategy chase pure uncertainty forever, ignoring where the answer might actually be good. This example uses the simplest strategy, maximum uncertainty. Real systems often balance uncertainty against the chance a point is actually promising — the exploration/exploitation trade-off, covered from a different angle in what is reinforcement learning, and in more detail in the researcher section below.
Forgetting that "expensive" is the whole point. It is tempting to run more experiments without limit in a toy example. In a real lab, each additional point in x_seen might represent a week of work — the entire value of this technique is measured in how few real experiments it needs to reach a given accuracy.
Try it yourself
Change the budget from 15 to 8 and rerun both strategies. The gap between active learning and random sampling should widen — with fewer experiments to spend, choosing them wisely matters even more.
What to learn next
- Gaussian process regression — the full mathematics behind the model used here.
- Prediction intervals for regression — the general skill of getting a trustworthy uncertainty estimate out of a model.
- Surrogate models for simulation — where this section started, and the model this lesson decides how to build more efficiently.
Researcher — Mathematics and papers.
The formal setting
Let f: X -> R be an expensive-to-evaluate function, and let D = {(x_i, y_i)} be the set of points measured so far. An active learning, or Bayesian optimisation, loop repeats:
1. Fit a probabilistic model p(f | D), typically a Gaussian process
2. Choose x_next = argmax_{x ∈ X} a(x; p(f | D))
3. Measure y_next = f(x_next) + noise
4. D <- D ∪ {(x_next, y_next)}p(f | D)— the posterior belief over the unknown function, given data seen so fara(x; ...)— the acquisition function, scoring how valuable it would be to measure atxnext
The developer example's pick_by_uncertainty uses the simplest acquisition function, pure predictive standard deviation, a(x) = σ(x), ignoring the predicted mean entirely — a strategy usually called uncertainty sampling in the active learning literature (Settles, 2009) and closely related to the maximum variance criterion in Bayesian optimisation.
Better acquisition functions
Pure uncertainty sampling explores well but ignores whether a region is likely to contain a good answer, which matters when the goal is optimisation rather than pure curve-fitting. Expected Improvement (Jones, Schonlau and Welch, 1998) instead scores each candidate by how much it is expected to improve on the best value seen so far:
EI(x) = E[ max(0, f(x) - y_best) ]
= σ(x) · [ z · Φ(z) + φ(z) ] , z = (μ(x) - y_best) / σ(x)μ(x), σ(x)— the Gaussian process posterior mean and standard deviation atxy_best— the best (e.g. highest) value observed so farΦ, φ— the standard normal CDF and PDF respectively
Upper Confidence Bound acquisition (Srinivas et al., 2010) takes a simpler, tunable route, a(x) = μ(x) + β·σ(x), where β explicitly controls the exploration/exploitation trade-off the common-mistakes section above referred to informally.
Complexity and cost
For n points measured so far and m candidates evaluated per acquisition step:
| Component | Cost |
|---|---|
| Gaussian process refit | O(n^3) (exact) or O(n·m_inducing^2) with sparse approximations |
| Prediction with uncertainty over the candidate pool | O(n · m) |
Full active learning loop, k iterations | O(k · n^3) if refit exactly each round, dominated by the cubic GP refit as n grows |
The cubic cost in n is why active learning loops in practice are usually run for tens to low hundreds of iterations, not thousands — exactly the regime where the real experiment, not the model refit, is the actual bottleneck, so the mismatch rarely matters in practice.
Papers
- Jones, D., Schonlau, M. and Welch, W. (1998). Efficient Global Optimization of Expensive Black-Box Functions. Journal of Global Optimization 13. The Expected Improvement acquisition function.
- Srinivas, N., Krause, A., Kakade, S. and Seeger, M. (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML. The Upper Confidence Bound approach with theoretical regret guarantees.
- Settles, B. (2009). Active Learning Literature Survey. University of Wisconsin-Madison Computer Sciences Technical Report 1648. A broad survey covering uncertainty sampling and related query strategies.
- Shahriari, B. et al. (2016). Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE 104. A widely cited, application-focused survey.
Current state
Bayesian optimisation and active learning are established, practical tools in materials discovery, drug candidate selection and hyperparameter tuning, with mature open-source implementations. Current research pushes into batch active learning — choosing several experiments to run in parallel per round, since real labs rarely run one experiment at a time and waiting for each result serially wastes the parallelism a real lab has available — and into settings with multiple competing objectives (a material that must be both strong and cheap), where a single acquisition function is no longer enough and the field increasingly borrows multi-objective optimisation techniques.
What to learn next
- Gaussian process regression — the full mathematical treatment of the model this lesson's uncertainty estimates depend on.
- Surrogate models for simulation — the broader idea this lesson makes more efficient by choosing training points deliberately.
- Conformal prediction — a distribution-free alternative way to get trustworthy uncertainty estimates out of any model, not only a Gaussian process.