Adaptive testing
An adaptive test picks each next question based on how the student answered the last one, converging on an accurate ability estimate in far fewer questions than a fixed test that asks everyone the same thing.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An adaptive test changes its next question based on how you answered the last one.
Think about a badminton coach feeding shots during practice. Hit an easy shot well, and the coach feeds a harder one next. Miss a few, and the coach eases up. The coach is not running through a fixed list — every shot depends on the last one.
A computer adaptive test (CAT) runs a written exam the same way, using item response theory to decide what "harder" and "easier" mean.
Why it exists
A fixed test gives everyone the same questions. That wastes time on a strong student, who breezes through easy questions that teach the test nothing new about their ability. It also demoralises a struggling student, who faces question after question far above their level.
Adaptive testing fixes both. Each new question is chosen to be the most informative one available, given everything learned about the student so far. That usually means a question near the student's current estimated ability. It should be hard enough to be a genuine test, not so hard it is a guaranteed miss.
How it works
Answer question -> Update ability estimate -> Pick next question
near that estimate
^ |
|______________________________________________________|The loop repeats until the estimate is precise enough, or a fixed number of questions has been asked. A student who is doing well quickly gets moved to harder questions. A student who is struggling gets moved to easier ones, without ever announcing that out loud.
Where you have already seen it
- The GRE and GMAT, both computer-adaptive, which is why two people can sit "the same" exam and never see the same set of questions.
- The NCLEX, the licensing exam for nurses in the US. It is also adaptive, and stops as soon as it is confident about a pass or fail decision.
- Adaptive practice apps that quietly get harder as you improve, without you ever choosing a difficulty level yourself.
An honest warning
Some adaptive tests decide real, consequential outcomes — whether someone becomes a licensed nurse, or which track a student is placed into. A wrong ability estimate in a high-stakes test like that has real cost for a real person.
Adaptive tests are also harder to inspect than a fixed paper. Two students see different questions, so comparing "which questions were on the test" no longer works as a fairness check. Any adaptive testing system used for a real, high-stakes decision needs its own dedicated fairness and validity review, not only accuracy numbers.
Remember this
- An adaptive test picks each next question based on the answers given so far, not a fixed order.
- It reaches an accurate ability estimate in far fewer questions than testing everyone on everything.
- High-stakes adaptive tests need real fairness auditing, since the usual "same test for everyone" check no longer applies.
What to learn next
- Item response theory — the model this lesson's item-selection logic is built on.
- Knowledge tracing — a related idea, tracking a changing skill instead of a fixed test-day ability.
- Predicting which students will struggle — another education model whose errors carry real consequences.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
import numpy as np
rng = np.random.RandomState(0)
def p_correct(ability, difficulty):
return 1 / (1 + np.exp(-(ability - difficulty)))
def estimate_ability(responses, difficulties, grid=np.linspace(-4, 4, 401)):
best_theta, best_loglik = 0.0, -np.inf
for theta in grid:
p = np.clip(p_correct(theta, difficulties), 1e-6, 1 - 1e-6)
loglik = np.sum(responses * np.log(p) + (1 - responses) * np.log(1 - p))
if loglik > best_loglik:
best_theta, best_loglik = theta, loglik
return best_theta
# A bank of 12 items spanning a wide difficulty range
item_bank = np.linspace(-3, 3, 12)
true_ability = 1.2 # this student's real, unobserved ability
# --- Adaptive test: always ask the item closest in difficulty to the current estimate ---
remaining = list(item_bank)
asked_d, asked_r = [], []
theta_est = 0.0 # start assuming an average student
for step in range(5):
next_item = min(remaining, key=lambda d: abs(d - theta_est))
remaining.remove(next_item)
correct = 1 if rng.random() < p_correct(true_ability, next_item) else 0
asked_d.append(next_item)
asked_r.append(correct)
theta_est = estimate_ability(np.array(asked_r), np.array(asked_d))
print(f"step {step + 1}: asked difficulty={next_item:5.2f} "
f"answer={'correct' if correct else 'wrong'} ability_estimate={theta_est:5.2f}")
print(f"\ntrue ability was {true_ability}, adaptive test estimate after 5 items: {theta_est:.2f}")
# --- Fixed test: always ask the first 5 items in the bank, regardless of the student ---
fixed_d = item_bank[:5]
rng2 = np.random.RandomState(0)
fixed_r = np.array([1 if rng2.random() < p_correct(true_ability, d) else 0 for d in fixed_d])
fixed_theta = estimate_ability(fixed_r, fixed_d)
print(f"fixed test (same 5 easy-to-hard items for everyone): estimate = {fixed_theta:.2f}")step 1: asked difficulty= 0.27 answer=correct ability_estimate= 4.00 step 2: asked difficulty= 3.00 answer=wrong ability_estimate= 1.64 step 3: asked difficulty= 1.91 answer=wrong ability_estimate= 0.84 step 4: asked difficulty= 0.82 answer=correct ability_estimate= 1.48 step 5: asked difficulty= 1.36 answer=correct ability_estimate= 1.94 true ability was 1.2, adaptive test estimate after 5 items: 1.94 fixed test (same 5 easy-to-hard items for everyone): estimate = 4.00
What actually happened
Two things worth noticing, and both are genuinely instructive, not code bugs.
- Step 1 jumps to 4.00, the edge of the search grid. After a single correct answer, the model has a perfect response record so far, and the maximum-likelihood estimate for a perfect record is unbounded — the true best answer is "infinitely able", so the search hits the edge of the grid it was given, and stops there. Real systems handle this with a Bayesian estimate that starts from a reasonable prior belief, instead of raw maximum likelihood, exactly to avoid this early instability.
- The fixed test also lands on 4.00 — for a different, worse reason. Its five questions are the easiest in the bank, so a student with true ability 1.2 answers all five correctly. The test never got hard enough to find this student's ceiling, so it cannot tell them apart from a much stronger student. The adaptive test, despite its own rough patch in step 1, ends up at 1.94 — closer to the true 1.2 than the fixed test's degenerate 4.00.
min(remaining, key=lambda d: abs(d - theta_est))is the entire item-selection rule: always ask whichever remaining question sits closest in difficulty to the current best guess.
Common mistakes
Using raw maximum likelihood with very few items. As shown in step 1, this is unstable with short, one-sided response histories. Real CAT systems use Bayesian ability estimates (a prior plus the evidence so far) for exactly this reason.
Reusing questions once a student has seen them. This toy example removes an item from remaining once asked, which real systems must also do — repeating a question a student has already seen breaks the model's assumptions and gives that student an unfair advantage.
Assuming more items always means more precision, without limit. Item information, covered in the researcher section of item response theory, does shrink uncertainty — but real test length is capped for reasons of student fatigue and exam-time cost, not only statistical ones.
Try it yourself
Change true_ability to 2.8, near the top of the item bank's range. Rerun both the adaptive and fixed tests and compare their final estimates. The fixed test's ceiling effect gets even worse for a stronger student, while the adaptive test keeps tracking reasonably well.
What to learn next
- Item response theory — the ability and difficulty model this whole selection process depends on.
- Prediction intervals for regression — expressing uncertainty around an estimate, not only a single number.
- Knowledge tracing — a related, time-varying alternative to a single test-day ability score.
Researcher — Mathematics and papers.
The adaptive testing loop, formally
1. Initialise theta_hat (commonly 0, or a prior mean).
2. Select the unadministered item i* that maximises information at theta_hat:
i* = argmax_i I_i(theta_hat)
3. Administer item i*, observe response u_i in {0, 1}.
4. Update theta_hat given all responses so far (MLE or Bayesian EAP/MAP).
5. Repeat from step 2 until a stopping rule is met.I_i(theta)— the item information function for itemiat abilitytheta, defined in the researcher section of item response theory.- Stopping rules in practice: fixed item count, standard error of
theta_hatbelow a threshold, or (for pass/fail tests like the NCLEX) sufficient confidence thattheta_hatis above or below the passing cutoff, regardless of overall precision — a variable-length stopping rule specifically documented for the NCLEX's CAT implementation.
Estimation stability
Maximum likelihood ability estimation is undefined or divergent for a response vector that is all-correct or all-incorrect, since the log-likelihood is then monotonic in theta with no interior maximum — exactly the failure the developer example's step 1 exhibits. Standard fixes:
- Bayesian modal (MAP) or expected a posteriori (EAP) estimation, incorporating a prior distribution over
theta(commonly standard normal), which keeps the posterior proper and finite even for extreme response patterns. - Step-size capping on early items, bounding how far
theta_hatcan move after a single response regardless of what MLE alone would suggest.
Item exposure control
Naively always selecting the maximum-information item causes a small subset of items to be shown far more often than others, which both wastes the rest of the item bank and creates test-security risk if those frequently-shown items leak. The Sympson-Hetter method (Sympson and Hetter, 1985) and randomesque item selection (choosing randomly among the top-k most informative items rather than always the single best) are standard mitigations, trading a small amount of measurement efficiency for materially better item-bank utilisation and security.
Evaluation
CAT systems are validated against their non-adaptive counterparts primarily on two axes: measurement precision (standard error of theta_hat at a given test length, typically substantially lower for CAT than a fixed test of the same length) and classification accuracy for pass/fail systems (agreement rate with the decision a full, non-adaptive administration of the entire item bank would have produced). Published CAT validation studies for licensure exams commonly report test-length reductions in the 40-60% range for equivalent measurement precision relative to fixed-form testing, though the exact figure is specific to the item bank and population studied and should not be treated as a universal constant.
Papers
- Weiss, D. J. (1982). Improving Measurement Quality and Efficiency with Adaptive Testing. Applied Psychological Measurement — foundational overview of CAT.
- Sympson, J. B., Hetter, R. D. (1985). Controlling Item-Exposure Rates in Computerized Adaptive Testing. Proceedings of the Military Testing Association.
- Wainer, H. (ed.) (2000). Computerized Adaptive Testing: A Primer. Standard textbook reference.
- van der Linden, W. J., Glas, C. A. W. (eds.) (2010). Elements of Adaptive Testing.
Current state
CAT is the operational standard for most major large-scale standardised testing programmes as of the mid-2020s. Its core selection and stopping-rule machinery is well established; active work concentrates on multidimensional CAT (adapting across several latent skills simultaneously, connecting to multi-skill knowledge tracing), and on fairness auditing methods specific to adaptive designs, where the traditional differential-item-functioning tools built for fixed tests need adjustment for the fact that different examinees see systematically different item sets.
What to learn next
- Item response theory — full detail on the difficulty and information model used for item selection.
- Choosing a threshold from costs — relevant to setting a pass/fail cutoff with an explicit cost trade-off.
- Knowledge tracing — the time-varying sibling of this fixed-ability-per-sitting model.