AI in Education

Item response theory

Item response theory scores a student's ability and a question's difficulty on the same shared scale, so a correct answer on a hard question counts for more than one on an easy question.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Item response theory scores how hard a question is and how able a student is, on one shared scale.

Think about a high-jump competition. As the bar rises, fewer athletes clear it. Every athlete has some jumping ability. Every bar height has some difficulty. Whether one athlete clears one height depends on both of those numbers together, not on either alone.

Item response theory (IRT) treats a test question exactly like a bar height, and a student exactly like an athlete's ability.

Why it exists

Counting correct answers treats every question as equally hard. Two students who both scored 7 out of 10 can be very different. One answered seven easy questions. The other answered seven including the three hardest ones on the test.

Raw scores also make different test forms hard to compare fairly. If one class got an easier paper than another, comparing raw scores directly is misleading. IRT fixes both problems by estimating item difficulty and student ability from the same response data, and placing both on one shared number line.

How it works

Low ability -----------------------------------> High ability
Low difficulty ---------------------------------> High difficulty

A student clears an item when their ability number
sits comfortably above that item's difficulty number.

Every answer in the data — correct or wrong, on every question, from every student — feeds into estimating both scales at once. A question many strong students get wrong is flagged as genuinely hard. A student who answers hard questions correctly is credited with high ability, more than someone who only answers easy ones.

Where you have already seen it

  • The GRE and GMAT, which use IRT-scored, computer-adaptive tests that give different students different questions and still produce comparable scores.
  • National assessments like NAEP in the US, which report student ability on a stable scale across different test forms and different years.
  • Any test description that mentions "scaled score" rather than a raw percentage — that scaling is usually IRT underneath.

An honest warning

IRT assumes a single test genuinely measures one ability, like "maths ability". A test mixing unrelated skills, or biased toward one group's background knowledge, produces a scale that looks precise. It can still be measuring the wrong thing, or an unfair one.

A polished ability number can look more objective than it actually is. The model is only ever as fair as the questions and the data that trained it.

Remember this

  • IRT places student ability and question difficulty on the same scale, estimated together from response data.
  • This makes different test forms and different years genuinely comparable, unlike raw scores.
  • A precise-looking ability score is only as fair as the test questions it was built from.

What to learn next

  • Knowledge tracing — a related model that tracks knowledge changing over time, instead of one fixed ability.
  • Adaptive testing — choosing each next question using a live ability estimate like this one.
  • Logistic regression — the general technique IRT's probability curve is built from.

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

irt_ability_estimate.py
import numpy as np

def p_correct(ability, difficulty):
    # 1-parameter logistic (Rasch) model: probability of a correct answer
    return 1 / (1 + np.exp(-(ability - difficulty)))

# Five questions, calibrated difficulties on the same scale as ability
item_difficulty = np.array([-1.5, -0.5, 0.0, 1.0, 2.0])
item_names =      ["easy_1", "easy_2", "medium_1", "hard_1", "hard_2"]

for name, d in zip(item_names, item_difficulty):
    print(f"{name:10s} difficulty={d:5.1f}   P(correct | ability=0.0) = {p_correct(0.0, d):.2f}")

# A real student's answers to these five questions
responses = np.array([1, 1, 1, 0, 0])  # got the two easy ones and the medium one right

# Estimate this student's ability by finding the theta that makes these
# responses most likely (maximum likelihood), via a simple grid search
theta_grid = np.linspace(-4, 4, 401)
best_theta, best_loglik = None, -np.inf

for theta in theta_grid:
    p = p_correct(theta, item_difficulty)
    p = np.clip(p, 1e-6, 1 - 1e-6)  # avoid log(0)
    loglik = np.sum(responses * np.log(p) + (1 - responses) * np.log(1 - p))
    if loglik > best_loglik:
        best_theta, best_loglik = theta, loglik

print(f"\nestimated ability (theta): {best_theta:.2f}")
print("P(correct) at estimated ability, per item:")
print(np.round(p_correct(best_theta, item_difficulty), 2))
Output
easy_1     difficulty= -1.5   P(correct | ability=0.0) = 0.82
easy_2     difficulty= -0.5   P(correct | ability=0.0) = 0.62
medium_1   difficulty=  0.0   P(correct | ability=0.0) = 0.50
hard_1     difficulty=  1.0   P(correct | ability=0.0) = 0.27
hard_2     difficulty=  2.0   P(correct | ability=0.0) = 0.12

estimated ability (theta): 0.72
P(correct) at estimated ability, per item:
[0.9  0.77 0.67 0.43 0.22]

What actually happened

p_correct is logistic regression with exactly one input variable: the gap between ability and difficulty. This is the same sigmoid curve covered in logistic regression, reused for a different purpose.

  • The grid search tries 401 candidate ability values and keeps whichever one makes the observed responses most probable — this is maximum likelihood estimation, the standard way IRT ability scores are actually computed, done here by brute force instead of calculus.
  • An estimated ability of 0.72 makes sense: this student cleared three questions up to difficulty 0.0, and missed the two harder ones. A theta near 0.72 is exactly the ability level where that specific pattern of hits and misses is most likely.
  • The final printed probabilities show the model's own view of how "expected" each answer was, given the estimated ability. Compare it against responses — the two misses (hard_1, hard_2) both had predicted probabilities under 50%, consistent with the student getting them wrong.

Common mistakes

Treating item difficulty as a teacher's guess instead of an estimate. Real IRT difficulty values come from fitting the model against real response data from many students, not from someone's intuition about how hard a question looks.

Estimating ability from too few items. Five questions, as used here, gives a rough estimate. Real testing programmes use dozens of items, or an adaptive sequence covered in adaptive testing, to narrow the estimate meaningfully.

Comparing raw response counts across students who took different items. The entire benefit of IRT disappears if you go back to counting correct answers directly — always compare estimated ability values, not raw scores, once an IRT model is in use.

Try it yourself

Change responses to [0, 1, 1, 1, 1] — missing the easiest question but getting everything else right, an unusual pattern for this model's assumptions. Rerun and look at the estimated ability. A single surprising miss like this is exactly the kind of pattern real IRT-based cheating and misconduct detection systems look for.

What to learn next

Researcher — Mathematics and papers.

The model family

The most common IRT models, in increasing order of complexity:

1PL (Rasch):   P(correct) = 1 / (1 + exp(-(theta - b)))
2PL:           P(correct) = 1 / (1 + exp(-a * (theta - b)))
3PL:           P(correct) = c + (1 - c) / (1 + exp(-a * (theta - b)))
  • theta — the examinee's latent ability, on an arbitrary but fixed scale (conventionally mean 0, standard deviation 1 in the calibration sample).
  • b — item difficulty, on the same scale as theta. P(correct) = 0.5 exactly when theta = b.
  • a — item discrimination, controlling how sharply probability changes near theta = b. A high a means the item cleanly separates students immediately above and below its difficulty; a low a means the item is uninformative, answered inconsistently regardless of ability.
  • c — the pseudo-guessing parameter, a lower asymptote representing the chance of a correct answer at very low ability, relevant for multiple-choice items where random guessing has a non-zero success rate.

The developer example implements the 1PL / Rasch model, the simplest and most restrictive: it assumes every item has equal discrimination and zero pseudo-guessing, which rarely holds exactly but yields a model whose sufficient statistic for ability is the number-correct score alone, giving it desirable measurement properties (specific objectivity) the 2PL and 3PL do not share.

Estimation

Item parameters and abilities are jointly unidentified from responses alone — the model is invariant to an affine rescaling of theta and a corresponding inverse rescaling of a and b. Estimation is done via marginal maximum likelihood (integrating over an assumed ability distribution, typically standard normal, to estimate item parameters first) followed by ability estimation per examinee conditional on fitted item parameters — expected a posteriori (EAP) or maximum likelihood, as implemented directly in the developer example.

Information and standard error

The item information function quantifies how much a given item narrows uncertainty about theta at a specific ability level:

I(theta) = a^2 * P(theta) * (1 - P(theta))      [for the 2PL model]

Total test information is the sum of item informations, and the standard error of an ability estimate is approximately 1 / sqrt(total information). This is the formal criterion adaptive testing uses to select the next item: choose whichever unadministered item maximises information at the examinee's current ability estimate.

Papers

  • Rasch, G. (1960). Probabilistic Models for Some Intelligence and Attainment Tests. Danish Institute for Educational Research. Origin of the 1PL model.
  • Birnbaum, A. (1968). Chapters in Lord and Novick's Statistical Theories of Mental Test Scores. Introduces the 2PL and 3PL logistic forms.
  • Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems. Standard reference for applied IRT and adaptive testing.
  • van der Linden, W. J., Hambleton, R. K. (eds.) (1997). Handbook of Modern Item Response Theory.

Current state

IRT remains the standard psychometric framework for large-scale standardised testing, valued for its interpretability and long track record of fairness auditing (differential item functioning analysis, which checks whether an item behaves differently for groups matched on overall ability — a direct, quantitative fairness check IRT enables that raw scoring does not). Multidimensional IRT and diagnostic classification models extend the single-ability assumption to multiple latent skills, closer to the multi-skill tagging used in knowledge tracing, at the cost of materially harder estimation.

What to learn next

What to learn next

These follow on from what you just read.

  • AI in Education

    Adaptive testing

    An adaptive test picks each next question based on how the student answered the last one, converging on an accurate ability estimate in far fewer questions than a fixed test that asks everyone the same thing.

  • AI in Education

    Spaced repetition scheduling

    Spaced repetition schedules a review right before you are about to forget, which turns out to need far less total study time than cramming while remembering the same amount.

  • AI in Education

    Automated essay scoring

    Automated essay scoring predicts the score a human grader would give, and it correlates well with human scores on average while remaining gameable by long, repetitive, well-punctuated nonsense, a genuine and well-documented weakness.