Knowledge tracing
Knowledge tracing updates a belief about whether a student knows a skill after every single answer, going up on a correct answer and down on a wrong one, the same way a coach revises an opinion after watching a few more attempts.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Knowledge tracing estimates whether a student has actually learned a skill, updating that estimate after every single answer.
Think about a dance teacher watching a student attempt a new step. One clean attempt does not prove they have learned it — it could be luck. One clumsy attempt does not prove they have not — nerves happen. The teacher watches several attempts and slowly forms a confident opinion, revising it up or down each time.
Knowledge tracing does the same thing with numbers instead of a teacher's eye, one question at a time.
Why it exists
A single test score is a snapshot. It cannot tell a tutoring app whether a student is improving right now, mid-lesson, question by question.
Knowledge tracing was built to track that in real time. After every answer, it updates one number: the estimated probability that the student currently knows the skill being tested. A correct answer usually pushes that number up. A wrong answer usually pushes it down. Over several questions, a clear picture emerges — even though any single answer, by itself, proves very little.
How it works
Belief before Student Belief after
this question + answers -> this question
"maybe knows it" correctly "more likely knows it"Two things make this trickier than it first sounds. A student can guess correctly without truly knowing the skill. A student can also know a skill and still slip, making a careless mistake. A good knowledge-tracing model accounts for both possibilities, instead of trusting each answer completely.
Where you have already seen it
- Adaptive maths apps like Khan Academy's mastery tracking, which shows a skill as "mastered" only after a consistent run of correct answers, not one lucky guess.
- Language apps that quietly re-introduce a word you got wrong recently, because their internal belief about whether you know it dropped.
- Any tutoring platform that says "you seem shaky on this topic" partway through a session, before you have finished it.
An honest warning
Knowledge tracing estimates a hidden mental state it can never directly observe. It only ever sees answers, never the student's actual understanding.
A confident-looking number like "85% likely to know this skill" is a model's best guess. It is built from a small handful of noisy answers. Treat it as a signal for what to teach next, not a certified fact about a child's ability.
Remember this
- Knowledge tracing updates a belief about student knowledge after every single answer, not only at the end of a test.
- It accounts for guessing and slipping, since neither one answer's correctness proves knowledge or its absence.
- The output is an estimate, not a diagnosis — it guides what to teach next, nothing more certain than that.
What to learn next
- What learning data looks like — the response log this model is built from.
- Item response theory — a related model that also estimates ability, from a different angle.
- Adaptive testing — using an estimate like this one to choose the next question.
Developer — Code and libraries.
Setup
No installation needed — this is plain Python arithmetic.
Minimal runnable code
This implements Bayesian Knowledge Tracing (BKT), the classic form of this model, from its two update rules directly.
def bkt_update(p_know, correct, p_transit=0.10, p_guess=0.25, p_slip=0.10):
# Step 1: update belief given the observed answer (Bayes' rule)
if correct:
numerator = p_know * (1 - p_slip)
denominator = numerator + (1 - p_know) * p_guess
else:
numerator = p_know * p_slip
denominator = numerator + (1 - p_know) * (1 - p_guess)
p_know_given_evidence = numerator / denominator
# Step 2: account for the chance the student learned something from this attempt
p_know_next = p_know_given_evidence + (1 - p_know_given_evidence) * p_transit
return p_know_next
# A student's first five attempts at one skill: wrong, wrong, right, right, right
answers = [0, 0, 1, 1, 1]
p_know = 0.30 # prior belief the student already knows this skill before attempt 1
print(f"before any attempts: P(knows) = {p_know:.2f}")
for i, correct in enumerate(answers, start=1):
p_know = bkt_update(p_know, correct)
print(f"after attempt {i} ({'correct' if correct else 'wrong'}): P(knows) = {p_know:.2f}")before any attempts: P(knows) = 0.30 after attempt 1 (wrong): P(knows) = 0.15 after attempt 2 (wrong): P(knows) = 0.12 after attempt 3 (correct): P(knows) = 0.40 after attempt 4 (correct): P(knows) = 0.73 after attempt 5 (correct): P(knows) = 0.92
What actually happened
Two wrong answers dropped the belief from 0.30 to 0.12. Three correct answers in a row then pulled it up to 0.92. Neither move happened instantly — the model updates gradually, weighing new evidence against its current belief.
p_guess(0.25) means a student who knows nothing still has a real chance of a lucky correct answer, so one correct answer alone cannot push belief straight to 1.0.p_slip(0.10) means even a student who knows the skill occasionally makes a careless mistake, so one wrong answer alone cannot push belief straight to 0.0.p_transit(0.10) is applied every attempt, correct or not, and represents the chance a student learns something from attempting the question alone. This is why belief keeps drifting upward slightly even without a correct answer, over a long enough sequence.- Notice belief never actually reaches exactly 1.0 or 0.0 in five attempts. That is intentional — the model always leaves room for the possibility it is wrong.
Common mistakes
Setting p_guess or p_slip to zero. This makes the model treat every single answer as perfectly certain proof, which collapses the entire point of using a probabilistic model instead of checking only the last answer.
Forgetting the running belief is the next call's input. p_know must be reassigned each loop iteration, using the previous output as the new input. Recomputing from the original 0.30 every time would ignore all prior evidence.
Trusting one skill's estimate for a different skill. BKT tracks one skill at a time. A high knowledge estimate on "fractions" says nothing about "algebra" unless a separate model is run for that skill too.
Try it yourself
Change the answer sequence to [1, 0, 1, 0, 1] — alternating right and wrong. Watch how the belief responds compared to the two-wrong-then-three-right pattern above. A model that swings wildly on an inconsistent student, rather than settling near 0.5, would be a sign the guess and slip parameters need tuning.
What to learn next
- What learning data looks like — the raw response log this model consumes.
- Item response theory — a model that estimates ability and item difficulty together, instead of skill mastery over time.
- Adaptive testing — choosing the next question using a live estimate like this one.
Researcher — Mathematics and papers.
The hidden Markov model formulation
BKT (Corbett and Anderson, 1995) is a two-state hidden Markov model per skill, with a binary latent state:
Latent state: L_t in {known, unknown}
Observed: C_t in {correct, incorrect}
Transition: P(L_t = known | L_{t-1} = unknown) = p_T (learning)
P(L_t = unknown | L_{t-1} = known) = 0 (no forgetting, standard BKT assumption)
Emission: P(C_t = correct | L_t = known) = 1 - p_S (slip)
P(C_t = correct | L_t = unknown) = p_G (guess)p_T— transition probability, chance of moving from unknown to known after one opportunity.p_G— guess probability, chance of a correct answer while in the unknown state.p_S— slip probability, chance of an incorrect answer while in the known state.p_L0— prior probability of starting in the known state, before any evidence.
Inference at each step is exact forward filtering in a two-state HMM, which is exactly the two-step update (Bayesian posterior update, then transition) implemented in the developer example — a closed-form special case that avoids needing general HMM forward-algorithm machinery because the state space has only two values.
Parameter fitting
Classical BKT fits (p_L0, p_T, p_G, p_S) per skill via expectation-maximisation or grid search over held-out log-likelihood of observed correctness sequences. A well-known identifiability problem (Beck and Chang, 2007) exists: distinct parameter combinations can produce near-identical predicted correctness sequences, meaning point estimates of these parameters carry real uncertainty even when predictive accuracy is good, a caution relevant whenever fitted p_G/p_S values are interpreted as literal guess/slip rates for reporting purposes.
Deep Knowledge Tracing and successors
Piech et al. (2015) introduced Deep Knowledge Tracing (DKT), replacing the two-state HMM with a recurrent neural network (typically an LSTM) that maps a sequence of (item, correctness) pairs to a predicted probability of correctness on the next item, across all skills jointly rather than one HMM per skill. This captures cross-skill transfer effects standard BKT cannot represent, at the cost of interpretability — DKT's hidden state has no guaranteed correspondence to "probability of knowing skill X".
Subsequent work identified that naive DKT can violate basic educational plausibility, predicting a drop in mastery immediately after a correct answer in some cases (Yeung and Yeung, 2018), motivating regularised variants that penalise non-monotonic mastery predictions. Attention-based architectures (SAKT, Pandey and Karypis, 2019; AKT, Ghosh et al., 2020) have since become competitive alternatives, applying the self-attention mechanism from attention from scratch to the interaction sequence directly.
Evaluation
Standard practice evaluates next-attempt correctness prediction via AUC on held-out students (never held-out attempts from students seen in training, which leaks information and inflates reported performance) — see group leakage for the general form of this mistake. Reported AUC in the education-data-mining literature commonly falls in the 0.70-0.86 range depending on dataset and model, with deep methods typically outperforming classical BKT by a small but consistent margin on large datasets, and BKT remaining competitive or preferable on small ones due to far fewer free parameters.
Papers
- Corbett, A. T., Anderson, J. R. (1995). Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge. User Modeling and User-Adapted Interaction. Original BKT formulation.
- Beck, J. E., Chang, K. (2007). Identifiability: A Fundamental Problem of Student Modeling. User Modeling.
- Piech, C. et al. (2015). Deep Knowledge Tracing. NeurIPS.
- Yeung, C. K., Yeung, D. Y. (2018). Addressing Two Problems in Deep Knowledge Tracing via Prediction-Consistent Regularization. L@S.
- Pandey, S., Karypis, G. (2019). A Self-Attentive Model for Knowledge Tracing. EDM.
- Ghosh, A., Heffernan, N., Lan, A. S. (2020). Context-Aware Attentive Knowledge Tracing. KDD.
Current state
Attention-based and deep sequence models lead published benchmarks, but classical BKT remains widely deployed in production tutoring systems (including Carnegie Learning's descendants of the original Cognitive Tutor) because its parameters map to interpretable, auditable quantities that a curriculum designer can inspect, unlike a trained neural network's internal state.
What to learn next
- Conjugate priors and updating — the general Bayesian updating machinery this model is a special case of.
- Item response theory — a complementary, non-sequential model of ability and difficulty.
- LSTM — the architecture underlying Deep Knowledge Tracing.