AI in Education

What learning data looks like

Education data is usually one row per question attempt, not one row per student, because that is the level of detail every model in this section is actually built on.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Learning data is usually one row per question attempt, not one row per student.

Think of a cricket scorecard. A final score line says "245 for 6". A full scorecard records every single ball — who bowled it, what happened, when. The second kind of record lets you replay the whole innings.

Education data works the same way. A report card is the final score line. The data this section is built on is the ball-by-ball version: every question a student attempted, whether they got it right, and how long it took.

Why it exists

A grade tells you a student scored 60%. It does not say which topic they struggled with, whether they were guessing, or whether they got faster over the term.

An attempt-by-attempt record answers all of that. Each row is one interaction — one moment where a student answered one question. Stack enough rows and patterns appear. Which skills has a student learned? Which questions are too easy or too hard? How is a class progressing week by week?

Every model later in this section reads this same row-by-row log as input — predicting what a student knows, scheduling revision, spotting who needs help.

How it works

student   item        skill        correct   time taken
asha      frac_1      fractions    1         12 sec
asha      frac_2      fractions    0         30 sec
rahul     frac_1      fractions    0         45 sec

One row, one attempt. A single student produces many rows over a term. A whole school produces millions.

Where you have already seen it

  • A coaching app that tells you exactly which chapter to revise, built from your attempt history, not your last test score alone.
  • Duolingo's streak and mistake tracking, which watches every single answer, not only lesson completion.
  • A teacher's online gradebook that shows question-level results, not only a final percentage.

An honest warning

Raw attempt data is messy. A student can get an answer right by guessing. A student can get it wrong because of a typo, not a lack of understanding. Response time can mean deep thinking or a distracted phone check.

No model in this section reads a student's mind directly. Every one infers understanding from noisy clues — correctness and time. That inference is never perfectly certain.

Remember this

  • Learning data is usually logged one row per attempt, not one row per student.
  • This finer detail is what makes it possible to see which specific skill a student struggles with.
  • Correctness and response time are useful clues, not direct proof of what a student understands.

What to learn next

  • Knowledge tracing — turning this kind of log into an estimate of what a student knows right now.
  • Item response theory — using the same data to measure how hard each question actually is.
  • Pandas — the tool used throughout this section to work with data shaped like this.

Developer — Code and libraries.

Setup

bash
pip install pandas

Minimal runnable code

response_log.py
import pandas as pd

# A tiny synthetic response log: one row per question attempt, not one row per student.
log = pd.DataFrame([
    {"student": "asha",  "item": "frac_1", "skill": "fractions", "correct": 1, "response_time_sec": 12},
    {"student": "asha",  "item": "frac_2", "skill": "fractions", "correct": 0, "response_time_sec": 30},
    {"student": "asha",  "item": "algebra_1", "skill": "algebra", "correct": 1, "response_time_sec": 20},
    {"student": "rahul", "item": "frac_1", "skill": "fractions", "correct": 0, "response_time_sec": 45},
    {"student": "rahul", "item": "frac_2", "skill": "fractions", "correct": 0, "response_time_sec": 50},
    {"student": "rahul", "item": "algebra_1", "skill": "algebra", "correct": 1, "response_time_sec": 15},
    {"student": "meera", "item": "frac_1", "skill": "fractions", "correct": 1, "response_time_sec": 10},
    {"student": "meera", "item": "frac_2", "skill": "fractions", "correct": 1, "response_time_sec": 14},
])

print("raw log (each row is one attempt at one question):")
print(log)

print("\naccuracy per student:")
print(log.groupby("student")["correct"].mean().round(2))

print("\naccuracy per skill:")
print(log.groupby("skill")["correct"].mean().round(2))

print("\naverage time spent, correct vs incorrect answers:")
print(log.groupby("correct")["response_time_sec"].mean().round(1))
Output
raw log (each row is one attempt at one question):
  student       item      skill  correct  response_time_sec
0    asha     frac_1  fractions        1                 12
1    asha     frac_2  fractions        0                 30
2    asha  algebra_1    algebra        1                 20
3   rahul     frac_1  fractions        0                 45
4   rahul     frac_2  fractions        0                 50
5   rahul  algebra_1    algebra        1                 15
6   meera     frac_1  fractions        1                 10
7   meera     frac_2  fractions        1                 14

accuracy per student:
student
asha     0.67
meera    1.00
rahul    0.33
Name: correct, dtype: float64

accuracy per skill:
skill
algebra      1.0
fractions    0.5
Name: correct, dtype: float64

average time spent, correct vs incorrect answers:
correct
0    41.7
1    14.2
Name: response_time_sec, dtype: float64

What actually happened

groupby("student")["correct"].mean() is the entire trick of this whole section, in one line. Grouping the log by any column and averaging correct turns raw rows into a summary at whatever level you ask for — per student, per skill, per question.

  • The per-skill accuracy shows something a grade never would: this small class understands algebra fully but is at 50% on fractions specifically. A single overall grade would have hidden that split.
  • The response-time comparison hints at something real: wrong answers took much longer on average than right ones here. That pattern, guessing aside, often shows up in real data too — struggle tends to be slower than confident knowledge.
  • Nothing here required a machine learning model. Grouping and averaging is often the first, most informative step on any new education dataset, before anything fancier.

Common mistakes

Treating one row as one student. The whole point of this format is many rows per student. Code that assumes one row per student will silently produce nonsense the moment real logged data is loaded.

Averaging correctness without checking sample size. meera's 1.00 accuracy is based on only two attempts. A tiny sample can look perfect or terrible by chance. Always check how many attempts a percentage is built from before trusting it.

Ignoring timestamps entirely. This example drops the order of attempts for simplicity, but real data needs it. Whether a student got a question right on their first try or their fifth try changes what it means, covered in knowledge tracing.

Try it yourself

Add a fourth student with five attempts on fractions, where the first three are wrong and the last two are correct. Group by student and print the running accuracy after each attempt using .expanding().mean() on that student's rows, sorted by attempt order. Watch the number climb — that climbing line is what knowledge tracing tries to model directly.

What to learn next

  • Knowledge tracing — modelling how that accuracy trajectory changes over time.
  • Item response theory — a model of the item side of this data, not only the student side.
  • Pandas — full detail on the grouping and reshaping operations used here.

Researcher — Mathematics and papers.

The canonical schema

Most learning-analytics systems converge on a small set of fields per interaction, regardless of platform:

(student_id, item_id, skill_id(s), timestamp, correct, response_time, attempt_number, hint_count)
  • student_id, item_id — identifiers, typically foreign keys into separate student and item metadata tables.
  • skill_id(s) — one item is frequently tagged to multiple skills via a Q-matrix, a binary item-by-skill matrix originating in cognitive diagnosis models (Tatsuoka, 1983), specifying which skills each item exercises.
  • correct — usually binary, though partial-credit and multi-attempt scoring schemes exist and complicate every downstream model that assumes a clean binary label.
  • attempt_number — position within a student's sequence of attempts, essential for any temporal model since interactions are not independent and identically distributed observations, unlike the i.i.d. assumption underlying standard supervised learning (see what is machine learning).

This format is standardised, loosely, as the basis for the PSLC DataShop schema and the ASSISTments dataset format, two of the most widely used public sources for education-data research.

Missingness is informative, not random

Unlike many domains, missing data in education logs is rarely missing at random. A student skips a question because it looks hard, stops a session when frustrated, or never sees advanced items because an adaptive system, covered in adaptive testing, never selected them. Treating missingness as noise to impute away discards a real signal about engagement and difficulty, and can bias downstream models trained naively on completed-only data.

Cost and scale

A single active learner on a modern adaptive platform generates on the order of tens to low hundreds of interaction rows per session. At the scale of a national platform with millions of learners, this reaches billions of rows annually, which motivates the columnar, append-only storage patterns covered generally in data pipelines rather than a naive relational design.

Key sources

  • Tatsuoka, K. K. (1983). Rule Space: An Approach for Dealing with Misconceptions Based on Item Response Theory. Journal of Educational Measurement. Origin of the Q-matrix concept.
  • Feng, M., Heffernan, N., Koedinger, K. (2009). Addressing the Assessment Challenge with an Online System that Tutors as it Assesses. User Modeling and User-Adapted Interaction — describes the ASSISTments platform and dataset.
  • Koedinger, K. R. et al. (2010). A Data Repository for the EDM Community: The PSLC DataShop. Handbook of Educational Data Mining.

Current state

Learning-analytics data infrastructure has converged substantially on this row-per-interaction model across commercial and research platforms, though skill tagging (the Q-matrix) remains largely manual and a recognised bottleneck — automating accurate skill tagging from item text is itself an active NLP research problem, related to text classification.

What to learn next

What to learn next

These follow on from what you just read.

  • AI in Education

    Knowledge tracing

    Knowledge tracing updates a belief about whether a student knows a skill after every single answer, going up on a correct answer and down on a wrong one, the same way a coach revises an opinion after watching a few more attempts.

  • AI in Education

    Item response theory

    Item response theory scores a student's ability and a question's difficulty on the same shared scale, so a correct answer on a hard question counts for more than one on an easy question.

  • AI in Education

    Adaptive testing

    An adaptive test picks each next question based on how the student answered the last one, converging on an accurate ability estimate in far fewer questions than a fixed test that asks everyone the same thing.