AI in Education

Predicting which students will struggle

An early-warning model flags students likely to fall behind, using attendance and grade trends, purely so a teacher can step in sooner, and a wrong flag can do real harm if it changes how that student gets treated.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An early-warning model predicts which students are likely to fall behind, before it becomes obvious.

Think about a fever thermometer. It catches a rising temperature before you feel fully sick, giving you time to rest before things get worse. It does not diagnose what is wrong. It only says "something here needs a closer look."

An early-warning system in education does the same job with attendance, grades, and assignment records, instead of body temperature.

Why it exists

By the time a struggling student's report card shows failing grades, months of falling behind have often already happened. A teacher managing thirty or more students cannot track every warning sign for every student by hand, every week.

Researchers, including work from the University of Chicago Consortium on School Research, have studied this. Simple, early signals — attendance, and whether a student has failed a core course — are strong predictors of staying on track toward graduation. Often, they matter more than a single test score. An early-warning model turns those signals into a flag a school can act on, while there is still time to help.

How it works

Attendance, grades,        ->   Model estimates      ->   Flag for a
assignment history               risk level                teacher or counselor
(so far this term)                                          to look closer, not
                                                              a final verdict

The output is a starting point for a conversation, not a decision on its own. A flagged student typically gets a check-in from a teacher or counselor, who then decides what support, if any, actually makes sense.

Where you have already seen it

  • Many US school districts run some version of this, often called an "early-warning system" or "on-track indicator".
  • University advising offices that flag first-year students for extra support based on early attendance and grade patterns.
  • Any "at risk of falling behind" nudge inside a school's learning platform.

An honest warning

A wrong flag has real cost, in both directions. Missing a student who genuinely needs help delays support they should have gotten sooner. Wrongly flagging a student who is actually fine can change how they are treated. A teacher's expectations of a student are known to shift how that student is taught. A flag can become a self-fulfilling problem, not a fixed observation.

These models also risk quietly leaning on unfair proxies. Attendance and grades can correlate with a student's home situation, income, or background, in ways that have nothing to do with their ability to learn. A flagging system that is not checked for this can end up systematically over-flagging some groups of students and under-flagging others.

Because of this, an early-warning model should never be the sole basis for a decision about a student. Any real deployment needs review by educators, and a real check for unfair patterns across student groups, not only an accuracy number.

Remember this

  • Early-warning models flag students for a closer look, using attendance and grade signals, well before failure is obvious.
  • A wrong flag has real cost either way: a missed student, or a wrongly labelled one.
  • These systems need human review and fairness checks across student groups — they should never make the final call alone.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Minimal runnable code

at_risk_model.py
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
import numpy as np

# Each row: [attendance_rate (0-100), failed_courses_so_far, late_assignments_pct]
# Synthetic numbers, illustrative only -- not real student records
X = [
    [95, 0, 5],   [92, 0, 10],  [88, 1, 15],  [70, 2, 40],
    [65, 3, 55],  [98, 0, 0],   [55, 3, 60],  [80, 1, 20],
    [90, 0, 8],   [60, 2, 50],  [75, 1, 25],  [50, 4, 70],
]
# 1 = did not stay on track this term, 0 = stayed on track
# (a simplified stand-in for an indicator like Chicago's "Freshman OnTrack")
y = [0, 0, 0, 1, 1, 0, 1, 0, 0, 1, 0, 1]

model = LogisticRegression()
model.fit(X, y)

print("coefficients [attendance, failed_courses, late_assignments]:")
print(np.round(model.coef_[0], 3))

# A new student partway through term
new_student = [[78, 1, 30]]
prob = model.predict_proba(new_student)[0][1]
print(f"\npredicted risk probability for new student: {prob:.2f}")

# Check the model against its own training data (a real evaluation needs a held-out set)
preds = model.predict(X)
print("\nclassification report on training data (for illustration only):")
print(classification_report(y, preds, target_names=["on track", "at risk"]))
Output
coefficients [attendance, failed_courses, late_assignments]:
[-0.186  0.028  0.413]

predicted risk probability for new student: 0.11

classification report on training data (for illustration only):
              precision    recall  f1-score   support

    on track       1.00      1.00      1.00         7
     at risk       1.00      1.00      1.00         5

    accuracy                           1.00        12
   macro avg       1.00      1.00      1.00        12
weighted avg       1.00      1.00      1.00        12

What actually happened

The coefficients make sense directionally: attendance is negative (more attendance, less risk), failed courses and late assignments are both positive (more of either, more risk). That is a healthy sign the model learned something sensible from this small example.

  • The classification report shows perfect scores — 1.00 across the board. Look at that number with suspicion, not pride. It is measured on the exact twelve rows the model was trained on, which tells you almost nothing about a real, unseen student. See overfitting and underfitting for why this number is close to meaningless here.
  • predict_proba returns a probability, not a hard yes/no. Choosing the cutoff probability that counts as "flagged" is itself a real decision, covered generally in choosing a threshold from costs — and here it should weigh the cost of a wrong flag against the cost of a missed one, rather than default to 0.5 automatically.
  • The new student, with mixed signals (fair attendance, one failed course, moderate lateness), got a relatively low risk score of 0.11 from this tiny example model. A real model trained on far more data, and validated properly, would be needed before trusting a number like this for a real decision.

Common mistakes

Trusting training-set accuracy. As shown above, perfect scores on the data a model was trained on prove nothing about real-world performance. Always evaluate on a genuinely held-out set, covered in train-test split.

Skipping a fairness check across student groups. A model can have good overall accuracy while performing worse for one group of students than another. Checking accuracy, false-positive rate, and false-negative rate separately by relevant student subgroups is a real, necessary step before deployment, not an optional extra.

Using the flag as a punishment trigger instead of a support trigger. The entire ethical basis for a system like this is that a flag leads to more support, never to lower expectations, reduced opportunities, or exclusion. A system used the other way around causes real harm and defeats its own stated purpose.

Try it yourself

Compute risk probabilities for two students with identical attendance and failed-course counts, but one has 10% late assignments and the other has 60%. Compare the two probabilities. That comparison is the model actually being useful — showing which single changeable behaviour (submitting work on time) moves the needle most for a specific student.

What to learn next

Researcher — Mathematics and papers.

The formal task and its stakeholders

At-risk prediction is standard binary classification — see classification — but with a stakeholder structure that changes what "good performance" means. The cost of a false negative (missed student) and a false positive (wrongly flagged student) are both real and non-monetary, and are not symmetric with each other or stable across contexts, which makes a single accuracy number an inadequate summary on its own — see choosing a threshold from costs for the general framework.

The Freshman OnTrack research

Allensworth and Easton (2005, 2007), at the University of Chicago Consortium on School Research, developed and validated the Freshman OnTrack indicator — whether a first-year high-school student has earned enough credits and failed no more than one semester of a core course — as a predictor of eventual on-time graduation in the Chicago Public Schools system. Their published research found this simple, two-input indicator outperformed prior test scores and demographic variables as a predictor of graduation, and the indicator has since been adopted, in adapted forms, by school districts and states across the US as a practical early-warning tool. This is a genuine research result specific to the Chicago Public Schools population studied; the exact predictive strength of the indicator in a different district, country, or grading system should be validated locally rather than assumed to transfer unchanged.

Fairness auditing

Standard fairness metrics applied to at-risk models include:

False positive rate parity:   FPR should not differ substantially across groups
False negative rate parity:   FNR should not differ substantially across groups
Calibration within groups:    a predicted risk of p should mean roughly the
                               same real-world risk in every group

These three criteria are provably impossible to satisfy simultaneously except in special cases (Kleinberg, Mullainathan, Raghavan, 2016; Chouldechova, 2017), when base rates differ across groups — a mathematical result, not an engineering oversight, meaning any deployed system embeds an explicit or implicit choice about which fairness criterion it prioritises, and that choice should be a deliberate, documented, and reviewable one, not a default left unexamined.

The labelling-effect risk

Rosenthal and Jacobson's classic "Pygmalion in the Classroom" study (1968) — and a substantial body of subsequent educational-psychology research on teacher expectation effects — documents that a teacher's expectations of a student can measurably influence that student's outcomes, independent of the student's actual ability. This is the empirical basis for the beginner section's caution about wrongly flagging a student: a risk label is not administered into a vacuum, it is administered into a system (a classroom, a teacher's attention and expectations) that can itself shift outcomes in response to the label, in either direction.

Papers

  • Allensworth, E. M., Easton, J. Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public High Schools. University of Chicago Consortium on School Research.
  • Kleinberg, J., Mullainathan, S., Raghavan, M. (2016). Inherent Trade-Offs in the Fair Determination of Risk Scores. ITCS.
  • Chouldechova, A. (2017). Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data.
  • Rosenthal, R., Jacobson, L. (1968). Pygmalion in the Classroom. The Urban Review.
  • Bowers, A. J., Sprott, R., Taff, S. A. (2013). Do We Know Who Will Drop Out? A Review of the Predictors of Dropping Out of High School. High School Journal — a broad review of early-warning-indicator research.

Current state

Early-warning systems are in production use across many school districts, generally as a support-routing tool paired with mandated human review, rather than an automated decision system. Active work concentrates on fairness auditing methodology specific to the education context, and on the causal question of whether flagging plus intervention actually improves outcomes versus only identifying students who would have struggled regardless — a distinction that predictive accuracy alone cannot answer, and one that needs the kind of controlled evaluation covered in randomisation and assignment.

What to learn next