AI in Education

Measuring learning, not engagement

Time spent, streaks and clicks are easy to measure and easy to game, and none of them are the same thing as learning, which can only really be checked by testing what a student can do before and after.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Time spent on an app is not the same thing as learning, and treating them as the same thing is a common, costly mistake.

Think about two people at a gym. One spends two hours there, scrolling their phone between sets. The other trains hard for thirty focused minutes. The person who spent more time is not automatically the one who got fitter.

Education apps face the exact same trap. Time spent, daily streaks, and clicks are all easy to measure. None of them are learning.

Why it exists

Real learning is slow and hard to measure. Checking whether a student actually understands more today than last month means testing them, comparing results, and waiting for the answer. That takes real effort and real time.

Engagement — minutes spent, streaks kept, lessons clicked through — is available instantly, updates in real time, and looks like it must be related to learning. It often is not closely related at all. Engagement numbers are easy to measure and easy to show to investors or schools. Platforms that optimise for them can end up optimising for something quite different from actual learning.

How it works

Easy to measure          Hard to measure
Time spent, streaks,     Actual learning gain
clicks, badges           (test score before vs after)
       |                          |
       v                          v
   Looks like              Is the thing that
   progress                actually matters

This mismatch has a name outside education too: Goodhart's Law — when a measure becomes a target, it stops being a good measure. Chase the streak, and a student can learn to keep the streak alive with minimal real effort, defeating the purpose the streak was meant to encourage.

Where you have already seen it

  • A language app streak you kept alive by doing the bare minimum lesson each day, without genuinely absorbing much.
  • A course platform proudly reporting "average time on platform" in its marketing, without reporting whether students actually learned more.
  • Any app that nudges you toward opening it daily, in ways that look designed for habit more than for education.

An honest warning

A platform reporting strong engagement numbers has not thereby shown it teaches well. The two can move together, move apart, or even move in opposite directions. Only measuring actual learning gain, not engagement alone, can tell you which is happening for a specific product.

Real learning-gain measurement needs a proper before-and-after comparison, ideally against a control group that did not use the product, the way any serious experiment does. Anyone claiming an education product "works" without that kind of evidence is making a claim engagement numbers alone cannot support.

Remember this

  • Time spent, streaks, and clicks are easy to measure, but they are not the same thing as learning.
  • A metric that becomes the target it is optimised for stops reliably measuring the thing it was meant to track.
  • Real evidence of learning needs a genuine before-and-after comparison, not engagement numbers alone.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

engagement_vs_learning.py
import numpy as np

# Ten students on a learning app over one month.
# Synthetic numbers, illustrative only -- not a real platform's data.
time_spent_min = np.array([420, 180, 610, 90, 340, 55, 500, 260, 130, 700])
streak_days    = np.array([28,  10,  30,  5,  20,  3,  29,  15,  8,   30])
# learning gain: post-test score minus pre-test score, out of 100
learning_gain  = np.array([12,  18,  6,   20, 15,  22, 5,   16,  19,  4])

print("correlation: time spent vs learning gain  ", round(np.corrcoef(time_spent_min, learning_gain)[0, 1], 2))
print("correlation: streak days vs learning gain  ", round(np.corrcoef(streak_days, learning_gain)[0, 1], 2))

top_engagement = np.argsort(-time_spent_min)[:3]
print("\ntop 3 students by time spent, and their actual learning gain:")
for i in top_engagement:
    print(f"  time_spent={time_spent_min[i]:4d} min   streak={streak_days[i]:2d} days   learning_gain={learning_gain[i]:2d}")

top_gain = np.argsort(-learning_gain)[:3]
print("\ntop 3 students by actual learning gain, and their engagement:")
for i in top_gain:
    print(f"  learning_gain={learning_gain[i]:2d}   time_spent={time_spent_min[i]:4d} min   streak={streak_days[i]:2d} days")
Output
correlation: time spent vs learning gain   -0.97
correlation: streak days vs learning gain   -0.95

top 3 students by time spent, and their actual learning gain:
  time_spent= 700 min   streak=30 days   learning_gain= 4
  time_spent= 610 min   streak=30 days   learning_gain= 6
  time_spent= 500 min   streak=29 days   learning_gain= 5

top 3 students by actual learning gain, and their engagement:
  learning_gain=22   time_spent=  55 min   streak= 3 days
  learning_gain=20   time_spent=  90 min   streak= 5 days
  learning_gain=19   time_spent= 130 min   streak= 8 days

What actually happened

This synthetic data was deliberately built to make the point vividly: the students who spent the most time and kept the longest streaks had the smallest learning gains, and vice versa. Read that as a demonstration of the risk, not a claim that real platforms always show correlations this strongly negative — real data is usually messier, with a weak or inconsistent relationship rather than a dramatic reversal like this one.

  • np.corrcoef(a, b)[0, 1] returns the correlation between two arrays, from -1 (perfectly opposite) to 1 (perfectly aligned). A value near 0 would already be a warning sign for a platform assuming engagement implies learning. The strongly negative values here are an exaggerated illustration of that same warning.
  • The "top by time spent" and "top by learning gain" lists barely overlap. If a school only looked at a dashboard of time-spent leaderboards, it would be praising exactly the students this data says learned the least.
  • This is the practical case for measuring learning gain directly, however inconvenient, rather than trusting engagement as a stand-in for it.

Common mistakes

Reporting correlation without checking causation or confounders. Even a real negative correlation here would not prove that spending more time causes less learning — struggling students might spend more time precisely because they are struggling, a confound covered generally in confounding.

Using engagement as the only success metric in a dashboard. If the only number a teacher or platform owner sees is time spent or streak length, that is the only number that will get optimised for, by design or by habit.

Assuming this pattern generalises without measuring it locally. The exact relationship between engagement and learning gain differs by product, subject, and student population. It needs to be measured for a specific platform, not assumed from an example like this one.

Try it yourself

Add a third feature, pretest_score, and check whether it explains part of the pattern — perhaps students who started with lower knowledge spent more time and gained less purely because the material was harder for them, not because of anything about "engagement" itself. That kind of check is the first step toward telling apart a real effect from a confound.

What to learn next

Researcher — Mathematics and papers.

Goodhart's Law and proxy metrics

Goodhart's Law, in the form popularised by Strathern (1997) from Goodhart's original 1975 monetary-policy observation — "when a measure becomes a target, it ceases to be a good measure" — describes a general phenomenon relevant well beyond education: optimising directly against a proxy metric degrades the proxy's correlation with the true underlying goal it was chosen to represent, because optimisation pressure exploits exactly the gap between the proxy and the goal. In education technology, engagement metrics (session length, streak length, daily active use) function as proxies for the true goal (learning), and platform design decisions that increase engagement (notifications, streak mechanics, variable-reward loops borrowed from mobile-game design) do not have a guaranteed, or even typical, positive relationship with the true goal.

What the empirical literature actually finds

The relationship between engagement and learning outcomes in the education-technology research literature is mixed and context-dependent, not uniformly positive or negative. Several studies of MOOC platforms (e.g., Kizilcec, Piech, Schneider, 2013, on disengagement patterns in MOOCs) find some engagement measures do correlate positively with completion and some learning measures, while others — particularly simple time-on-platform — show weak or inconsistent correlation with assessed learning gain once course difficulty and prior knowledge are controlled for. The honest summary is that engagement is not a reliable proxy for learning in general, and the specific relationship must be measured per platform and per metric rather than assumed in either direction.

Proper measurement of learning gain

The standard, defensible design is a pre-test/post-test comparison with a control condition:

Learning gain (treatment)  =  posttest_treatment - pretest_treatment
Learning gain (control)    =  posttest_control - pretest_control
Estimated effect           =  Learning gain (treatment) - Learning gain (control)

A treatment-only pre/post comparison, without a control group, cannot separate the platform's effect from simple maturation (students would have improved somewhat regardless, from time passing, other coursework, or repeated testing itself — see novelty and primacy effects and general regression-to-the-mean concerns), which is why a genuine randomised or quasi-experimental control condition, covered in randomisation and assignment, is the methodological floor for a credible learning-outcome claim, not an optional refinement.

Effect size interpretation

Education-intervention effect sizes are conventionally reported as standardised mean difference (Cohen's d or Hedges' g), since raw test-score-point differences are not comparable across different tests and scales:

d = (mean_treatment - mean_control) / pooled_standard_deviation

The What Works Clearinghouse and similar evidence-review bodies commonly treat d around 0.20 as a small but potentially meaningful effect in education research, though appropriate thresholds are debated and depend heavily on intervention cost and scale of deployment — a large-d study on 20 students carries far less evidential weight than a small-d finding replicated across a large, randomised, multi-site trial.

Papers

  • Goodhart, C. (1975), as popularised by Strathern, M. (1997). "Improving Ratings": Audit in the British University System. European Review — the widely cited restatement of Goodhart's Law.
  • Kizilcec, R. F., Piech, C., Schneider, E. (2013). Deconstructing Disengagement: Analyzing Learner Subpopulations in Massive Open Online Courses. LAK.
  • Reich, J. (2020). Failure to Disrupt: Why Technology Alone Can't Transform Education. Harvard University Press — a book-length treatment of the engagement-versus-learning gap in edtech specifically.
  • What Works Clearinghouse (Institute of Education Sciences). Standards Handbook — the standard US reference for effect-size interpretation and evidence-quality tiers in education research.

Current state

Rigorous, controlled learning-gain evidence remains scarcer than engagement-metric reporting across the commercial education-technology sector, a gap explicitly discussed in the education-research literature (e.g., Reich, 2020) as a structural feature of the industry's incentives rather than a temporary data-collection lag. Reading any specific product's efficacy claim critically — checking whether it reports a genuine controlled learning-gain comparison, or only engagement numbers — remains, as of the mid-2020s, a necessary skill for evaluating education technology rather than a niche academic concern.

What to learn next