AI in Education

Automated essay scoring

Automated essay scoring predicts the score a human grader would give, and it correlates well with human scores on average while remaining gameable by long, repetitive, well-punctuated nonsense, a genuine and well-documented weakness.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Automated essay scoring predicts the score a human grader would have given an essay, from the text alone.

Think about judging a wedding gift by how heavy the box feels, without ever opening it. A heavy box often does hold something substantial. A brick wrapped in gift paper would fool you completely.

Automated essay scoring (AES) judges writing quality using measurable clues like length, vocabulary, and grammar. These clues usually track quality. They can also be faked by anyone who knows what they measure.

Why it exists

Grading thousands of essays by hand, consistently, is slow and expensive. Two human graders do not always agree either. The same essay can score differently, depending on who reads it and how tired they are that day.

AES was built to grade at scale, quickly, and consistently with itself every time. It learns from a large set of essays a human already scored. It finds the textual patterns — vocabulary range, sentence structure, grammar, length — that tend to go with a high or low human score.

How it works

Essay text   ->   Measured features        ->   Predicted score
                  (length, vocabulary,
                   grammar, structure)

Older systems used hand-picked features like these directly. Newer ones read the whole essay with a language model and predict a score from that richer understanding. Either way, the system is predicting what a human grader would likely say, not directly judging truth or insight.

Where you have already seen it

  • The GRE's Analytical Writing section, scored partly by ETS's e-rater system alongside a human rater.
  • TOEFL and other standardised English tests, several of which use automated scoring for at least part of the writing section.
  • Some school platforms that give instant draft feedback on an essay before a teacher ever reads it.

An honest warning

AES has a well-documented, serious weakness: it can be fooled. Researchers have shown that long, grammatically correct but meaningless essays, stuffed with sophisticated vocabulary, can score highly on some systems. The system is reading surface patterns, not meaning.

AES also struggles with anything a length-and-vocabulary model cannot see. Is an argument true? Is an idea genuinely original? Is a personal essay honest? A wrong or unfair score can affect a real student's exam result or college application. That is why high-stakes deployments almost always use AES as one of two scorers. A human rater checks or overrides it — AES is never the sole grader.

Remember this

  • AES predicts the score a human would give, based on textual patterns, not a direct judgement of meaning or truth.
  • It can be gamed by fluent-sounding, meaningless writing — a genuine, published weakness, not a rare edge case.
  • High-stakes exams pair it with a human rater rather than trusting it alone, because the cost of a wrong score is real.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Minimal runnable code

essay_score_demo.py
from sklearn.linear_model import LinearRegression

# Each row: [word_count, unique_word_ratio x100] for one past scored essay
# unique_word_ratio is how many distinct words appear, as a percent of total words
# Synthetic numbers, illustrative only -- not a real scored essay dataset
X = [
    [120, 55], [340, 60], [410, 45], [180, 70], [500, 40],
    [260, 65], [90,  50], [430, 58], [300, 62], [150, 68],
]
# human-assigned score, out of 10, for each essay above
y = [4, 7, 7, 5, 8, 7, 3, 8, 7, 5]

model = LinearRegression()
model.fit(X, y)

print("coefficient on word_count:      ", round(model.coef_[0], 4))
print("coefficient on unique_word_ratio:", round(model.coef_[1], 4))
print("R-squared on training data:      ", round(model.score(X, y), 3))

# Two essays of very different quality, matched on length
short_thoughtful = [[150, 70]]   # shorter, but a rich, varied vocabulary
long_repetitive =  [[500, 30]]   # much longer, but repeats the same simple words

print("\npredicted score, short but varied essay:", round(model.predict(short_thoughtful)[0], 2))
print("predicted score, long but repetitive essay:", round(model.predict(long_repetitive)[0], 2))
Output
coefficient on word_count:       0.0135
coefficient on unique_word_ratio: 0.0629
R-squared on training data:       0.964

predicted score, short but varied essay: 5.17
predicted score, long but repetitive essay: 7.38

What actually happened

The long, repetitive essay outscored the short, varied one — 7.38 against 5.17 — despite having a far worse vocabulary ratio. This is the length-gaming problem, reproduced directly.

  • The per-unit coefficient on unique_word_ratio (0.0629) really is larger than the one on word_count (0.0135). That is not the whole story.
  • word_count swings by 350 between the two essays (150 to 500). unique_word_ratio swings by only 40 (70 to 30). Multiply each coefficient by its actual swing, and word count's contribution to the score wins, even though its per-unit weight looks smaller.
  • This is a realistic replay of the actual, published criticism of early AES systems: because length correlates with score in typical human-graded training data, a model can end up rewarding length more than it rewards the quality signals it was meant to detect.

Common mistakes

Reading a single coefficient's size as "how much this feature matters." As shown above, a feature's real effect depends on both its coefficient and its typical range of values. Comparing coefficients directly, without accounting for scale, is a common and misleading mistake — see linear regression.

Training only on features a student could easily learn to fake. Word count and vocabulary ratio are exactly the kind of surface feature a student could pad without writing anything better. Real production systems add many more features specifically to make simple gaming harder, though none eliminate the problem entirely.

Treating a high R-squared as proof the model understands quality. 0.964 here reflects how well the model fits ten training rows, not whether it has learned to detect genuine insight or argument quality. A model can fit surface patterns very well while still being fooled by writing designed to exploit exactly those patterns.

Try it yourself

Add a third feature, grammar_errors, to a few of the training rows (make the long, repetitive essay have a high error count) and refit the model. Check whether the model's predicted gap between the two example essays narrows. This is exactly the kind of feature-engineering escalation real AES research has gone through, one gaming discovery at a time.

What to learn next

Researcher — Mathematics and papers.

Feature-based and neural approaches

Classical AES (e.g. ETS's e-rater, Pearson's Intelligent Essay Assessor) extracts a feature vector per essay — length, syntactic complexity, grammar and mechanics error counts, discourse structure markers, vocabulary sophistication measures — and fits a regression or ordinal classification model against human scores. Neural approaches (Taghipour and Ng, 2016, an early LSTM-based AES model; and subsequent transformer fine-tuning approaches) instead learn features directly from token sequences, typically outperforming feature-based systems on held-out quadratic weighted kappa while being harder to audit for which textual properties actually drove a given score.

Evaluation metric

AES model quality is standardly reported as quadratic weighted kappa (QWK) against human rater scores:

kappa = 1 - (sum_ij w_ij * O_ij) / (sum_ij w_ij * E_ij)
  • O_ij — observed count of essays scored i by the model and j by the human rater.
  • E_ij — expected count under an assumption of independence between model and human scores.
  • w_ij = (i - j)^2 / (N - 1)^2 — a quadratic penalty weight, penalising large disagreements more than small ones, where N is the number of score categories.

QWK is used specifically because it accounts for chance agreement and penalises the size of a disagreement, not only whether two raters agreed exactly — appropriate for the ordinal, multi-point score scales typical of essay grading.

The BABEL generator and adversarial gaming

Powers et al. (2001), in a study conducted with ETS's own e-rater system, built a generator (nicknamed BABEL) producing syntactically well-formed, semantically nonsensical essays specifically designed to exploit known e-rater feature weightings, and reported these nonsense essays received scores substantially higher than their content warranted. Perelman (2014) and related later work extended and publicised similar critiques against subsequent AES systems, arguing surface-feature-based scoring remains structurally vulnerable to essays optimised against its known feature set, regardless of genuine content quality. These are not isolated anecdotes — they are the AES field's own documented adversarial-robustness literature, and any claim that a specific deployed AES system is immune to gaming needs to be checked against published red-teaming results for that specific system, not assumed.

Fairness and bias

Multiple studies (e.g. Bridgeman et al., 2012, on e-rater and TOEFL) have examined systematic score differences by first-language background and writing style, since a model trained predominantly on one variety of English can systematically under- or over-score writing that deviates from that variety in structurally consistent, non-error ways. This is an active fairness concern distinct from the gaming problem above, and a documented reason several major testing programmes retain a human rater as a check or override on the automated score rather than relying on AES alone, particularly for population subgroups underrepresented in the AES training data.

Papers

  • Page, E. B. (1966). The Imminence of Grading Essays by Computer. Phi Delta Kappan. The earliest AES system, Project Essay Grade.
  • Powers, D. E. et al. (2001). Stumping E-Rater: Challenging the Validity of Automated Essay Scoring. ETS Research Report — the BABEL generator study.
  • Taghipour, K., Ng, H. T. (2016). A Neural Approach to Automated Essay Scoring. EMNLP.
  • Bridgeman, B., Trapani, C., Attali, Y. (2012). Comparison of Human and Machine Scoring of Essays: Differences by Gender, Ethnicity, and Country. Applied Measurement in Education.
  • Perelman, L. (2014). When "the state of the art" is counting words. Assessing Writing.

Current state

AES is deployed at scale for large-standardised-test writing sections, essentially always alongside a human rater rather than as sole scorer, with automated-human disagreement above a threshold routed to additional human review. Transformer-based scoring is now common in research and increasingly in production, improving raw agreement with human scores while not resolving the underlying gaming and fairness concerns documented above, which are properties of learning from surface textual correlation with human scores rather than properties of any one model architecture.

What to learn next

What to learn next

These follow on from what you just read.

  • AI in Education

    Building an LLM tutor

    A good LLM tutor is designed to guide a student toward their own answer instead of handing it over, the same habit a good human tutor already has, and it needs that design because a confidently wrong answer can teach a wrong fact as easily as a right one.

  • AI in Education

    Predicting which students will struggle

    An early-warning model flags students likely to fall behind, using attendance and grade trends, purely so a teacher can step in sooner, and a wrong flag can do real harm if it changes how that student gets treated.

  • AI in Education

    Why AI-text detectors do not work

    AI-text detectors guess whether writing came from a language model by how predictable the word choices are, and that same test flags plain, simple human writing too, which is why they have a documented history of falsely accusing real students, especially non-native English speakers.