Automated essay scoring
Automated essay scoring predicts the score a human grader would give, and it correlates well with human scores on average while remaining gameable by long, repetitive, well-punctuated nonsense, a genuine and well-documented weakness.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Automated essay scoring predicts the score a human grader would have given an essay, from the text alone.
Think about judging a wedding gift by how heavy the box feels, without ever opening it. A heavy box often does hold something substantial. A brick wrapped in gift paper would fool you completely.
Automated essay scoring (AES) judges writing quality using measurable clues like length, vocabulary, and grammar. These clues usually track quality. They can also be faked by anyone who knows what they measure.
Why it exists
Grading thousands of essays by hand, consistently, is slow and expensive. Two human graders do not always agree either. The same essay can score differently, depending on who reads it and how tired they are that day.
AES was built to grade at scale, quickly, and consistently with itself every time. It learns from a large set of essays a human already scored. It finds the textual patterns — vocabulary range, sentence structure, grammar, length — that tend to go with a high or low human score.
How it works
Essay text -> Measured features -> Predicted score
(length, vocabulary,
grammar, structure)Older systems used hand-picked features like these directly. Newer ones read the whole essay with a language model and predict a score from that richer understanding. Either way, the system is predicting what a human grader would likely say, not directly judging truth or insight.
Where you have already seen it
- The GRE's Analytical Writing section, scored partly by ETS's e-rater system alongside a human rater.
- TOEFL and other standardised English tests, several of which use automated scoring for at least part of the writing section.
- Some school platforms that give instant draft feedback on an essay before a teacher ever reads it.
An honest warning
AES has a well-documented, serious weakness: it can be fooled. Researchers have shown that long, grammatically correct but meaningless essays, stuffed with sophisticated vocabulary, can score highly on some systems. The system is reading surface patterns, not meaning.
AES also struggles with anything a length-and-vocabulary model cannot see. Is an argument true? Is an idea genuinely original? Is a personal essay honest? A wrong or unfair score can affect a real student's exam result or college application. That is why high-stakes deployments almost always use AES as one of two scorers. A human rater checks or overrides it — AES is never the sole grader.
Remember this
- AES predicts the score a human would give, based on textual patterns, not a direct judgement of meaning or truth.
- It can be gamed by fluent-sounding, meaningless writing — a genuine, published weakness, not a rare edge case.
- High-stakes exams pair it with a human rater rather than trusting it alone, because the cost of a wrong score is real.
What to learn next
- Text classification — the general technique most AES systems are built from.
- Building an LLM tutor — a related use of language models in education, with its own honest limits.
- Why AI-text detectors do not work — a related, even less reliable use of AI on student writing.
Developer — Code and libraries.
Setup
pip install scikit-learnMinimal runnable code
from sklearn.linear_model import LinearRegression
# Each row: [word_count, unique_word_ratio x100] for one past scored essay
# unique_word_ratio is how many distinct words appear, as a percent of total words
# Synthetic numbers, illustrative only -- not a real scored essay dataset
X = [
[120, 55], [340, 60], [410, 45], [180, 70], [500, 40],
[260, 65], [90, 50], [430, 58], [300, 62], [150, 68],
]
# human-assigned score, out of 10, for each essay above
y = [4, 7, 7, 5, 8, 7, 3, 8, 7, 5]
model = LinearRegression()
model.fit(X, y)
print("coefficient on word_count: ", round(model.coef_[0], 4))
print("coefficient on unique_word_ratio:", round(model.coef_[1], 4))
print("R-squared on training data: ", round(model.score(X, y), 3))
# Two essays of very different quality, matched on length
short_thoughtful = [[150, 70]] # shorter, but a rich, varied vocabulary
long_repetitive = [[500, 30]] # much longer, but repeats the same simple words
print("\npredicted score, short but varied essay:", round(model.predict(short_thoughtful)[0], 2))
print("predicted score, long but repetitive essay:", round(model.predict(long_repetitive)[0], 2))coefficient on word_count: 0.0135 coefficient on unique_word_ratio: 0.0629 R-squared on training data: 0.964 predicted score, short but varied essay: 5.17 predicted score, long but repetitive essay: 7.38
What actually happened
The long, repetitive essay outscored the short, varied one — 7.38 against 5.17 — despite having a far worse vocabulary ratio. This is the length-gaming problem, reproduced directly.
- The per-unit coefficient on
unique_word_ratio(0.0629) really is larger than the one onword_count(0.0135). That is not the whole story. word_countswings by 350 between the two essays (150 to 500).unique_word_ratioswings by only 40 (70 to 30). Multiply each coefficient by its actual swing, and word count's contribution to the score wins, even though its per-unit weight looks smaller.- This is a realistic replay of the actual, published criticism of early AES systems: because length correlates with score in typical human-graded training data, a model can end up rewarding length more than it rewards the quality signals it was meant to detect.
Common mistakes
Reading a single coefficient's size as "how much this feature matters." As shown above, a feature's real effect depends on both its coefficient and its typical range of values. Comparing coefficients directly, without accounting for scale, is a common and misleading mistake — see linear regression.
Training only on features a student could easily learn to fake. Word count and vocabulary ratio are exactly the kind of surface feature a student could pad without writing anything better. Real production systems add many more features specifically to make simple gaming harder, though none eliminate the problem entirely.
Treating a high R-squared as proof the model understands quality. 0.964 here reflects how well the model fits ten training rows, not whether it has learned to detect genuine insight or argument quality. A model can fit surface patterns very well while still being fooled by writing designed to exploit exactly those patterns.
Try it yourself
Add a third feature, grammar_errors, to a few of the training rows (make the long, repetitive essay have a high error count) and refit the model. Check whether the model's predicted gap between the two example essays narrows. This is exactly the kind of feature-engineering escalation real AES research has gone through, one gaming discovery at a time.
What to learn next
- Linear regression — full detail on reading and misreading model coefficients.
- Text classification — the general NLP technique underlying most real AES systems.
- Model evaluation — how AES systems are actually checked against human raters.
Researcher — Mathematics and papers.
Feature-based and neural approaches
Classical AES (e.g. ETS's e-rater, Pearson's Intelligent Essay Assessor) extracts a feature vector per essay — length, syntactic complexity, grammar and mechanics error counts, discourse structure markers, vocabulary sophistication measures — and fits a regression or ordinal classification model against human scores. Neural approaches (Taghipour and Ng, 2016, an early LSTM-based AES model; and subsequent transformer fine-tuning approaches) instead learn features directly from token sequences, typically outperforming feature-based systems on held-out quadratic weighted kappa while being harder to audit for which textual properties actually drove a given score.
Evaluation metric
AES model quality is standardly reported as quadratic weighted kappa (QWK) against human rater scores:
kappa = 1 - (sum_ij w_ij * O_ij) / (sum_ij w_ij * E_ij)O_ij— observed count of essays scorediby the model andjby the human rater.E_ij— expected count under an assumption of independence between model and human scores.w_ij = (i - j)^2 / (N - 1)^2— a quadratic penalty weight, penalising large disagreements more than small ones, whereNis the number of score categories.
QWK is used specifically because it accounts for chance agreement and penalises the size of a disagreement, not only whether two raters agreed exactly — appropriate for the ordinal, multi-point score scales typical of essay grading.
The BABEL generator and adversarial gaming
Powers et al. (2001), in a study conducted with ETS's own e-rater system, built a generator (nicknamed BABEL) producing syntactically well-formed, semantically nonsensical essays specifically designed to exploit known e-rater feature weightings, and reported these nonsense essays received scores substantially higher than their content warranted. Perelman (2014) and related later work extended and publicised similar critiques against subsequent AES systems, arguing surface-feature-based scoring remains structurally vulnerable to essays optimised against its known feature set, regardless of genuine content quality. These are not isolated anecdotes — they are the AES field's own documented adversarial-robustness literature, and any claim that a specific deployed AES system is immune to gaming needs to be checked against published red-teaming results for that specific system, not assumed.
Fairness and bias
Multiple studies (e.g. Bridgeman et al., 2012, on e-rater and TOEFL) have examined systematic score differences by first-language background and writing style, since a model trained predominantly on one variety of English can systematically under- or over-score writing that deviates from that variety in structurally consistent, non-error ways. This is an active fairness concern distinct from the gaming problem above, and a documented reason several major testing programmes retain a human rater as a check or override on the automated score rather than relying on AES alone, particularly for population subgroups underrepresented in the AES training data.
Papers
- Page, E. B. (1966). The Imminence of Grading Essays by Computer. Phi Delta Kappan. The earliest AES system, Project Essay Grade.
- Powers, D. E. et al. (2001). Stumping E-Rater: Challenging the Validity of Automated Essay Scoring. ETS Research Report — the BABEL generator study.
- Taghipour, K., Ng, H. T. (2016). A Neural Approach to Automated Essay Scoring. EMNLP.
- Bridgeman, B., Trapani, C., Attali, Y. (2012). Comparison of Human and Machine Scoring of Essays: Differences by Gender, Ethnicity, and Country. Applied Measurement in Education.
- Perelman, L. (2014). When "the state of the art" is counting words. Assessing Writing.
Current state
AES is deployed at scale for large-standardised-test writing sections, essentially always alongside a human rater rather than as sole scorer, with automated-human disagreement above a threshold routed to additional human review. Transformer-based scoring is now common in research and increasingly in production, improving raw agreement with human scores while not resolving the underlying gaming and fairness concerns documented above, which are properties of learning from surface textual correlation with human scores rather than properties of any one model architecture.
What to learn next
- Text classification — general technique underlying most AES feature and neural pipelines.
- Model evaluation — agreement metrics and their proper use, in general.
- Why AI-text detectors do not work — a related, harder problem in the same space.