AI Can Grade a Million Essays Overnight. That Doesn't Fix Testing.
· AI Evaluation & Safety · 8 min read
Scale is the easy problem. Construct validity is the one no smarter model will fix.
The analysis
Automated grading solves a genuine constraint. Marking at scale is slow, expensive and inconsistent between human raters, and a model that produces defensible scores overnight removes a real bottleneck from any large assessment system.
What it does not solve is whether the thing being scored is the thing you care about. Construct validity — the question of whether the test measures the capability it claims to measure — is a property of the assessment design, not of the marker. Automating the marker faster does not make an essay prompt a better measure of reasoning; it makes a possibly poor measure cheap enough to apply universally.
The second failure mode is reflexive. Once scoring is automated and its behaviour is stable, teaching adapts to it. Students and institutions optimise for what the grader rewards, which is observable surface structure far more reliably than it is insight. The measure becomes the target and its usefulness as a measure decays — an old dynamic, but one that automation accelerates because the feedback loop is now instant and universal rather than slow and local.
The recommendation for anyone deploying this in education or in corporate certification is to invest the savings from automation back into assessment design and periodic human audit of a sampled subset, rather than banking the whole saving. Otherwise the organisation buys speed and quietly loses the ability to know what its scores mean.
Full essay on Substack: AI Can Grade a Million Essays Overnight. That Doesn't Fix Testing..
More on AI Evaluation & Safety
- Hiring AI Promised Objectivity. It Delivered a Single Point of Failure.
- I'm Actually Glad UC Berkeley Failed Them
Discuss this with Prasanna: pw@prima-partners.com · book a 30-minute call.