© 2026 FUTURE PROOF™
AI & Tutoring · Automated Scoring

Can AI grade writing?

Machine scoring of essays is older than the web, its agreement with human raters matches human-to-human agreement, and its best-known critic gamed it with eloquent nonsense. All three facts are true. What five decades of automated-scoring research actually established, and how Future Proof™ scores open responses without pretending the hard parts are solved.

TL;DR

The finding: On standardized essay tasks, automated scoring engines agree with trained human raters roughly as well as those raters agree with each other — demonstrated publicly in a large multi-vendor benchmark. Automated writing feedback measurably improves student revision and writing quality. But engines score proxies of quality, can be gamed by sophisticated nonsense, and inherit whatever the human ratings they imitate contained.

The mechanism: Scoring engines model the features that co-vary with human judgments. Within the distribution they were built for, the proxies track quality tightly; outside it — adversarial writing, novel tasks, unusual voices — the proxies and the construct come apart. The critical variable is not the algorithm but the validation regime around it.

The product: Future Proof uses automated scoring where its evidence is strongest — formative feedback, low-stakes checks, first-pass scoring with human escalation — anchored to expert-built rubrics, monitored for drift, with confidence routing anything anomalous to people.

In this article

  1. 01Fifty years of the same idea
  2. 02The Perelman problem
  3. 03How the engines actually work — and why it matters to buyers
  4. 04Feedback: the use case with the cleanest win
  5. 05What changes when open response gets cheap
  6. 06The validation frame that settles the argument
  7. 07What the evidence doesn’t show
  8. 08What this means for practice
© 2026 FUTURE PROOF™
The route. 8 sections, from “Fifty years of the same idea” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Ask what someone knows, and a machine can score the answer. Ask them to show what they know — in their own words, structured their own way — and for fifty years the scoring required a person. That asymmetry quietly shaped everything about how firms measure. It shaped what gets asked, how often, and at what depth. It decided which skills get certified on recognition alone, because production was too costly to read.

Open response is where assessment goes to get expensive. Multiple choice scales to millions but caps what can be asked. The moment you want a candidate to explain a decision, summarize an incident, or argue a position, a human must read it. And human reading costs money, takes days, and, as our rater-effects review documents, disagrees with itself. The dream of machine-scored writing is as old as educational computing. It has spent fifty years swinging between overclaim and dismissal.

The swing has costs in both directions. Firms that believed the overclaim deployed unvalidated engines into weighty decisions. Firms that believed the dismissal kept paying the open-response tax — or, worse, stopped asking open questions at all. Both extremes are wrong in telling ways. The measured position — what the benchmarks, the meta-analyses, and the critics each actually showed — is one of the more useful things a test buyer can carry. Writing is precisely the format modern AI has made cheap to score, and cheap is where discipline matters most.

Fifty years of the same idea

The idea’s age is its most underrated credential. Machine essay scoring predates the internet, the personal computer, and most of the people now arguing about it. So its core claims have survived not one hype cycle but several — each with better technology and the same basic questions.

The founding demo came in 1966. Ellis Page showed that a computer regression over crude text features — essay length, word variety, punctuation — could predict teachers’ essay grades disturbingly well, and he predicted that machine grading was imminent (Page, 1966). He was right about the correlation and early by several decades on the deployment. The modern engines that grew from that line score grammar, organization, coherence, and content against models trained on human-scored essays. They became working infrastructure in large testing programs, with published engineering and validity documentation (Attali & Burstein, 2006).

The number

1966 The year Ellis Page showed a computer regression over crude text features — length, word variety, punctuation — could predict teachers’ essay grades disturbingly well. Right about the correlation; early by several decades on the deployment (Page, 1966).

The public benchmark arrived in 2012: a contest across eight essay sets and many vendor and open engines. The headline result: machine scores agreed with human raters at levels on par with human–human agreement — on some prompts slightly better, on others slightly worse (Shermis & Hamner, 2013). That result deserves its precise reading. It does not say machines understand writing. It says that on standardized prompts, with training data, machines reproduce the scores trained humans give about as reliably as a second trained human does. Given what our rater-effects review says about human consistency, this is a meaningful bar and a modest one at the same time.

The Perelman problem

Every measurement technology gets the critic it deserves, and automated scoring got a superb one. The best-known critique came from MIT’s Les Perelman, who showed that scoring engines could be reliably gamed. Essays built to be long, ornate in vocabulary, and conventional in shape scored highly while being empty or factually absurd. His machine-written gibberish generators could produce top-scoring “essays” on demand (Perelman, 2014). The stunts were showmanship in service of a serious point: the engines score correlates of writing quality, and correlates can be made without the quality. Length alone carries an awkward share of the predictive weight in many systems.

The catch

The engines score correlates of writing quality, and correlates can be manufactured without the quality — Perelman’s gibberish generators produced top-scoring “essays” on demand (Perelman, 2014). The hazard is adversarial: it bites hardest where test-takers have the means, motive, and coaching to optimize against the scorer, which is why stakes decide the deployment tier.

The right lesson is narrower than “machines can’t grade”. Proxy-gaming is a problem of adversarial settings. It matters where test-takers have the means, motive, and coaching to optimize against the scorer — high-stakes admissions, certification. It matters far less in formative use, where gaming the feedback engine defeats the learner’s own purpose. The design responses are also known: mix construct-relevant features, audit how much length drives scores, red-team the engine with hostile entries, and keep humans on anomalies. The critique sets the engineering bar; it does not void the benchmark results.

How close engine scores come to a second human rater human–human parityessay set 1 essay set 2 essay set 3 essay set 4 essay set 5 essay set 6 essay set 7 essay set 8 collapses — far off this scale engineered text −10 −5 0 +5 +10 points of agreement above or below parity © 2026 FUTURE PROOF™
Figure 1. The public benchmark’s finding, drawn as distance from parity: across its eight essay sets, engine–human agreement sat about level with human–human agreement, a little above on some prompts and a little below on others (Shermis & Hamner, 2013). Adversarial, engineered text is the exception that matters — there the engine’s verdict and a reader’s part company (Perelman, 2014). Per-set positions are illustrative rather than measured; the published result is parity on average with prompt-to-prompt variation in both directions. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

How the engines actually work — and why it matters to buyers

A buyer does not need the math, but the engine’s design matters because it predicts the failure modes. The classical engines are feature machines: they extract measurable traits — length and its disguises, fancy vocabulary, varied syntax, discourse markers, word overlap with high-scoring essays. A fitted model then maps those traits to the human scores in a training set (Attali & Burstein, 2006). Everything such an engine knows about quality passes through those features. That is precisely why Perelman’s ornate nonsense worked: he fed the features without the substance. The engines are honest about being correlational; the trouble begins when deployments treat correlation as comprehension.

LLM-based scoring changes the design, and with it the failure profile. The model reads the answer against a rubric rather than counting engineered features. The gaming vector weakens: empty eloquence no longer reliably fools a model that can paraphrase the argument back. New vectors open in its place — sensitivity to prompt wording, drift across model versions, the odd confident misreading, and the chance of instructions hidden in submitted text. The validation framework absorbs all of this unchanged, because it never depended on the design: agreement on your population, generalization checks, adversarial testing, escalation paths (Williamson et al., 2012). What changes is only which tests the framework makes you run.

Feedback: the use case with the cleanest win

Grading — the final verdict — is where machine scoring gets debated. It is not where the technology has done its best work, and the difference matters for anyone deciding what to deploy first. While the grading debate absorbed the attention, the quieter application piled up the better evidence.

Automated writing evaluation as feedback — instant, specific, revision-focused comments during drafting — shows measured gains in how students revise and in writing quality across studies (Stevenson & Phakiti, 2014), (Wilson & Roscoe, 2020). The gain comes precisely because it changes the economics of the practice loop. Learners draft, get comments in seconds rather than weeks, and revise while the draft is alive. The formative case rides every advantage this library documents — feedback while the question is warm, more practice cycles per unit of teacher time, low stakes damping both gaming and anxiety. And it dodges the adversarial hazard almost entirely.

The feedback results also answer the objection that machine comments must be shallow to be fast. What the studies measure is not the depth of any single comment but the effect of the loop — and the loop’s advantage is structural. A mediocre comment delivered during drafting beats an insightful one delivered after the grade is spent, for the same reasons our feedback review documents. Human writing teaching at its best remains better. Human writing teaching as it is actually available — one pass, days later, if at all — is the comparison that matters. Against it, the automated loop wins on the variable the learning depends on: cycles.

Short-answer scoring, the workhorse format between multiple choice and the essay, matured in parallel. Systems that match student answers against target concepts can score explanation-style items at scale. A systematic review logs steady accuracy gains across eras of the technology (Burrows, Gurevych & Stein, 2015). And the LLM era has moved both uses again. Modern models score from rubric-conditioned prompts rather than per-prompt training sets, write the explanations older engines could not, and extend to formats — case analyses, reflective writing — the feature-built generation never handled. The capability jump is real; the validation duties it inherits are exactly the ones this article describes, now applied to systems whose failure modes are newer and less mapped.

What changes when open response gets cheap

The strategic consequence of workable machine scoring is not that essays get graded faster. It is that test design stops being constrained by scoring cost. For decades, the dominance of multiple choice was an economic fact wearing a teaching costume. Recognition items were what could be scored at scale, so curricula measured recognition. Bloom’s upper levels — explain, analyze, evaluate — lived in the small corners a human reader could afford. Every argument in our Bloom’s-taxonomy review about testing ceilings was, underneath, an argument about the price of reading.

Cheap open-response scoring repeals that constraint. Explanation can become a routine item type rather than a luxury. The self-explanation effect our tutoring cluster documents becomes usable as a test, not just as a study technique. The writing-to-learn loop — draft, feedback, revise — can run on the same cadence as flashcards. The firms that benefit first will not be the ones that bolt an essay scorer onto existing tests, but the ones that redesign what they ask because the answer format is finally affordable. The technology’s deepest effect is on the question bank, not the grading queue.

The same repeal applies to a measurement gap this library keeps flagging: the difference between recognizing an answer and producing one. Recognition items systematically overstate usable knowledge — the fluency illusions of our testing-effect review live partly in the format. Production items were always the honest measure nobody could afford to score. That excuse has expired. With it goes some of the comfort of high pass rates built on recognition alone.

The validation frame that settles the argument

After the benchmarks and the stunts, the field needed an adult in the room, and the measurement field supplied one. The mature position is neither camp’s slogan but a framework: a machine score is an interpretation that needs evidence. The evidence has named parts: agreement with trusted human judgment on the target population, and generalization across prompts and demographic groups. It also demands resistance to construct-irrelevant strategies, and written procedures for the cases the engine cannot score (Williamson, Xi & Breyer, 2012).

Under this frame, “can AI grade writing?” dissolves into answerable questions. This engine, this task, this population, these stakes — validated how, monitored how, with what escalation path? Vendors who answer those questions are selling measurement. Vendors who answer with the 2012 benchmark alone are selling a headline. The difference between the two is where every deployment’s risk actually lives.

Feedback latency, and the cycles it buysTime from submission to feedback automated secondshuman marking days to weeks 1 s 1 min 1 hour 1 day 1 week log scaleCycles per assignment window human loop 1 cycleautomated loop 4+ cycles 0 1 2 3 4 5 draft–feedback–revise cycles per assignment © 2026 FUTURE PROOF™
Figure 2. The formative win is arithmetic rather than eloquence. Machine commentary returns in seconds where human marking returns in days to weeks — note the log scale on the upper axis — so the same assignment window holds several draft–feedback–revise cycles instead of one, and cycles are the variable the learning depends on (Stevenson & Phakiti, 2014), (Wilson & Roscoe, 2020). Latencies are the contrast the studies describe; the cycle counts are illustrative — the measured claim is more cycles, not a particular number. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
The machine reproduces the scores trained humans give — which is a meaningful bar and a modest one at the same time. The precise reading of the automated-scoring benchmarks, after Shermis & Hamner (2013).

What the evidence doesn’t show

  • It doesn’t show machines judge originality or truth. Engines validated against pooled human scores inherit the pool’s insensitivities; distinctive voices and unconventional-but-excellent arguments are where machine and construct part company, and where human review earns its cost (Perelman, 2014).
  • Agreement numbers are prompt-specific. Engine performance in the benchmark varied across essay sets; a validation on one task family does not transfer to another without new evidence (Shermis & Hamner, 2013).
  • Bias auditing is not optional. Engines trained on human ratings can reproduce and systematize human raters’ group-level patterns; subgroup agreement analysis is a named requirement of the validity framework, not an enhancement (Williamson et al., 2012).
  • LLM fluency is not LLM validity. That a modern model writes persuasive commentary does not establish that its scores track the rubric; the plausible-sounding wrong score is the new failure mode, and it requires the old validation machinery plus drift monitoring.

Where the evidence stops

  1. 1It doesn’t show machines judge originality or truth
  2. 2Agreement numbers are prompt-specific
  3. 3Bias auditing is not optional
  4. 4LLM fluency is not LLM validity
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

The workforce translation of all this is concrete. The knowledge your teams most need to show — why this escalation path, what this clause requires, how you would triage this incident — has always been explanation-shaped. Your testing program has been standing in for it with recognition items, because reading ten thousand explanations was impossible. It no longer is.

Deploy by stakes, in three tiers. In formative use — practice, drafting, low-stakes checks — machine scoring and feedback are the closest thing this library has to a free lunch. They buy more writing practice and faster loops, at costs that make open response affordable everywhere multiple choice used to be the ceiling. In moderate-stakes internal testing, run the engine as first reader with confidence-based routing. Watch the spread of scores, escalate anomalies and near-cut cases to humans, and run regular blind double-scoring to catch drift. In high-stakes decisions, the engine assists and humans decide. The engine’s real job there is consistency pressure on the humans — whose own agreement numbers, as this library keeps noting, are nothing to protect.

Sequence the rollout by evidence strength. Formative feedback goes first, where the win is cleanest and the risk lowest. Internal assessment comes second, once the monitoring habits exist. High-stakes goes last, if at all, and never without the human layer. Firms that run the sequence backwards — leading with the high-stakes rollout because that is where the cost savings headline — are choosing the setup with the weakest evidence and the strongest adversaries.

Whatever the tier, buy validation, not vibes. Require agreement numbers on your people and your tasks, not the vendor’s showcase; require the score–length analysis; require subgroup reporting; require the escalation path in writing. And keep the construct honest at the design stage. Decide what the writing task is actually measuring — the reasoning, the communication, the knowledge — and check that the engine’s features touch that construct rather than its packaging. Machine scoring is no longer the question. Machine scoring governance is — and the firms that get value from the technology are the ones that inherited the testing industry’s validation habits along with its tools.

Applied research

How Future Proof™ applies this: tiered scoring, human escalation.

Open-response items in the platform are scored against expert-built rubrics, with the deployment tiered exactly as the evidence prescribes: instant formative feedback in practice, first-pass scoring with confidence-based human escalation in internal assessment, and human decision-making wherever stakes are high. Every engine is validated on the organization’s own tasks before it counts, score-length dependence and subgroup agreement are monitored continuously, and anomalous or adversarial submissions route to people. The writing gets scored in seconds. The pretending doesn’t happen at all.

See the AI Engine
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 1966Page
  • 2006Attali
  • 2012Williamson
  • 2013Shermis
  • 2014Perelman
  • 2014Stevenson
  • 2015Burrows
  • 2015Graham
  • 2020Wilson
© 2026 FUTURE PROOF™
The evidence base. The 9 sources cited here span 1966–2020, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Page, E.B. (1966). The imminence of… grading essays by computer. Phi Delta Kappan 47(5): 238–243. PDF
  2. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater V.2. Journal of Technology, Learning, and Assessment 4(3). PDF
  3. Shermis, M.D., & Hamner, B. (2013). Contrasting state-of-the-art automated scoring of essays. In Handbook of Automated Essay Evaluation (Shermis & Burstein, eds.): 313–346. PDF
  4. Perelman, L. (2014). When “the state of the art” is counting words. Assessing Writing 21: 104–111. PDF
  5. Stevenson, M., & Phakiti, A. (2014). The effects of computer-generated feedback on the quality of writing. Assessing Writing 19: 51–65. PDF
  6. Wilson, J., & Roscoe, R.D. (2020). Automated writing evaluation and feedback: Multiple metrics of efficacy. Journal of Educational Computing Research 58(1): 87–125. PDF
  7. Burrows, S., Gurevych, I., & Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education 25(1): 60–117. PDF
  8. Williamson, D.M., Xi, X., & Breyer, F.J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice 31(1): 2–13. PDF
  9. Graham, S., Hebert, M., & Harris, K.R. (2015). Formative assessment and writing: A meta-analysis. The Elementary School Journal 115(4): 523–547. PDF
Try the AI engine

Make open response affordable everywhere.

Book a 20-minute demo. We’ll score real writing from your domain against your rubric — and show you the validation report, not just the scores.

9 citations Reviewed August 2026 Open peer review welcomed

Where this shows up in the platform