© 2026 FUTURE PROOF™
Assessment Science · Formative Assessment

Formative assessment: the 0.4 that became 0.2.

Inside the Black Box turned formative assessment into a global movement on the strength of a famous range — effect sizes of 0.4 to 0.7. When a meta-analysis finally pooled the qualifying studies, it found roughly 0.20. Both numbers are honest answers to different questions, and the gap between them is the whole practical story.

TL;DR

The finding: Black and Wiliam’s 1998 review — and its pamphlet, Inside the Black Box — made formative assessment a movement on a claimed range of roughly d ≈ 0.4–0.7. The first formal meta-analysis (Kingston & Nash, 2011) pooled the studies that qualified and found ≈ 0.20. The variation by subject and delivery was large enough to swallow the average. The label now covers practice from excellent to inert.

The mechanism: The versions that work share four features. Short elicit–interpret–act loops run minute to minute. Feedback aims at the task, the process and the learner’s self-regulation rather than the person. Checks double as retrieval practice. And teachers get sustained development rather than purchased products. The versions that don’t work share the inverse — generic tools, grades standing in for feedback, and evidence that is collected but never used.

The product: Future Proof™ builds the working version in by default — every learning interaction is a retrieval-based check, feedback is generated at the process level rather than the person level, and the next step adapts to what the check found, so the formative loop closes automatically instead of heroically.

In this article

  1. 01The pamphlet that moved a policy world
  2. 02Where 0.4–0.7 actually came from
  3. 03The meta-analysis that halved it
  4. 04The definitional sprawl
  5. 05Feedback: the engine and its failure modes
  6. 06The versions that demonstrably work
  7. 07What the evidence doesn’t show
  8. 08Formative practice by the evidence
© 2026 FUTURE PROOF™
The route. 8 sections, from “The pamphlet that moved a policy world” to “Formative practice by the evidence”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Few documents in modern education have moved as much policy per page as Inside the Black Box. The pamphlet appeared in 1998 as the public face of a long research review. In it, Paul Black and Dylan Wiliam argued that minute-to-minute classroom assessment was one of the strongest levers for raising achievement — and that it was being wasted (Black & Wiliam, 1998). The lever they meant: the questions teachers ask, the feedback they give, the use they make of what they learn about their students. The claim came with numbers attached. Programmes that strengthened formative assessment produced typical gains of roughly 0.4 to 0.7 on the effect-size scale — the standard measure of impact — larger than almost anything else education research had on offer (Black & Wiliam, 1998).

Ministries listened. “Assessment for learning” entered national strategies, district improvement plans and vendor catalogues. Formative assessment became that rarest of things — a research phrase every working teacher can quote.

Thirteen years later came the first formal meta-analysis of formative-assessment interventions — a study that pools the results of all the qualifying studies. It reported a pooled effect of roughly 0.20, about half the bottom of the famous range (Kingston & Nash, 2011). And the variation underneath was large enough to swallow the headline entirely.

This article is about what happened between those two numbers. Where did 0.4–0.7 actually come from? What did the 2011 meta-analysis measure, and what did it miss? Why did the sharpest critique in the literature conclude that the question itself was badly posed (Bennett, 2011)? And what does the surviving evidence say separates the formative practice that moves learning from the much larger set of things sold under the same name?

The pamphlet that moved a policy world

In 1998, Black and Wiliam published two versions of the same argument. The scholarly version was “Assessment and classroom learning,” a review running to nearly seventy pages in Assessment in Education. It synthesised roughly 250 studies of classroom assessment practice: feedback, questioning, self- and peer-assessment, mastery-learning approaches (Black & Wiliam, 1998). The public version, Inside the Black Box, squeezed the review into a short pamphlet for teachers and policymakers. Its title named the problem it saw. Standards policy obsessed over inputs and outputs while treating the classroom processes between them as an unexamined black box (Black & Wiliam, 1998).

The argument had two central claims. First, there was firm evidence that strengthening formative assessment raises achievement. The authors put the typical gains at effect sizes of roughly 0.4 to 0.7 — at the top end, among the largest effects ever reported for a sustained educational programme (Black & Wiliam, 1998). Second, prevailing practice was poor: marking assigned scores without guidance. Questioning left wait times too short for thinking to happen. Feedback invited students to compare themselves with each other rather than to improve the work (Black & Wiliam, 1998).

What followed was one of the fastest research-to-policy transfers on record. England built “assessment for learning” into statutory guidance. American districts wrote it into improvement plans. A whole industry of teacher development grew up around the phrase, and assessment vendors discovered that the word “formative” sold products. Almost none of this machinery carried the review’s hedges along with it. What got carried forward, mostly, was the number.

Where 0.4–0.7 actually came from

It is worth being precise about what the famous range was — and what it was not. It was not the output of a meta-analysis. The 1998 review explicitly declined to pool its studies. The reason: the underlying literatures were too mixed — different interventions, different populations, different outcome measures — for a single average to mean anything defensible (Black & Wiliam, 1998). The 0.4–0.7 figure was a description in words, not a computed average. Across these varied literatures, the authors judged, effects of roughly this order kept appearing.

As an orientation, that was legitimate — and the review said so carefully. But second-hand citation stripped the hedges within a few years. The pamphlet had offered an illustration — gains of this size would lift a country markedly in international comparisons (Black & Wiliam, 1998) — and the illustration was repeated as a promise. Bennett later traced where the range came from. He found it resting on studies of feedback, mastery learning and classroom questioning — studies that were never “formative assessment programmes” in any working sense (Bennett, 2011). So the review that made the number famous never justified treating 0.4–0.7 as the expected return on buying something called formative assessment.

None of this makes the underlying practice hollow. It means something narrower and more uncomfortable: for its first decade as a movement, formative assessment was running on a headline number that no study of the thing itself had produced.

The meta-analysis that halved it

Kingston and Nash set out to compute the missing number (Kingston & Nash, 2011). They searched the K-12 literature for studies of interventions described as formative assessment. Then they applied the standard filter: a measurable contrast — control group or equivalent design — from which an effect size could be computed. Hundreds of candidate studies had piled up over the boom years. Only around a dozen qualified, yielding roughly forty effect sizes. That was the first striking finding, before any pooling happened at all: a movement of this size was standing on an evidence base this thin.

The number

≈ 12 studies Out of hundreds of candidates from the boom years, roughly a dozen qualified for the first formal meta-analysis — around forty effect sizes in total (Kingston & Nash, 2011). The movement was larger than its evidence base by orders of magnitude.

The pooled estimate came out at roughly 0.20, with a median in the same neighbourhood (Kingston & Nash, 2011). Read plainly: real, positive, and worth having. Plenty of things schools spend money on do worse. But it is roughly half the floor of the range the movement had been quoting for a decade.

The moderators — the factors that split the average — mattered more than the mean. By subject, effects for English language arts ran around 0.3, mathematics around 0.17, and science smaller still — on the fewest estimates (Kingston & Nash, 2011). By delivery route, programmes run through sustained teacher development came in around 0.3, beating the other routes (Kingston & Nash, 2011). When the spread is that large next to the pooled mean, the average stops being the interesting number. What you buy, in which subject, delivered how, matters more than whether the label is present.

The second half of the paper’s title — “a call for research” — was the authors’ own reading of the situation. Not that formative assessment fails. Rather, the evidence base was thin, uneven, and unable to license the claims being made on its behalf (Kingston & Nash, 2011). Their estimate has been argued over ever since: which designs should have qualified, whether the inclusion rules were too loose or too strict. The definitional critique below explains why that argument cannot be settled as posed (Bennett, 2011). But the reset stands: when the measured versions of the label were pooled, they delivered roughly 0.2, not 0.4 to 0.7.

Black & Wiliam 1998 (claimed range) ≈0.4–0.7 Kingston & Nash 2011 (pooled) ≈0.20 English language arts ≈0.32 Mathematics ≈0.17 Science (fewest studies) ≈0.09 0 0.2 0.4 0.6 0.8 Approximate standardised effect sizes — claim vs. pooled measurement © 2026 FUTURE PROOF™
Figure 1. The claim and the measurement: the range that built the movement, the pooled estimate that finally tested it, and the subject-level spread that swallows both. Schematic; values are approximate. The 0.4–0.7 range was a narrative characterisation across heterogeneous literatures, not a pooled estimate (Black & Wiliam, 1998); the ≈0.20 pooled figure and its subject moderators come from a small qualifying study pool with contested inclusion rules (Kingston & Nash, 2011). Read both as orientation, not precision. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

The definitional sprawl

Bennett’s 2011 critical review is the closest thing the field has to an audit, and its central diagnosis is about definition (Bennett, 2011). “Formative assessment” names a process: the minute-to-minute cycle in which a teacher elicits evidence of thinking, interprets it, and acts on it. But it also names a product category: the interim and benchmark testing systems marketed under the same adjective. One term is being asked to cover both a teacher’s next question and a district’s quarterly test. A category that broad cannot have a single effect size, because it is not a single thing (Bennett, 2011).

The sprawl is not a cosmetic problem. It is the machinery by which evidence earned in one corner of the category gets spent in another. Claims generated by studies of classroom questioning and feedback are inherited, in marketing, by products those studies never examined. Both headline numbers mislead if read as properties of the label. The 0.4–0.7 drew together literatures that were never programmes. The ≈0.20 averaged across variants that share little besides the name (Bennett, 2011) (Kingston & Nash, 2011).

The catch

The sprawl is the mechanism of the overclaim: evidence earned by classroom questioning and feedback gets spent, in marketing, by benchmark-testing products those studies never examined. Neither headline number — the 0.4–0.7 nor the ≈0.20 — is a property of the label on the box.

Bennett’s second diagnosis follows from the first: formative assessment depends on the subject. Acting on evidence means knowing what a given error means and what to do about it next — what researchers call pedagogical content knowledge, subject-specific teaching skill, not generic technique (Bennett, 2011). A teacher can run the surface moves — traffic lights, exit tickets, mini-whiteboards — and still have nothing formative happen. The formative part is the interpretation and the response. This is also the simplest reading of the delivery pattern in the meta-analytic record. Development-heavy programmes beat tool purchases because the expensive ingredient is judgment (Kingston & Nash, 2011).

Feedback: the engine and its failure modes

Strip the label away, and the engine inside every formative loop is feedback plus adjustment. The feedback literature is the whole story in miniature. Hattie and Timperley’s synthesis put the average effect of feedback among the largest in education — roughly d ≈ 0.8 across the accumulated meta-analyses (Hattie & Timperley, 2007). It also carried some of the widest variability on record, including a sizeable minority of cases in which feedback made performance worse. A tool whose average is excellent and whose variance includes harm is not a tool you can buy by name. It is a design space.

Their model gives the design its coordinates. Feedback operates at four levels — the task, the process behind the task, the learner’s self-regulation, and the self (Hattie & Timperley, 2007). The first three carry the effects. Feedback aimed at the self — praise, person-level judgment — carries little, and it is the usual suspect when effects go negative. Effective feedback answers three questions: where am I going, how am I going, where to next. The third is the one most classroom marking never reaches (Hattie & Timperley, 2007).

For formative practice this is the operative law. The same generic activity — “giving feedback” — can be the best or the worst intervention in the building depending on where it points. Grades used as feedback pull attention to the self level and invite social comparison; comments about the work and the next step operate at the task and process levels (Hattie & Timperley, 2007). That is precisely the marking critique the original pamphlet led with (Black & Wiliam, 1998). The average effect of feedback is not a thing an organization can purchase. A specific feedback design is.

Why it matters

The same generic activity — giving feedback — can be the best or the worst intervention in the building depending on where it points. Aim at the task, the process, and the learner’s self-regulation; the person level is where the negative tail lives, and a grade used as feedback points there by default.

The versions that demonstrably work

Three studies, at three scales, show what the working variants have in common. The first is the founders testing their own claim. Wiliam, Lee, Harrison and Black followed two dozen secondary mathematics and science teachers through six months and more of structured development. The teachers built their own formative repertoires — questioning, comment-only marking, self- and peer-assessment. Outcomes were then measured on external tests and examinations. The mean effect was roughly 0.3 (Wiliam, Lee, Harrison & Black, 2004).

The study is notable for its honesty. It was run by the people with the most to lose, and measured on outcomes they did not control. And it landed near the meta-analytic estimate for development-based programmes — not the pamphlet range (Kingston & Nash, 2011) (Wiliam, Lee, Harrison & Black, 2004).

The second study finds the active ingredient at the grain of conversation. Ruiz-Primo and Furtak videotaped middle-school science-inquiry lessons. They coded the lessons into assessment conversations: the teacher elicits a response, the student responds, the teacher recognises what the response reveals, and then uses it to adjust the next move. Teachers who completed more full cycles — especially the final, using step — had students who performed markedly better on the unit’s outcomes (Ruiz-Primo & Furtak, 2007). The sample was small and the design correlational — a link, not proof of cause — so the estimate deserves hedging. But the location is the point: assessment becomes formative at the use step, not the asking step — evidence that is elicited and admired is not an intervention (Ruiz-Primo & Furtak, 2007).

The third supplies the memory floor underneath the whole practice. Roediger and Karpicke showed that retrieval is itself a learning event. Students tested on studied material retained far more a week later than students who spent the same time restudying it (Roediger & Karpicke, 2006). So frequent low-stakes checking earns part of its effect before any feedback arrives. A classroom thick with real questions is running retrieval practice under another name (Roediger & Karpicke, 2006). Put the three together and the shape is unmistakable: short loops, retrieval, interpretation, action — aimed at the work, at the level of the work, close to the moment of the work.

pooled label average Formative label, pooled Via teacher development Founders’ field trial Feedback (the engine) 0.20 0.30 0.30 0.80 0 0.2 0.4 0.6 0.8 1.0 Effect size (d) © 2026 FUTURE PROOF™
Figure 2. What the delivery route is worth. Pooled across everything sold under the label, formative assessment returns roughly 0.20; programmes delivered through sustained teacher development run around 0.3 (Kingston & Nash, 2011), and the founders’ own development-based field trial landed in the same neighbourhood (Wiliam, Lee, Harrison & Black, 2004). The engine inside every formative loop — feedback — averages about d ≈ 0.8 across the accumulated meta-analyses (Hattie & Timperley, 2007): the ceiling the label itself never reaches. Different literatures share one axis here, and the feedback average carries some of the widest variance on record, a minority of it negative. Read the ordering, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
Formative assessment: A meta-analysis and a call for research. Kingston & Nash, Educational Measurement: Issues and Practice, 2011

What the evidence doesn’t show

The honest reading of this literature is positive: feedback-rich, minute-to-minute formative practice works. But the boundaries do real work. Buyers should hold them as firmly as the headline:

  • A meta-analytic 0.4–0.7. The famous range was a narrative characterisation of heterogeneous literatures; the review that produced it explicitly declined to pool them (Black & Wiliam, 1998) (Bennett, 2011). Quoting it as the expected return on a purchase misreads its own source.
  • A settled 0.20. The pooled estimate rests on roughly a dozen qualifying studies with contested inclusion rules, skewed toward packaged, measurable implementations (Kingston & Nash, 2011). It resets the claim; it does not end the argument.
  • That products inherit the evidence. Interim and benchmark systems sold as “formative” are the least-tested variant of the category; the strongest evidence concentrates in classroom practice and teacher development (Bennett, 2011).
  • That feedback is safely positive. The feedback literature’s average is large, but a substantial minority of measured effects run negative, and person-level feedback is reliably the weakest (Hattie & Timperley, 2007). More feedback is not the prescription; better-aimed feedback is.
  • Durability and transfer. Outcomes in the qualifying studies are mostly proximal and short-horizon; whether formative gains persist across years or transfer across domains is essentially unmeasured (Kingston & Nash, 2011).
  • Subject equivalence. The gap between English language arts and mathematics estimates is large enough to change a buying decision; evidence from one subject does not price another (Kingston & Nash, 2011).

Where the evidence stops

  1. 1A meta-analytic 0.4–0.7
  2. 2A settled 0.20
  3. 3That products inherit the evidence
  4. 4That feedback is safely positive
  5. 5Durability and transfer
  6. 6Subject equivalence
© 2026 FUTURE PROOF™
The boundary. 6 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Formative practice by the evidence

Read as one body of work, the literature becomes an operating manual — shorter than the brand’s, and far more demanding.

Make every check a retrieval event. The memory benefit of answering real questions accrues even before feedback arrives (Roediger & Karpicke, 2006). Checks that merely re-present material — or that can be passed by recognition — spend the class’s time without buying the effect.

Aim feedback at the task, the process and self-regulation — never the person. That is where the large average effects live, and the person level is where the negative tail lives (Hattie & Timperley, 2007). Comments that answer “where to next” beat scores; scores used as feedback are the classic failure mode the original review documented (Black & Wiliam, 1998).

Close the loop or it wasn’t formative. The active ingredient is the use step — evidence must change the next instructional move: the next question, the next grouping, the decision to reteach (Ruiz-Primo & Furtak, 2007). Dashboards full of elicited-but-unused evidence are the process’s most common modern failure.

Buy development, not adjectives. The delivery pattern is the clearest buying signal in the meta-analytic record: development-based programmes at roughly 0.3 against a pooled 0.2 (Kingston & Nash, 2011). And the founders’ own field trial delivered its ≈0.3 through months of teacher development and ownership, not through tooling (Wiliam, Lee, Harrison & Black, 2004).

Evaluate the practice, not the label. The category is too sprawling for its average to price any specific programme (Bennett, 2011). Pre-specify the outcomes, expect subject-sized differences (Kingston & Nash, 2011), and treat any vendor quoting 0.4–0.7 as making a marketing claim about a number their product has never produced.

Applied at Future Proof

How Future Proof™ applies this.

The evidence says formative assessment works when checks are retrieval events, feedback points at the process rather than the person, and the loop actually closes. That is the platform’s default grammar: every learning interaction doubles as a low-stakes retrieval check. The feedback generated answers “where to next” at the task and process level, never as a grade-shaped judgment. And what each check finds changes what happens next — the following question, the spacing of review, the route through the material — automatically, so the use step no longer depends on heroic attention. The analytics report the loop itself: what was elicited, what changed because of it, and what it did to performance — not how much “formative” activity occurred.

See the platform
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 1998Black
  • 1998Black
  • 2004Wiliam
  • 2006Roediger
  • 2007Hattie
  • 2007Ruiz-Primo
  • 2011Kingston
  • 2011Bennett
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 1998–2011, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice 5(1): 7–74. DOI
  2. Black, P., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan 80(2): 139–148. PDF
  3. Kingston, N., & Nash, B. (2011). Formative assessment: A meta-analysis and a call for research. Educational Measurement: Issues and Practice 30(4): 28–37. PDF
  4. Bennett, R.E. (2011). Formative assessment: A critical review. Assessment in Education: Principles, Policy & Practice 18(1): 5–25. PDF
  5. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research 77(1): 81–112. PDF
  6. Wiliam, D., Lee, C., Harrison, C., & Black, P. (2004). Teachers developing assessment for learning: impact on student achievement. Assessment in Education 11(1): 49–65. PDF
  7. Ruiz-Primo, M.A., & Furtak, E.M. (2007). Exploring teachers’ informal formative assessment practices and students’ understanding in the context of scientific inquiry. Journal of Research in Science Teaching 44(1): 57–84. PDF
  8. Roediger, H.L., & Karpicke, J.D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science 17(3): 249–255. PDF
Try the AI engine

Assessment that closes its own loop.

Book a 20-minute demo. We’ll show you retrieval-based checks, process-level feedback, and learning paths that change because of what the last check found — the formative loop, closed by default.

8 citations Reviewed August 2026 Open peer review welcomed