© 2026 FUTURE PROOF™
The Uncomfortable Evidence · Training Evaluation

Smile sheets: why happy isn’t learned.

The training industry’s default success metric — did participants like it? — correlates with actual learning at roughly .09. The meta-analyses behind that number, what the Kirkpatrick pyramid got right and wrong, and what an evaluation stack looks like when it measures outcomes instead of applause. How Future Proof™ builds the honest version.

TL;DR

The finding: Meta-analytically, trainees’ reactions to a course — satisfaction, enjoyment, perceived usefulness — correlate with measured learning at around r ≈ .09, and with job behavior barely better. The four-level evaluation framework everyone cites assumed the levels form a causal chain; the correlational evidence found the links between them weak to absent. Most organizations evaluate almost exclusively at the level with the least signal.

The mechanism: Reactions measure the experience — the instructor’s charisma, the pacing, the lunch — while learning depends on effortful conditions that often depress enjoyment. The desirable-difficulties literature predicts the disconnect: what works can feel worse.

The product: Future Proof measures the levels that carry signal by default — retention over time, demonstrated skill, behavioral indicators — and keeps reactions in their proper role: diagnosing the experience, never certifying the learning.

In this article

  1. 01The four levels and the assumption inside them
  2. 02Why liking and learning come apart
  3. 03Why the industry kept the instrument anyway
  4. 04What honest evaluation looks like
  5. 05A worked example: one program, two verdicts
  6. 06The rehabilitated smile sheet
  7. 07What the evidence doesn’t show
  8. 08What this means for practice
© 2026 FUTURE PROOF™
The route. 8 sections, from “The four levels and the assumption inside them” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

At the end of most corporate training in the world, the same instrument appears: a short survey asking whether the session was useful, engaging, well-paced, and worth recommending. Its results decide whether the course runs again, whether the vendor is renewed, and whether the program appears in the quarterly deck as a success. The industry calls these reaction forms; the trade’s affectionate slur is smile sheets.

The stakes scale with the industry. Global corporate training spend runs to hundreds of billions a year, and for most programs, the reaction form is the only measurement that will ever be taken. That means the world’s training portfolio is steered, in aggregate, by an instrument whose link to learning is an empirical question with a very specific answer.

That question — do smiles predict learning? — was answered decades ago, with unusual clarity. The answer reshapes what training measurement should look like. This article walks the evidence: the framework that made reactions respectable, the meta-analyses that measured what they really tell us, and the evaluation stack that survives contact with both.

The four levels and the assumption inside them

To see how the industry got here, meet the framework that organized its thinking for sixty years. Its taxonomy was sound; its popular reading was not; its author bears limited blame for the difference. The framework is so common it functions as furniture. Donald Kirkpatrick’s four levels — reactions, learning, behavior, results — organized training evaluation from the late 1950s onward. As a list of what one could measure, it remains perfectly serviceable (Kirkpatrick, 1994).

The trouble was the assumption that traveled with it: that the levels form a causal ladder, each feeding the next. Enjoyable training produces learning; learning produces behavior change; behavior produces results — so measuring the bottom rung tells you about the ones above. Under that assumption, the smile sheet is a leading indicator. Without it, the smile sheet is a survey about the room.

The assumption was testable, and thirty years after the framework’s debut, researchers tested it. Pooled across the available studies, the correlations among levels ranged from weak to trivial — nothing like the causal chain practice presumed (Alliger & Janak, 1989). The follow-up meta-analysis put the flagship number on it. Reactions correlated with immediate learning at about r ≈ .07–.09 — for practical purposes, no relationship (Alliger, Tannenbaum, Bennett, Traver & Shotland, 1997). Utility-type reactions — perceived job relevance — did modestly better against later measures than affect-type reactions like enjoyment. But everything sat at sizes that forbid using any reaction as a proxy for any outcome.

The number

r ≈ .09 The meta-analytic correlation between trainees’ reactions and their measured learning — for practical purposes, no relationship, from the instrument most renewal decisions run on (Alliger, Tannenbaum, Bennett, Traver & Shotland, 1997).

A later, larger synthesis confirmed the verdict with more data and a sharper breakdown. Reactions relate mainly to affect — how people felt — and to motivation. They carry little information about cognitive learning, and they work best as measures of the training experience rather than its effects (Sitzmann, Brown, Casper, Ely & Zimmerman, 2008). None of this was hidden in obscure journals; it has been the settled quantitative position for a generation. The industry simply kept the smile sheet — cheap, immediate, flattering to everyone involved — because nothing in the renewal process ever asked for more. Instruments survive on their incentives, not their information.

What reactions actually predict (meta-analytic correlations, approximate)reactions ↔ immediate learning ≈ .09reactions ↔ job behavior ≈ .07learning ↔ job behavior ≈ .2 range(for scale: retest of the same exam) ≈ .8 © 2026 FUTURE PROOF™
Figure 1. The information content of the industry’s favorite metric: satisfaction predicts learning at roughly the level of noise. Approximate values from Alliger et al. (1997) and Sitzmann et al. (2008). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Why liking and learning come apart

A near-zero link between two things people assume go together demands a mechanism. This one has a good one — not sloppy measurement, but a genuine split between what the two gauges track. The result stops being surprising the moment this library’s other findings enter the room. Learning depends on effortful conditions — retrieval attempts, spaced struggle, interleaved confusion, errors with feedback — and those reliably depress in-session enjoyment. Our desirable-difficulties review shows learners rating the less effective conditions as better, study after study.

Meanwhile, the things that drive satisfaction — presenter charisma, production polish, pace without friction, jokes that land — are either neutral to learning or, per the seductive-details evidence, mildly negative. The smile sheet is not a noisy measure of learning. It is a fairly good measure of a different construct — one that sometimes anti-correlates with learning in exactly the region where courses are being tuned.

That last clause is the practical danger. A program whose renewal depends on reaction scores faces constant selection pressure toward comfort: fewer tests (they stress people), fewer struggles (they frustrate), more stories and gloss (they delight). Each step is invisible at the outcome level, because outcomes were never measured. Over several cycles the portfolio evolves toward maximum applause and minimum durable change — a drift no one chose and everyone funded. Learner self-report cannot stop the drift. Study after study shows learners misjudging their own learning in favor of what feels smooth.

The catch

A program renewed on reaction scores is under continuous selection pressure toward comfort: fewer tests, fewer struggles, more polish. Each step is invisible at the outcome level because outcomes were never measured — the portfolio drifts toward maximum applause and minimum durable change, and nobody chose the drift.

Why the industry kept the instrument anyway

Why level-one-only evaluation persists is no mystery once its incentives are listed — and listing them is the honest prelude to fixing them. The smile sheet is immediate: results the same afternoon, where learning measures need delays and behavior measures need the workplace. It is cheap — no item writing, no follow-up logistics, no baseline. It is diplomatically safe: it judges the vendor and the experience rather than the participants, so nobody’s competence is on the line. That is also why people fill it in honestly and at once. And it is reliably flattering — reaction scores cluster high, so the metric almost always supports the program, which suits the people whose budgets depend on the program being supported.

Notice that every advantage is real. The smile sheet genuinely is the fastest, cheapest, safest, most agreeable instrument available. It merely fails at the one property the renewal decision needs — information about whether anyone learned anything. So the reform problem is not persuading anyone the correlation is .09; the number has been public since 1997. It is re-engineering the decision so that convenience stops outranking validity.

That is an infrastructure story. Measurement beyond level one stayed rare while it was costly. The math changes exactly when retention checks and skill measures fall out of the learning system as by-products. This is the modern situation — and the reason this article’s advice is newly practical rather than forever pious.

What honest evaluation looks like

Criticizing the smile sheet is a decades-old sport. The field’s real contribution is the replacement, built and written up by the same researchers who measured the problem. The constructive tradition rebuilt evaluation from the learning-outcomes side. The influential taxonomy split outcomes into cognitive, skill-based, and affective, each with fitting measures and timing (Kraiger, Ford & Salas, 1993). The point: “did it work?” breaks into answerable questions with instruments — not into one survey.

The applied review of training science turns that into design advice (Salas, Tannenbaum, Kraiger & Smith-Jentsch, 2012). Specify intended outcomes before the training. Measure them after a delay, because immediate post-tests overstate durable learning — the entire message of our forgetting-curve review. Include behavior where you can. And treat evaluation as part of the training design, not an admin afterthought.

The training-effectiveness meta-analyses close the loop on why any of this is worth the effort. Training designed and evaluated against specified outcomes shows solid average effects on learning and job behavior — the investment genuinely works when built and measured properly (Arthur, Bennett, Edens & Bell, 2003). That makes the industry’s measurement habits a case of a functioning product wrapped in a broken gauge.

The transfer research adds the level most programs skip entirely. Whether a trained skill reaches the job depends, measurably, on what happens after the course — the chance to use it, manager support, follow-up cues. So a behavior-level check doubles as a read on the workplace, not just the course (Blume, Ford, Baldwin & Huang, 2010). And where the full causal question matters — whether the program caused the outcome — the evaluation inherits every guard our measurement articles describe. Comparison groups protect against regression artifacts; delayed measurement against fluency illusions; pre-specified metrics against post-hoc storytelling.

A worked example: one program, two verdicts

Consider a negotiation course as the stack sees it. Level one: 4.7 of 5, glowing comments about the facilitator. Level two, measured the old way — a quiz at the closing session: 88% average, everyone certified. Level two, measured honestly — the same items delivered cold three weeks later through the retrieval system: 61%, with the framework’s central sequencing rule recalled by fewer than half. Level three, from the CRM: discount depth in deals run by trained reps, unchanged against a matched comparison group. The program is simultaneously a triumph (by the deck it currently ships in) and inert (by everything that matters), and no one involved is lying — they are reading different instruments.

The rebuilt version of the same course starts from the delayed numbers. The decaying concepts get spaced reinforcement. The unused sequencing rule gets an in-workflow prompt at the moment of quoting. The next cohort’s evaluation pre-registers the three-week retention target and the CRM metric. Satisfaction stays at 4.7 or drops to 4.2, and nobody’s renewal hangs on it. That is the entire reform in miniature: same course, same budget, different question — what remained, and what changed? — asked by instruments that can actually answer it.

Design rule

Specify the intended outcomes before the training runs, and measure them after a delay — immediate post-tests overstate durable learning, and the delayed check is what separates a certified cohort from a capable one. Evaluation is part of the training design, not an administrative afterthought (Salas, Tannenbaum, Kraiger & Smith-Jentsch, 2012).

The rehabilitated smile sheet

The critique invites an overcorrection worth heading off: ripping the reaction form out entirely. That discards real information and needlessly angers everyone who relies on its legitimate uses. None of the evidence makes reaction data worthless. It makes it data about reactions — which has honorable uses once demoted from proxy to instrument. Reactions flag delivery failures fast: a module nobody understood, a broken pace. And utility judgments, the strongest sub-family in the meta-analyses, give a coarse early read on perceived job relevance (Alliger et al., 1997).

Reactions also predict attendance, completion, and word-of-mouth — outcomes an internal program rightly cares about. The rehabilitation has one rule: reaction data may inform the experience and may never certify the learning. A course can be pleasant and inert, unpleasant and transformative, and every combination between. Only outcome measurement says which. The entire reform of this article sits in refusing to let one instrument answer the other’s question.

Worked example — illustrative, not measured 100% 75 50 25 0 percent of scale as reported measured after a delay difference, not a level 4.7 of 5 88% 61% <50% 0 — no change Reactions end of session Quiz at close same day Same items, cold 3 weeks later Key rule recall 3 weeks later Deal behavior vs matched group © 2026 FUTURE PROOF™
Figure 2. One negotiation course, five instruments: four of them on a common percent-of-scale axis, the fifth fenced off to the right because it is a difference rather than a level. The two numbers the deck ships with sit near the top, and every measure taken after a delay or off the survey falls away — 4.7 of 5 on the reaction form, 88% on the closing quiz, 61% when the same items are delivered cold three weeks later, fewer than half recalling the framework’s central sequencing rule, and no measurable movement in discount depth against a matched comparison group. The last column sits outside the percent grid because it is a difference against a comparison group rather than a level on the same scale; its flat marker is a zero, not a short bar. Illustrative worked example from this article, not a measured study; it shows the shape the evidence predicts when reactions are read as an outcome proxy (Alliger et al., 1997) instead of measuring the specified outcomes after a delay (Salas, Tannenbaum, Kraiger & Smith-Jentsch, 2012). Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
A fairly good measure of a different construct — collected at the moment it is least informative, from the witness least able to judge. The smile sheet, per the meta-analytic record — Alliger et al. (1997), Sitzmann et al. (2008).

What the evidence doesn’t show

  • It doesn’t show experience is irrelevant. Misery has real costs — dropout, avoidance of future learning, brand damage to the L&D function. The argument is against reactions as an outcome proxy, not against caring how learners are treated.
  • It doesn’t invalidate the four levels as a menu. Kirkpatrick’s taxonomy still names the right measurement families; the correction is to the assumed chain between them, and to the practice of stopping at level one (Alliger & Janak, 1989).
  • Utility reactions earn a footnote. Perceived job relevance, asked concretely, carries modest signal about later use — the one corner of the smile sheet worth keeping in the evaluation conversation (Alliger et al., 1997).
  • Results-level measurement is genuinely hard. Attributing business outcomes to training is a causal-inference problem, not a dashboard problem; honest programs measure learning and behavior well rather than claiming ROI they cannot identify.

Where the evidence stops

  1. 1It doesn’t show experience is irrelevant
  2. 2It doesn’t invalidate the four levels as a menu
  3. 3Utility reactions earn a footnote
  4. 4Results-level measurement is genuinely hard
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

The transition plan matters as much as the destination. An evaluation reform announced as an audit of past programs will be fought as an accusation. The framing that works is forward-looking and even-handed. From the next cycle, every program — internal and vendor alike — pre-registers its intended outcomes and its delayed measures, and everyone’s numbers start from the same honest baseline. Grandfather the past; instrument the future.

Rebuild the evaluation stack in proportion to signal. Retention and demonstrated skill, measured after a delay, become the primary outcome — cheap now that spaced retrieval infrastructure exists, and immune to the fluency illusions that corrupt both smile sheets and immediate post-tests. Behavior reads come next, where the job makes them visible. Reactions stay — one screen, not five — reframed toward utility (“what will you use, where will it be hard?”) and read as experience checks. And the renewal decision, the point where measurement becomes money, moves from the applause metric to the outcome metrics — because whatever the renewal rewards is what the portfolio will evolve toward.

Set the delayed-measurement windows on purpose, not by calendar convenience. Put the first retention check where the forgetting curve does its steepest work — one to three weeks out. Put the second where the knowledge must actually survive — a quarter out. Key the behavioral reads to the job’s natural cycles: the next audit, the next release, the next negotiation season. Windows chosen this way turn each program’s evaluation into a retention curve rather than a single point, and curves are what the improvement conversation actually needs.

Expect two awkward finds in the first honest cycle. Some beloved, five-star programs will show flat retention curves — the drift this article describes, finally visible. And some middling-rated programs will show strong outcomes, usually the ones with the retrieval practice and the difficulty the ratings punished. Both are the system working. The training function that survives them has something rarer than good scores: evidence. And evidence, unlike applause, compounds.

Applied research

How Future Proof™ applies this: outcomes as the default metric.

The platform’s dashboards lead with the levels that carry information: retention curves from spaced retrieval, demonstrated skill against role maps, movement between diagnostic baselines — all collected automatically, all measured after the delays that make them honest. Reaction data is captured lightly and displayed in its lane, labeled as experience feedback, never aggregated into a “success score” with learning it cannot predict. When a program is renewed on Future Proof data, it is renewed on what learners retained and can do — the question the smile sheet was never able to answer.

See outcome analytics
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 1989Alliger
  • 1993Kraiger
  • 1994Kirkpatrick
  • 1997Alliger
  • 2003Arthur
  • 2005Brown
  • 2008Sitzmann
  • 2010Blume
  • 2012Salas
© 2026 FUTURE PROOF™
The evidence base. The 9 sources cited here span 1989–2012, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Kirkpatrick, D.L. (1994). Evaluating Training Programs: The Four Levels. Berrett-Koehler. PDF
  2. Alliger, G.M., & Janak, E.A. (1989). Kirkpatrick’s levels of training criteria: Thirty years later. Personnel Psychology 42(2): 331–342. PDF
  3. Alliger, G.M., Tannenbaum, S.I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology 50(2): 341–358. PDF
  4. Sitzmann, T., Brown, K.G., Casper, W.J., Ely, K., & Zimmerman, R.D. (2008). A review and meta-analysis of the nomological network of trainee reactions. Journal of Applied Psychology 93(2): 280–295. PDF
  5. Kraiger, K., Ford, J.K., & Salas, E. (1993). Application of cognitive, skill-based, and affective theories of learning outcomes to new methods of training evaluation. Journal of Applied Psychology 78(2): 311–328. PDF
  6. Salas, E., Tannenbaum, S.I., Kraiger, K., & Smith-Jentsch, K.A. (2012). The science of training and development in organizations: What matters in practice. Psychological Science in the Public Interest 13(2): 74–101. PDF
  7. Blume, B.D., Ford, J.K., Baldwin, T.T., & Huang, J.L. (2010). Transfer of training: A meta-analytic review. Journal of Management 36(4): 1065–1105. PDF
  8. Brown, K.G. (2005). An examination of the structure and nomological network of trainee reactions: A closer look at “smile sheets.” Journal of Applied Psychology 90(5): 991–1001. PDF
  9. Arthur, W., Bennett, W., Edens, P.S., & Bell, S.T. (2003). Effectiveness of training in organizations: A meta-analysis of design and evaluation features. Journal of Applied Psychology 88(2): 234–245. DOI
Try the AI engine

Renew programs on evidence, not applause.

Book a 20-minute demo. We’ll show you retention curves and skill movement for real content — the metrics a five-star rating cannot fake.

9 citations Reviewed August 2026 Open peer review welcomed