中文EN
← Deep Research
Deep Research · Deep dive

The Evidence Hierarchy of Learning Science: What Actually Works? (Deep Dive)

This is the deep-dive edition · read the plain-language edition →
TL;DR
The common study methods, re-seated by evidence strength: retrieval practice and spacing carry the thickest dossiers (classroom-valid, five bias checks clean); interleaving is strong within narrow borders (vocabulary reverses); active learning is directionally solid with soft numbers (0.47 SD atop 88% quasi-experiments); deliberate practice explains just 14% of variance post-corrigendum, and in preregistered replication the best violinists hadn't practiced more; learning-styles matching keeps failing qualified tests while ~9 in 10 educators believe it. Every effect size must be read with its control group and setting. Twelve testable claims close the essay.
72 votes · 24/24 surviveddeliberate practice: 14% of variancestyles: 89% belief vs 26% crossover12 testable claims

The empirical citations in this essay are graded: the 24 load-bearing claim groups (founding-experiment figures, meta-analytic effect sizes, verbatim statements from both sides of each academic battle) were each challenged by 3 independent verifiers (word-for-word checks against primary sources, searches for counter-evidence and corrigenda). All 24 survived refutation, with 30+ calibration amendments applied — including the recovery of a rarely cited official correction (the famous "12%" from Macnamara 2014 is the pre-corrigendum figure). Citations that did not enter verification are marked 【unverified, source】.

0. The popular ranking vs. the evidence ranking

On "how to learn," popular culture has a well-worn ranking: ten thousand hours makes a master; everyone has a learning style and should be taught to match; highlighting and rereading are the basics. Learning science has another ranking — some of its effects have been measured for over a century — and the two barely overlap. That is not a new discovery; the problem is that the measured conclusions and the popular ones travel on separate tracks.

This essay does one straightforward thing: it re-seats spaced repetition, retrieval practice, deliberate practice, active learning, and learning styles by strength of evidence. The rules up front: only checkable primary numbers count, every effect size is read together with its measurement conditions, counter-evidence stays in, and anyone selling something gets labeled.

1. How to read the numbers: calibrate the ruler first

Ranking requires effect sizes (d or g — the group difference divided by the standard deviation), but effect sizes in this field carry three traps. Unstated, they will corrupt every number that follows.

Trap one: the control group decides the number. The same "practice testing works" differs sixfold depending on the opponent: in the 2021 meta-analysis by Yang and colleagues covering 222 classroom studies and 48,478 students, class quizzing scores g=0.610 against "no activity/filler," shrinks to g=0.330 against "restudying," and drops to g=0.095 — non-significant (p=.062) — against other elaborative strategies such as concept mapping. 【verified】 Marketing always quotes the first number; the strict comparisons are the latter two.

Trap two: laboratory rulers don't measure classrooms. Matthew Kraft (2020) compiled 1,942 effect sizes from 747 education RCTs (standardized-achievement outcomes): the median is 0.10 SD. His field benchmarks: below 0.05 is small, 0.05 to 0.20 medium, 0.20 and above large — and in his words, "effects that are small by Cohen's standards are large relative to the impacts of most field-based interventions." By 5th grade (and beyond), a student's total gain over an entire academic year is about 0.40 SD or less, of which schooling contributes roughly 40%. 【verified】 Promising classroom gains from a laboratory g=0.5 is this field's most common conversion fraud.

Trap three: the league table itself can be broken. Education's most famous effect-size league table — John Hattie's Visible Learning (800+ meta-analyses pooled) — was called out by name by statistician Pierre-Jérôme Bergeron: its CLE index (a probability, in essence) was miscalculated in ways "flagrant to the point of giving negative probabilities or probabilities superior to 100%" (first noticed by Topphol in 2012), on top of averaging incommensurable effect sizes; "We must therefore absolutely qualify Hattie's methodology as pseudoscience." Hattie later conceded the CLE computation error; the league-table method stayed. 【verified】 Beneath that sits a structural defect: education research barely replicates — Makel and Plucker (2014) searched the complete publication history of the top 100 education journals by impact factor and found replications make up 0.13% of articles, with success rates significantly lower when the original authors are not involved. 【verified】

So every number below comes with its three-part tag: compared to what, measured where, reported by whom.

One “quizzing works”, three controls, three numbers vs no activity / fillerg=0.610vs restudying (strict)g=0.330vs elaborative strategiesg=0.095 (ns)
Schematic: the effect of class quizzing by control-group type (Yang et al. 2021; 222 classroom studies, 48,478 students); 0.095 vs elaborative strategies is non-significant (p=.062) — ads quote the first number, strict comparisons are the other two

2. Retrieval practice: the best-evidenced throne, and its borders

The founding experiment's numbers deserve a close look. Roediger and Karpicke (2006; Washington University, N=120): after reading a prose passage, one group restudied it, the other took a recall test. A week later the tested group retained 56% versus 42% for restudying (d=0.83). Experiment 2 is starker: a repeated-testing group that read the passage only about 3.4 times remembered 61% a week later; a repeated-reading group that read it 14.2 times remembered 40% (d=1.26). Two details that popularizations tend to cut: first, at a 5-minute delay the direction reverses — restudying 81% beats testing 75% — the testing dividend only appears after a delay; second, the group that read the most was the most confident about its future memory and performed the worst — the felt sense of effort and the reality of learning run in opposite directions. 【verified】

The meta-analytic lineage: shrinking from lab to classroom, but not vanishing. Laboratory ledger (Rowland 2014, 159 effects): g=0.50 against restudying, with high heterogeneity; with feedback 0.73 vs. 0.39 without; retention interval ≥1 day 0.69 vs. 0.41 same-day — feedback and delay are the two switches. 【verified】 Classroom ledger (Yang 2021, above): overall g=0.499, shrinking to 0.330 against the strict restudy control; effective at every school level; publication bias probed with five methods (PET-PEESE, trim-and-fill, 3PSM, and more), none revealing noteworthy bias — trim-and-fill estimated zero missing studies, a rare clean record in psychology meta-analysis. 【verified】 The worry that frequent quizzing breeds anxiety has also been measured: a 2023 meta-analysis (24 studies, 3,374 participants) found practice tests reduce test anxiety to a medium extent (g=−0.52, BF10>25,000), with the authors cautioning the field is young. 【verified】

The soft spot is transfer and boundaries. Pan and Rickard's 2018 transfer meta-analysis (192 transfer effects, N=10,382): does the benefit travel to new contexts? Overall d=0.40 against a restudy control, but after bias correction the intercept estimates drop substantially — often indicating no positive transfer when none of the favorable moderators (response congruency, elaborated retrieval practice, high initial test performance) are present. There is a real step between "recites the questions" and "can use the knowledge." 【verified】 Material complexity is another unfinished fight: van Gog and Sweller (2015) argued the testing effect vanishes as element interactivity rises; Karpicke and Aue countered that the concept resists quantification. Yang's 2021 classroom data show g=0.453 for problem-solving outcomes, leaning toward the latter, but the dispute is open 【unverified, source: Educ Psychol Rev 27 special issue】. And 2026 added a new border: two Prolific crowdsourced experiments failed to replicate the delayed testing effect, suggesting low-engagement online self-study may be a failure zone 【unverified, source: Frontiers in Psychology, PMC12894256】.

3. Spacing and interleaving: a century-old effect's modern ledger

Spacing has the thickest laboratory dossier in the field. Ebbinghaus found in 1885 that 38 repetitions spread over 3 days roughly equaled 68 repetitions massed in a single day 【unverified, source: 1913 translation, ch. 8】. The 2006 meta-analysis by Cepeda and colleagues gathered 184 articles and 839 assessments; across 271 direct massed-vs-spaced comparisons, spacing won broadly, with an average 15% advantage in studies with retention beyond one month — while the authors themselves noted most data points had retention under one day, a systematic gap between lab evidence and the month/year scales education needs. Their 2008 study filled it with 1,354 people and delays up to 350 days: the optimal gap as a share of test delay falls as delay grows — roughly 20–40% at a one-week delay, about 5–10% at one year; studying at the optimal gap beat zero gap by up to 111% in recall. 【verified】

"Intervals must expand" is a creed without evidence. Expanding review intervals are the default in Anki/SuperMemo culture. The 2021 meta-analysis by Latimier, Peyre, and Ramus tested it directly: expanding vs. uniform schedules, g=0.034, non-significant (16 studies, 54 effects) — the evidence does not support expanding schedules being generally superior, with only a non-significant trend favoring expansion when items are tested many times. In the same paper, the spaced-vs-massed retrieval subset (11 studies) showed uncorrected g=1.01 with a significant Egger test signaling publication bias, corrected to g=0.74 — direction solid, headline number discounted by about a quarter. 【verified】

The classroom ledger is less pretty. In 2024, Bego and colleagues ran spaced retrieval practice across 9 introductory university STEM courses: an average gain of 2.06 percentage points, dropping to 1.50 points and non-significance once calculus was excluded — to date the most systematic cold shower for "carry the lab effect into a regular course at light dosage." 【verified】 A 2025 meta-analysis restricted to authentic course material and timescales found d=0.54, but 3,000+ screened records yielded only 22 eligible reports, with 92% heterogeneity and visible small-study bias — read it as an upper bound 【unverified, source: Behavioral Sciences 15(6):771】. In mathematics specifically, spacing shrinks to g=0.28 and retrieval practice to g=0.18 with a CI crossing zero 【unverified, source: Educ Psychol Rev 2025, author-site PDF】 — echoing Donovan and Radosevich's 1999 warning that the spacing advantage shrinks as task complexity rises 【unverified, source: J Applied Psych 84(5)】.

Interleaving is the rare "survives the classroom intact" case — with narrow borders. The Rohrer team's trajectory: a 2015 classroom experiment with 126 seventh-graders showed the interleaving advantage growing with delay (d=0.42 after 1 day; 74% vs. 42%, d=0.79 after 30 days); in 2020 they upgraded to a preregistered cluster-randomized controlled trial — 54 classes, 787 students, four months of intervention — and a surprise test one month later scored 61% for interleaved versus 38% for blocked, d=0.83. A large effect, preregistered, in real classrooms: education research rarely gets all three at once. The cost is also in the paper: every teacher reported the interleaved assignments took longer, so the per-unit-time advantage shrinks. 【verified】 Brunmair and Richter's 2019 meta-analysis drew the borders: overall g=0.42, strongest for paintings (0.67), moderate for math (0.34), non-significant for expository text, and reversed for vocabulary — blocking wins (g=−0.39). Interleaving works by forcing discrimination between confusable categories; without confusability, only its cost remains. 【verified】

The lab's ruler vs the classroom's ledger retrieval, lab (Rowland 2014)g=0.50retrieval, classroom (Yang 2021)g=0.33interleaving RCT (Rohrer 2020)g≈0.83: the rare survivor Spacing added lightly to 9 regular university courses: avg +2.06 percentage points (Bego 2024)
Schematic: on the same vs-restudy calibration, retrieval practice loses about a third from lab to classroom; interleaving is the rare exception (7th-grade math, 54-class preregistered cluster RCT); light-dose spacing in regular courses shrinks to ~2 percentage points

4. The deliberate-practice war: a complete record of a strong claim being shot through by numbers

This is learning science's most famous battle, worth reading as a timeline, because it shows how a field self-corrects — and how slowly.

1993: the strong claim. Ericsson, Krampe, and Tesch-Römer studied 30 violin students at the Berlin conservatory and coined "deliberate practice." The original wording is unambiguous: "individual differences in ultimate performance can largely be accounted for by differential amounts of past and current levels of practice," and "we reject any important role for innate ability" (excepting height and body size). The data cornerstone: the best group had accumulated about 7,410 hours of solo practice by age 18 versus 5,301 for the good group, reported as significant — retrospective self-estimates, N=30. 【verified】

2008: the myth takes over. Malcolm Gladwell packaged it in Outliers as "ten thousand hours is the magic number of greatness" 【unverified, source: Outliers, 2008】. Ericsson cut ties in Peak (2016): "there is nothing special or magical about ten thousand hours" — half the violinists hadn't accumulated ten thousand hours by age 20, and Gladwell had conflated deliberate practice with anything labeled practice 【unverified, source: Peak excerpt, Salon 2016-04】. Myths don't need the original author's consent to spread.

2014: the meta-analytic counterattack. Macnamara, Hambrick, and Oswald pooled 88 studies (157 effect sizes, N=11,135): deliberate practice explained 12% of performance variance overall — the headline of the year. Adversarial verification here surfaced a fact almost nobody cites: the paper's official 2018 corrigendum revises it to 14% (r=.38), with domains at games 24%, music 23%, sports 20%, education 5%, professions 1% (still non-significant). The direction stands: about 86% of the variance lies outside practice amounts. The "19%" often quoted by the Ericsson camp is a sensitivity illustration — assuming both measures have reliability .80 — a share of reliable variance the authors never actually corrected for. 【verified】 The 2016 sports-specific meta-analysis ran colder: 18% overall, 1% among elite performers 【unverified, source: Perspectives on Psychological Science 11(3)】.

2016–2019: the definition war. Ericsson attacked the meta-analysis's method: summing every hour of any practice type assumes all practice affects performance equally, which the evidence contradicts. 【verified】 Macnamara's side counter-charged definitional drift: Ericsson himself wrote in 1998 that deliberate practice could be designed "by a teacher or the performers themselves," yet now insisted on teacher design — "For a theory to be falsifiable, definitions must be used consistently" 【unverified, source: Macnamara et al. 2016 Reply, Purdue archive PDF】. In 2019 Ericsson and Harwell renamed the broad definition "structured practice" and reanalyzed the 14 effect sizes from the 2014 meta-analysis meeting their criteria (14 of the full set of 157 effect sizes survived the three-criteria filter): uncorrected r=0.54 (about 29% of variance), rising to 61% (of reliable variance) after attenuation correction assuming practice-report reliability of 0.6 and performance-measure reliability of 0.8 — note these are two steps of one analysis, standing on a double assumption. 【verified】

2019: the replication verdict. Macnamara and Maitra ran a preregistered replication of the 1993 violin study (N=39; double-blind: experimenters unaware of skill groups, participants unaware of the study's purpose and of the existence of multiple groups). Results: practice explained 26% of variance, half the original's 48%. More fatal: the best violinists had accumulated 8,224 hours of practice alone by 18 — not more than the good group's 9,844 (p=0.364); the good group's mean was actually higher. The original's cornerstone — "the best practiced most" — simply did not hold in replication. 【verified】

The talent side scored in parallel. Mosing and colleagues (2014), 10,500 Swedish twins: the amount of music practice is itself heritable (40–70%), and within identical-twin pairs the twin who practiced more was no better — controlling for genetics, more practice was no longer associated with better music skills (measured as rhythm/melody/pitch discrimination, not elite performance). 【verified】 A 2024 genotyped study (~3,800 people) supplied the closing shape: practice benefits musical achievement more strongly in people with higher polygenic scores for cognitive performance — talent and practice interact rather than compete 【unverified, source: Heliyon 2024, PMC11292230】.

Verdict: "practice matters greatly" survives; "practice amounts largely explain who becomes a master" is dead. Two honest footnotes: both camps have been re-coding the same datasets for a decade, so every single percentage carries a camp's calibration; and Ericsson died in 2020 — the debate ended in a multifactor model, not in a concession.

The deliberate-practice war: a 30-year timeline 1993 strong claim practice “largely accounts” (N=30) 2008 Outliers 10,000 hrs = “magic number” 2014 Macnamara meta: 12% (corrected to 14% in 2018) 2016 definition war each side disputes coding 2019 preregistered replication best 8,224h ≯ good 9,844h 2024 interaction model talent × practice amplify
Schematic: from the 1993 strong claim to the failed 2019 preregistered replication — “practice matters” survives, “practice explains it all” does not; both camps re-code the same datasets

5. Learning styles: the myth that lost every test and still runs the classroom

The verdict was written in 2008. Pashler, McDaniel, Rohrer, and Bjork, commissioned to review learning styles, first set the bar — qualifying evidence must show a crossover interaction (assess styles, randomize teaching method, common exam; method A best for style A and not for style B) — then delivered: "We found virtually no evidence for the interaction pattern," while the few methodologically adequate studies (e.g., Massa & Mayer 2006, nearly 20 visualizer–verbalizer measures) returned null or flatly contradictory results (the review concedes at most one arguable positive, with serious methodological problems). Their famous line: "The contrast between the enormous popularity of the learning-styles approach within education and the lack of credible evidence for its utility is, in our opinion, striking and disturbing" — alongside naming "a thriving industry" publishing style tests and teacher guidebooks. 【verified】 Earlier, the UK's commissioned Coffield report (2004) examined 13 major models out of 71 identified; exactly one passed all four minimal psychometric criteria 【unverified, source: LSRC report】. Direct tests since (Rogowsky et al. 2015 with adults, 2020 with fifth-graders) found no style-by-modality interaction 【unverified, source: JEP 107(1); Frontiers in Psychology 11:164】.

Evidence never beat belief. Newton and Salvi's 2020 review pooled 33 studies (37 samples, 18 countries, 15,405 educators): a weighted 89.1% believe matching instruction to learning styles works; pre-service teachers at 95.4% versus in-service at 87.8%, a non-significant difference — the belief is installed before the career starts. 【verified】 Nancekivell, Shah, and Gelman (2020), US survey: 93.7% of respondents believed in learning styles (full pre-exclusion sample of 383), educators (90.7%) actually slightly below non-educators (95.9%); a substantial cluster of believers holds an essentialist reading — styles as innate, unchanging, wired into the brain. 【verified】 A 2025 survey of 1,257 primary teachers across 11 countries found over 90% endorsement, with teachers reporting the misconception came mainly from training programs and professional workshops — the vector is the training system itself 【unverified, source: Trends in Neuroscience and Education, PMID 40889831】.

The honest 2024 amendment doesn't amend the conclusion. Clinton-Lisell and Litzinger meta-analyzed 21 matching experiments: g=0.31 (CI lower bound 0.05, marginal) — the most favorable quantitative result learning styles has ever received. But only 26% of outcomes showed the theoretically required crossover interaction, only 5 of 21 studies met What Works Clearinghouse standards, and the authors' own conclusion reads: benefits "too small and too infrequent to warrant widespread adoption." 【verified】 The state of the defense is itself informative: VARK designer Neil Fleming conceded in 2012 that "we do not have any reliable, valid research that would predict that knowing one's learning style is beneficial for learning" 【unverified, source: vark-learn.com self-published; direct commercial interest】; engineering education's Richard Felder (whose own questionnaire is "accessed by millions of users") defended in 2020 by retreat — the critics attack a meshing straw man; styles should merely inform teaching that "balances" preferences 【unverified, source: Advances in Engineering Education 8(1); interested party】 — and that weak claim, teach diversely, no critic ever opposed.

Learning styles: the belief-evidence scissors educators who believe matching works89.1% (15,405 educators, 18 countries)outcomes showing the crossover26% (friendliest 2024 meta)qualifying evidence, 2008 review“virtually no evidence”
Schematic: educator belief across 18 countries (Newton & Salvi 2020) vs the rate of theoretically required crossover interactions in the friendliest meta-analysis (Clinton-Lisell 2024; only 5/21 studies met quality standards) vs the Pashler et al. 2008 verdict

6. Active learning vs. lecture: direction solid, numbers soft

The headline numbers are loud. Freeman et al.'s 2014 PNAS meta-analysis (225 undergraduate STEM studies): active learning raises exam performance by 0.47 SD on average; failure rates run 33.8% under traditional lecture versus 21.8% with active learning (odds ratio 1.95); the effect holds across STEM disciplines and is largest in small classes. 【verified】 Deslauriers et al. (2019) added the mechanism experiment at Harvard (N=149, randomized, one 90-minute class per method in a crossover design, identical handouts): the active group learned more by 0.46 SD on tests yet felt they learned less by 0.56 SD — students punish effective, effortful teaching in evaluations. This is the same phenomenon as Section 2's "read it 14 times, most confident, remembered least": the feeling of effort gets booked by the brain as "not learning," when effort is precisely the signal that learning is happening. 【verified】 Active learning is also reported to narrow achievement gaps for underrepresented groups (exam gap −33%, passing gap −45%, only under high-intensity implementation) 【unverified, source: Theobald et al. 2020, PNAS】.

But the evidence base has been audited by name. Martella et al. (2023) re-examined the 176 codeable articles behind the Freeman meta-analysis: 88.1% quasi-experimental, only 10.2% randomized; 92.6% reported no implementation fidelity for the experimental group; counting fidelity, zero articles met all 12 internal-validity controls; a random sample of 84 articles from 2015–2022 showed no improvement (8.3% randomized). 【verified】 Note what the audit disputes: causal strength, not direction — no evidence says lecture is better — but 0.47 SD stands on a pile of quasi-experiments and should be read discounted. "Active learning" as a term was itself judged by an authoritative review to be an umbrella "not particularly useful": courses coded "active" can contain up to ~90% lecture, and "lecture" courses can contain activities 【unverified, source: Lombardi et al. 2021, PSPI; Martella 2022 blog calibration】.

The war's real map is not binary. The cognitive-load school (Kirschner, Sweller & Clark 2006) declared half a century of evidence shows "minimally guided instruction" doesn't work 【unverified, source: Educational Psychologist 41(2)】; Hmelo-Silver et al. countered that PBL and inquiry learning are heavily scaffolded and were wrongly filed under minimal guidance 【unverified, source: Educational Psychologist 42(2)】. Mayer's 2004 "three strikes" review supplies the reconciliation: across three decades of literature, guided discovery consistently beats pure discovery — what works is cognitive engagement plus instructional guidance, not behavioral busyness 【unverified, source: American Psychologist 59(1)】. The Direct Instruction camp's meta-analysis (Stockard et al. 2018, 328 studies, overall 0.54) has handsome numbers, but the first author was research director of NIFDI, the DI advocacy institute, and the literature is criticized for comparing "something with nothing" 【unverified, source: Rev Educ Res 88(4), interested party; Eppley & Dudley-Marling 2019】. Each side burned the other's straw man; the one enemy they share is unguided pure discovery.

Same students, two rulers 0 measured learning (test) +0.46 SD feeling of learning −0.56 SD
Schematic: Deslauriers et al. 2019 (intro physics at Harvard, N=149, randomized, identical handouts) — the active group measurably learned 0.46 SD more yet felt it learned 0.56 SD less; the same metacognitive illusion as “read it 14 times, most confident, remembered least”

7. Why hard works: two theories that converge and constrain each other

Six sections keep replaying one pattern: feelings and outcomes point in opposite directions (rereading feels confident and underdelivers; blocking feels smooth and fades; lectures feel great and teach less). Robert Bjork named it in 1994 — desirable difficulties: "Manipulations that speed the rate of acquisition during training can fail to support long-term posttraining performance, while other manipulations that appear to introduce difficulties for the learner during training can enhance posttraining performance"; spacing, variability, reduced feedback, and tests-as-practice share the property of introducing difficulty 【unverified, source: Bjork 1994 chapter, gwern-archived PDF】. Soderstrom and Bjork (2015) systematized it as learning ≠ performance: what you can observe during training is performance, an often unreliable index of whether learning occurred 【unverified, source: Perspectives on Psychological Science 10(2)】.

But "difficulty helps" is not a master key; cognitive load theory installs the limiter. For novices, studying worked examples beats solving problems — in Sweller and Cooper's 1985 experiments, the worked-example group solved similar problems in about half the time with about one-fifth the errors (though gains didn't extend to varied problems) 【unverified, source: Cognition and Instruction 2(1), figures via secondary transcription】; and the expertise reversal effect (Kalyuga et al. 2003) shows guidance nearly essential for novices can become ineffective or harmful for the experienced 【unverified, source: Educational Psychologist 38(1)】. Together they form the complete operating rule: difficulty must land within the learner's reach — novices need guidance and examples first; only the experienced earn the "desirable difficulties" of generation, retrieval, and interleaving. Which also resolves Section 6: active learning was never meant to be the same medicine for students with different prior knowledge.

8. The re-seated ranking: an evidence-tier table

Putting the 24 verified claim groups against four criteria (effect robustness, classroom evidence, independent replication, bias checks):

  1. Tier 1 (thick evidence, holds in classrooms, clean bias checks): retrieval practice and spaced practice. Conditions locked in: retrieval wants feedback and delayed outcomes, steadiest on factual/vocabulary material; spacing at light dosage in a regular course may deliver only ~2 percentage points. Most reliable ≠ unconditional.
  2. Tier 2 (direction solid, size or borders in question): interleaving (strong in math/category learning, reversed for vocabulary), active learning (read as "guided cognitive engagement"; discount the 0.47 SD), worked examples (novices only; reverses with expertise).
  3. Tier 3 (sound core, mythical packaging): deliberate practice. Structured practice is necessary and high-return, but the post-corrigendum ledger reads: practice amounts explain ~14% of performance variance, non-significant in professions; "ten thousand hours" is a bestseller's invention, and in preregistered replication the best group hadn't practiced more at all.
  4. Tier 4 (failed qualified testing repeatedly): learning-styles matching and unguided pure discovery. The friendliest meta-analysis styles ever got shows a marginal small effect, a 26% interaction rate, and low-quality studies; pure discovery is the one shared enemy of both opposing schools.
  5. Footnote: rereading and highlighting are "low utility," not "useless" — Dunlosky et al.'s 2013 calibration is that their benefits are conditional and non-general (and they are what students report using most). That rating has never received a formal ten-year update; the evidence for the three moderate-utility techniques (elaborative interrogation, self-explanation, interleaving) has thickened considerably since 2013. 【verified】
The re-seated ranking: evidence tiers Tier 1: retrieval practice · spaced practiceconverging metas · classroom-valid · clean bias checks; conditions: feedback + delayTier 2: interleaving · active learning · worked examplesdirection solid, size or borders in question (vocab reverses · discount 0.47 SD · reverses with expertise)Tier 3: deliberate practicestructured practice is necessary, but “10,000 hours” is packaging — 14% of variance post-corrigendumTier 4: styles matching · pure discoveryfails qualified tests repeatedly; the latter is the one shared enemy of both opposing schools
Schematic: seating by four criteria — effect robustness × classroom evidence × independent replication × bias checks; rereading/highlighting are “low utility”, not useless

9. Twelve testable claims

Ordered by evidence strength:

  1. Retrieval practice and spacing are the field's two best-evidenced techniques: converging meta-analyses, classroom validity, five-method publication-bias checks coming back clean. (Strong: Rowland/Yang/Cepeda/Latimier all verified)
  2. The control group decides the advertised number: the same technique scores 0.610 against nothing, 0.330 against rereading, 0.095 and non-significant against elaborative strategies — always ask "compared to what." (Strong: Yang 2021 verified)
  3. Testing's dividend requires delay (same-day, rereading wins), and confidence runs opposite to outcomes — performance during learning is not learning. (Strong: R&K 2006 and Rowland moderators verified)
  4. "Review intervals must expand" has no empirical support: expanding vs. uniform g=0.034, non-significant; the "optimal gap" is only a coarse ratio that shifts with retention needs (~20–40% of delay at one week, ~5–10% at one year). (Strong: Latimier/Cepeda 2008 verified)
  5. Lab-to-classroom shrinkage is real and technique-specific: retrieval 0.50→0.33 against rereading; light-dose spacing in regular courses down to ~2 percentage points; interleaving the rare exception (preregistered classroom RCT, d=0.83). (Medium-strong: Bego/Rohrer verified; applied-meta upper bound unverified)
  6. Interleaving has hard borders: material must be confusable — vocabulary reverses (g=−0.39) — and interleaved work takes longer. (Strong: Brunmair & Richter verified)
  7. Deliberate practice's explanatory power is limited: 14% overall post-corrigendum (86% lies elsewhere), non-significant in professions; in preregistered replication the best violinists had not accumulated more practice (p=0.364). (Strong: 2018 corrigendum and 2019 replication verified)
  8. Practice amounts are themselves highly heritable (40–70%), and controlling for genetics, practice no longer tracks music ability — the correct form of "talent vs. practice" is an interaction model. (Medium-strong: Mosing verified; 2024 interaction study unverified)
  9. Learning-styles matching fails qualified tests repeatedly while 89–94% of educators believe it, with teacher training as the main vector — the field's most extreme popularity-to-evidence inversion. (Strong: Pashler/two prevalence sources/2024 amendment meta all verified)
  10. Active learning: direction solid, numbers soft — 0.47 SD stands on 88% quasi-experiments and zero fully controlled articles; students who measurably learned 0.46 SD more felt they learned 0.56 SD less, so evaluations punish effective teaching. (Strong: Freeman/Martella/Deslauriers all verified)
  11. The two warring schools actually agree: what works is cognitive engagement plus guidance; novices need worked examples and structure, and the same method reverses with expertise. (Medium: theoretical synthesis; individual citations unverified)
  12. Education research's meta-level defects are systemic: 0.13% replication rate, median field effect 0.10 SD, and the most famous effect-size league table contains arithmetic-grade errors — audit the calibration of any ranking before citing it. (Strong: Makel & Plucker/Kraft/Bergeron all verified)

Worth watching: whether the Dunlosky ratings get a formal update (interleaving and self-explanation would likely rise); a second preregistered classroom RCT for interleaving outside middle-school math; whether any WWC-grade positive evidence for style-matching emerges after Clinton-Lisell; and whether AI learning tools ship the three verified switches — retrieval, spacing, feedback — as defaults, or keep selling "personalized learning styles."


Appendix: primary sources

Measurement & meta-level: Kraft 2020, Educational Researcher 49(4) · Makel & Plucker 2014, Educational Researcher 43(6) · Bergeron & Rivard 2017, McGill Journal of Education 52(1) · Topphol 2012 (Norwegian original unchecked; via Bergeron)

Retrieval practice: Roediger & Karpicke 2006, Psychological Science 17(3) · Rowland 2014, Psychological Bulletin 140(6) · Yang, Luo, Vadillo, Yu & Shanks 2021, Psychological Bulletin 147(4) · Adesope et al. 2017, Rev Educ Res 87(3) (secondary citation) · Agarwal, Nunes & Blunt 2021, Educ Psychol Rev 33 (first author runs retrievalpractice.org; interested party) · Pan & Rickard 2018, Psychological Bulletin 144(7) · van Gog & Sweller 2015 and Karpicke & Aue 2015, Educ Psychol Rev 27(2) · Yang et al. 2023, Educ Psychol Rev 35:87 · King-Shepard et al. 2025, Educ Psychol Rev · Sigayret et al. 2026, Frontiers in Psychology

Spacing & interleaving: Ebbinghaus 1885/1913 · Murre & Dros 2015, PLOS ONE · Cepeda et al. 2006, Psychological Bulletin 132(3) · Cepeda et al. 2008, Psychological Science 19(11) · Latimier, Peyre & Ramus 2021, Educ Psychol Rev 33 · Carpenter et al. 2012, Educ Psychol Rev 24 · Donovan & Radosevich 1999, J Applied Psychology 84(5) · Bego et al. 2024, Int J STEM Educ 11:9 · Mawson & Kang 2025, Behavioral Sciences 15(6) · Murray et al. 2025, Educ Psychol Rev · Rohrer & Taylor 2007, Instructional Science 35 · Rohrer, Dedrick & Stershic 2015, JEP 107(3) · Rohrer et al. 2020, JEP 112(1) · Brunmair & Richter 2019, Psychological Bulletin 145(11) · Settles & Meeder 2016, ACL (Duolingo; interested party)

The deliberate-practice war: Ericsson, Krampe & Tesch-Römer 1993, Psychological Review 100(3) · Gladwell 2008, Outliers · Ericsson & Pool 2016, Peak · Macnamara, Hambrick & Oswald 2014, Psychological Science 25, and 2018 Corrigendum, Psychological Science 29(7) · Macnamara, Moreau & Hambrick 2016, Perspectives on Psychological Science · Ericsson 2016, Perspectives on Psychological Science · Ericsson & Harwell 2019, Frontiers in Psychology 10:2396 · Macnamara & Maitra 2019, Royal Society Open Science 6:190327 · Mosing et al. 2014, Psychological Science 25 · Wesseldijk et al. 2024, Heliyon

Learning styles: Pashler, McDaniel, Rohrer & Bjork 2008, PSPI 9(3) · Massa & Mayer 2006 (via the Pashler review) · Coffield et al. 2004, LSRC report · Rogowsky et al. 2015, JEP 107(1); 2020, Frontiers in Psychology 11 · Willingham, Hughes & Dobolyi 2015, Teaching of Psychology 42(3) · Newton 2015, Frontiers in Psychology 6 · Newton & Salvi 2020, Frontiers in Education 5 · Nancekivell, Shah & Gelman 2020, JEP 112(2) · Fleming 2012, vark-learn.com (interested party) · Felder 2020, Advances in Engineering Education 8(1) (interested party) · Clinton-Lisell & Litzinger 2024, Frontiers in Psychology 15 · Bresnahan et al. 2024, Frontiers in Psychology 15 · Adiguzel et al. 2025, Trends in Neuroscience and Education

Active learning & guidance: Freeman et al. 2014, PNAS 111(23) · Deslauriers et al. 2019, PNAS 116(39) · Theobald et al. 2020, PNAS 117(12) · Martella et al. 2023, Educ Psychol Rev 35:104 · Lombardi et al. 2021, PSPI 22(1) · Kirschner, Sweller & Clark 2006, Educational Psychologist 41(2) · Hmelo-Silver, Duncan & Chinn 2007, Educational Psychologist 42(2) · Mayer 2004, American Psychologist 59(1) · Stockard et al. 2018, Rev Educ Res 88(4) (first author formerly NIFDI research director; interested party) · Eppley & Dudley-Marling 2019, J Curriculum and Pedagogy 16(1)

Frameworks: Dunlosky, Rawson, Marsh, Nathan & Willingham 2013, PSPI 14(1) · Bjork 1994, in Metcalfe & Shimamura (Eds.), Metacognition · Soderstrom & Bjork 2015, Perspectives on Psychological Science 10(2) · Sweller & Cooper 1985, Cognition and Instruction 2(1) · Kalyuga, Ayres, Chandler & Sweller 2003, Educational Psychologist 38(1) · Donoghue & Hattie 2021, Frontiers in Education 6 (interested party) · Pan, Dunlosky, Xu & Ouwehand 2024, Educ Psychol Rev 36(1)

Research materials and all verification rulings are archived in the research base (6 investigation lines, 106 claims, 24 load-bearing claim groups × 3 votes).