Methodology note: the load-bearing claims in this essay went through 24 groups × 3 adversarial verification votes (72 votes: 1 group survived intact, 23 groups carried caliber corrections, 0 groups overturned), plus contradiction-search and methods-audit seats for 5 single-source empirics (10 verdicts: 1 upgraded to multi-source, 9 restricted-use with conditions, and 2 numbers vetoed from load-bearing). Verification produced 60+ caliber corrections — and the distinctive feature of this edition is that the corrections landed on both sides of the debate in roughly equal measure. Evidence grades: 【multi-source】≥2 independent sources concur; 【single-source, verified】traceable to a primary document; 【direction contested】independent sources conflict; 【interested-party account】self-description by a stakeholder — direction informative, magnitude not load-bearing. All figures as of July 2026. This essay is an evidence review, not medical or parenting advice.
Jonathan Haidt's The Anxious Generation (2024) is one of the most influential social-science bestsellers of the decade: it directly propelled the wave of US state school-phone legislation, Australia's under-16 social media ban, and the UK equivalent announced in June 2026 and planned for 2027 【official announcement; state-law counts vary by caliber, so no figure is quoted here】. The critics are equally heavyweight: Candice Odgers ruled in Nature that the book's thesis is "not supported by science," and Andrew Przybylski and Amy Orben computed, from datasets with a nominal 355,000 adolescents, associations in the same league as eating potatoes.
Seven years on, the fight has not converged. The first reason is that it is actually three independent questions, and the two camps routinely declare victory on different ones:
The second reason is the construct ladder: beneath the "same topic," the exposure variable spans at least five rungs (screen time → digital technology use → social media → a single platform → smartphone access), and the outcome variable another five (school loneliness → self-reported symptoms → clinical diagnosis/visits → self-harm hospitalization → suicide). Most of the 60+ corrections in this audit come from someone — frequently both camps at once — treating one rung as another. Before trusting any number below, ask: which rung of exposure, which rung of outcome.
The circulating version is "Haidt says social media caused the teen mental health crisis." The original is more careful — and the careful parts are precisely what supporters and critics alike tend to delete.
The book's thesis sentence (end of Chapter 1, US hardcover p. 44) is two sentences that must not be split 【single-source, verified via three independent verbatim reprints】:
"Between 2010 and 2015, the social lives of American teens moved largely onto smartphones with continuous access to social media, online video games, and other internet-based activities. This Great Rewiring of Childhood, I argue, is the single largest reason for the tidal wave of adolescent mental illness that began in the early 2010s."
Three routinely deleted qualifiers: the subject is the composite construct of "the Great Rewiring" (social media + online games + other internet activities), not social media alone; "I argue" marks an author's argued position; and "single largest reason" is not "the only reason." The introduction's self-declared central claim (p. 9) is weaker still: overprotection in the real world and underprotection in the virtual world — two trends — are "the major reasons" (plural). Treating p. 44 as the book's central claim systematically strengthens Haidt.
His written testimony to a Senate Judiciary subcommittee in May 2022 (its own title says "a Major Contributing Cause") contains the causal sentence 【multi-source, official PDF】:
"Correlational, experimental, and eye-witness testimony points to social media as a major cause of the crisis. I do not believe that social media is the only cause of the crisis, but there is no alternative hypothesis that can explain the suddenness, enormity, and international similarity that I laid out in part 1 of this document."
Note the structure: this is a burden-shifting argument ("if not social media, then what?"), not a statistical estimate — the next line demands platform spokespeople answer "OK, then what do YOU think caused this?". The testimony's evidence base is the rolling Google Docs collaborative reviews Haidt maintains with Jean Twenge — public and auditable, but unversioned and not peer-reviewed 【interested-party account】.
The action agenda is the TAG movement site's four norms (retrieved 2026-07): "No smartphones before high school / No social media before 16 / Phone-free schools, from bell to bell / More independence, free play, and responsibility in the real world."
On numbers, the testimony's headline is that from 2009 to 2019 the increases "are generally between 50% and 150%." The audit finds this range source-dependent 【multi-source】: on NSDUH (structured interviews approximating diagnosis), 12-17 major depressive episodes went 8.1%→15.8% (+95%), girls +105% — inside the range; on CDC's YRBS (anonymous self-report), persistent sadness +40.6%, considered suicide +36.2%, suicide plan +44.0%, attempt +41.3%, attempt requiring medical care +31.6% — all five below the 50% floor. "50-150%" can be quoted as Haidt's sentence, not used as a factual range; its underlying computation points to an unversioned Google Doc and cannot be reproduced 【untraced】.
One self-correction deserves the record: the book's official errata page lists six factual errors, and in its Postscript Haidt concedes the p. 128 "Brain Drain" claim (mere phone presence impairs cognition) — citing a 2024 meta-analysis (d=−0.02 [−0.06, 0.01], N=4,368), he writes "my initial claim... may be incorrect" 【interested-party account, primary errata page】.
Three mutually independent US anonymous surveys, immune to help-seeking behavior, point the same way 【multi-source】:
MTF's constant-wording reversal in 2012 is the strongest single rebuttal to the "pure measurement artifact" reading: anonymous questionnaires are untouched by screening guidelines, help-seeking, or coding rules. Norway's Ungdata (N=560,712) ran the formal test: measurement invariance across time essentially holds, i.e., the trend stands without correction, and the reporting-drift explanation "was not supported" — girls' latent means rise from 2014 【multi-source】.
Suicide deaths are immune to reporting inclination. The official NCHS joinpoint series 【multi-source, official】:
"The suicide rate for people aged 10–14 declined from 2001 through 2007 (from 1.3 deaths per 100,000 to 0.9), tripled from 2007 through 2018 (from 0.9 to 2.9), and then did not change significantly through 2021."
Critics like to note the takeoff is 2007, before smartphones — true, but split by age: in the 10-14 annual series 2012 stands at 1.5, so about 1.9× of the tripling happened after 2012; for ages 20-24, NCHS itself reports 2012-2021 rising 4% per year versus 1% over 2001-2012 — that group's acceleration kink sits exactly at 2012; the 15-19 rise runs 2009-2017 (+57%). Both "2007 refutes the 2012 story" and "the hockey stick starts precisely in 2012" over-read the joinpoints.
Recent direction: 2021→2022 shows the only officially significant youth decline (ages 10-14: 598→493 deaths, −17.6%), then the series flattens (10-14 rates 2022-2024: 2.4/2.3/2.3). The circulating "all youth groups fell in 2023→2024, driven by males, females unchanged" is overturned by final data: the combined changes are statistically indistinguishable from zero; by sex, the only significant fall is males 15-19 (14.6→13.3) while females 15-19 (4.7→5.4) are the only significant rise among the six cells (own computation, z≈2.2-2.5, not an official test) 【multi-source, final NVSS】.
Pin the confounder too: the circulating pair "firearm suicides +59% vs non-firearm +29%" exists in no source — verification exhausted window combinations without reproducing it; the verifiable figures are 2010-2019 firearm +40-42% vs non-firearm +28-29% (a ratio of ~1.4-1.5, not 2), and baseline-sensitive — firearm rates bottomed in 2006-07, so starting at 2010 means starting at the trough 【multi-source incl. independent NVSS recomputation; Everytown is a gun-control advocacy group, figures independently confirmed】.
The strongest blow against the crisis narrative comes from Corredor-Waldron & Currie (Journal of Human Resources 2024): in New Jersey's statewide 2008-2019 hospital data, of the +50% rise in suicide-related visits, 24.9 of 25.3 points came from "suicidal ideation" diagnoses, 18.5 of those as secondary diagnoses; the two jumps track screening expansion (USPSTF's 2009 grade-B recommendation, insurance mandate 2011) and the October 2016 coding reversal (ICD-10's Exclude 1 → Exclude 2 for R40-R46, first instructing clinicians to record ideation alongside a psychiatric principal diagnosis) — while self-harm, attempts, and completions were "essentially flat" 【single-source verified, R3 dual-seat audited】.
The audit drew the load-bearing boundary precisely:
A second, independent artifact line: the October 2015 ICD-9→10 switch produced an immediate jump in the 12-17 self-harm care series (nine-system interrupted time series: 38.5/100k — note the abstract carries no sign) 【single-source verified】; contemporaneous literature reads the net effect as "totals largely unchanged, intent categories swapped" (Stewart 2017: "Marked changes in coding of intent... almost certainly represent artifacts" — the qualifier of intent cannot be dropped). And the artifact story runs into one dataset it cannot touch: Mercado et al. (JAMA 2017) uses NEISS ED data where self-harm is identified by reading injury narratives, not billing codes — girls 10-14 rose 18.8% per year from 2009 (109.8→317.7/100k). That takeoff is 2009, before every coding event.
Bottom line: the visit explosion contains major artifacts (multi-source confirmed), and the underlying deterioration is also real (anonymous surveys + suicide rates + narrative-coded ED data). Both hold; neither swallows the other.
The deterioration in Anglosphere girls' self-report and US hard outcomes is real — the artifact reading is boxed in by constant-wording MTF, invariance-tested Ungdata, and coding-independent data. But "globally synchronized" outruns the evidence (German girls neither fell nor clearly worsened; East Asian data are thin; PISA measures a different construct), and so does "starting precisely in 2012" (10-14 suicide bottomed in 2007; the ideation coding jump is 2016). One more curve neither camp likes: US indicators improved or flattened over 2021-2024 while usage did not retreat (see 5.4) — a new explanatory burden for a dose-response mechanism, and equally unexplained by anyone calling the crisis an artifact.
Orben & Przybylski 2019 (Nature Human Behaviour) ran specification-curve analysis over three big datasets 【multi-source】:
"The association we find between digital technology use and adolescent well-being is negative but small, explaining at most 0.4% of the variation in well-being. Taking the broader context of the data into account suggests that these effects are too small to warrant policy change."
Calibers pinned by the audit: 0.4% is the largest of the three datasets' median-specification partial η² (the MCS analysis carrying it has a median analytic n≈7,968, not the nominal 355,358; subanalyses in the same paper reach 1.1%). "Same league as eating potatoes" (×0.86) is from YRBS while "worse than wearing glasses" (×1.45) is from MCS — different datasets; and in the circulating "sleep and breakfast are 1.7-44.2× more positive," ×44.2 is a near-zero-denominator artifact (MTF's all-technology β=−0.006; "listening to music" ×32.68 and "religious activity" ×16.29 blow up identically). Citable as table values; not usable as substantive argument.
The response (testimony §§2.2-2.6 + Twenge et al. 2020 in NHB) rests on four pillars; the audit wounded each in turn:
The within-person longitudinal anchor (Orben, Dienlin & Przybylski 2019, PNAS, RI-CLPM): "Both median longitudinal effects were trivial in size (social media predicting life satisfaction, β = −0.05; life satisfaction predicting social media use, β = −0.02)" — with three qualifiers: these are medians over 2,268 specifications; the median analytic n is 1,699 (not 12,672); and the between-person contrast ψ=−0.13 must be shown alongside — the paper's real finding is "small negative between persons, near-zero within" 【multi-source】.
The skeptics' own meta-analyses carry their own conditions: Ferguson's two correlational metas (2022: r+=.052, 37 effect sizes/33 studies, construct = screen media overall, the adolescent×social-media subgroup only k=6; 2025: β=.061, 79 effect sizes/46 studies). Note the "girls .075 / boys .044" gender moderation is not significant (p=.074) — "girls are more affected" may not be written from it; the two papers share all four authors and 15% of samples — not independent corroboration 【single-source verified + direction contested】.
Haidt's testimony Figure 5 (redrawn from Kelly et al. 2018, UK Millennium Cohort): girls at 0 hours of social media, 11.2% with clinically relevant depressive symptoms; at 5+ hours, 38.1% (3.4×); boys 7.4%→14.5% (~2×). The numbers match the source table. What collapses is three calibers:
The circulating "each extra hour raises depression risk 13%" also needs its source corrected: not Vidal but Liu et al. 2022 (OR=1.13 [1.09-1.17]), whose dose-response rests on only 5 studies — and whose pool includes Kelly's MCS. Quoting the Kelly chart and "13%/hour" as two pieces of evidence counts the same data twice.
At full-sample, wide-construct caliber the correlation is genuinely small (r≈.03-.06); narrowing to girls×social media genuinely enlarges it — but by how much, the camps still disagree by 2-3×. Both sides' headline rhetoric took corrections in this audit: O&P's potato framing carries a near-zero-denominator artifact and dataset splicing; Haidt's "2-6×," "r=.15-.22," and lead-exposure benchmark fail on the lower bound, the missing source, and the attribution respectively. One structural reason the war never ends: a genuine adversarial collaboration — jointly preregistered, each camp re-running the other's data — has never happened in this field.
Exploiting Facebook's staggered 2004-2006 college rollout (420 colleges, 359,827 NCHA responses, 4 expansion groups), Facebook's arrival raised the poor-mental-health index by 0.085 SD 【multi-source for the magnitude's position; design details single-source verified + R3 audited】. The audit drew its load-bearing boundary finely:
Multiplicity, stated plainly: under sharpened-FDR correction of 14 individual outcomes, nothing survives at q<0.05; five items survive at q<0.10, all in the affective cluster; suicidal ideation and eating-disorder items are all null.
Allcott et al. 2020 (AER): paid Facebook deactivation for four weeks; subjective well-being +0.09 SD (SE 0.04). Three pins: estimated on the ~1,637 completing endline (not the circulating 2,743 — that is how many received offers); the estimand is a LATE (compliers willing to deactivate for $102), not ITT; the high-frequency text-message measures are not significant — the significant result is the retrospective endline survey. Its most valuable sentence 【multi-source】:
"the magnitudes of our causal effects are far smaller than those we would have estimated using the correlational approach of much prior literature."
(The correlational approach would have implied ~0.23 SD; the experiment finds 0.09 — correlational literature overestimates by roughly 3×.) The circulating "25-40% of the effect of psychological interventions" should be retired: the denominator is a positive-psychology intervention meta whose own outcomes all show publication bias; corrected, 0.09's relative share goes up — the opposite of the "so it's small" intent 【direction contested】.
Allcott et al. 2025 (NBER WP 33697, not peer-reviewed): the largest deactivation experiment ever run (with Meta, before the 2020 election; 17,802 Facebook and 13,480 Instagram users completing endline). Six weeks off Facebook (five net) improved the emotional-state index 0.060 SD (q=0.002 after correction); Instagram 0.041 SD — p=0.016 alone, q=0.139 after correction: "does not meet our pre-registered p = 0.05 significance threshold." The widely circulated "the Instagram effect is driven by women under 25" is a non-preregistered exploratory analysis: that cell shows 0.111 SD (p=0.002), but the four age×gender cells' difference test gives p=0.062 — the abstract's "driven by" is stronger than the underlying test; quote it only with the numbers attached 【single-source verified】. A counterintuitive mechanism finding: about three-quarters of the freed time (Instagram: all of it) flowed to other phone apps, not offline life. The interest structure cuts both ways: several co-authors are Meta employees and Meta executed sampling and deactivation; but the checks are in writing ("Meta could not block any results from being published") and the results cut against Meta.
Castelo et al. 2025 (PNAS Nexus): blocking mobile internet on phones for two weeks; headline dz=0.45/0.56/0.23 (well-being/mental health/sustained attention). What survives citation is the objective-attention dz=0.23 and the significant interactions; 0.56 needs caution. The three headline dz are within-subject pre-post changes pooled across both arms (not randomized between-group contrasts); at the randomization level the three Condition main effects are all non-significant while Condition×Time interactions are significant; the "25.5% met the preregistered compliance criterion" claim fails against the primary preregistration file (the OSF .docx contains no "10 of 14 days" definition); and "larger than antidepressants" reverses under matched calibers (its comparator, Kirsch's d=0.32, is a drug-minus-placebo between-group difference; the matched within-drug-arm improvement is d≈1.24 — more than double 0.56) 【single-source verified; multiple vetoes】.
The reduction/abstinence meta melee, after caliber sorting, is less of a melee:
May∩Lemahieu share zero studies and no outcomes — "depression improves g≈0.2-0.3" and "momentary affect ≈0" can both be true.
Nagata et al. 2025 (JAMA Network Open; ABCD cohort N=11,876; RI-CLPM; exposure self-reported by the adolescent, outcome reported by the caregiver — a cross-informant design that rules out the crudest common-method bias): of three one-year-lag paths, the latter two are significant (β=0.07 and 0.09), and all three reverse paths (depression → more use) are null. The R3 contradiction seat upgraded the directionality to multi-source via a power analysis: detecting β≈0.07 takes N≈1,500-1,600, and the commonly cited nulls (Coyne N=500, Steinsbekk N=810, Jensen N=388) are all underpowered; among adequately powered studies, two of three are positive (Nagata; Boers N=3,826) 【multi-source (direction)】.
But the magnitude must be given naked: β=0.09 translates to "about 1.5 extra hours per day → caregiver-rated CBCL depression raw score higher by ~0.21 points" — under a fifth of one symptom item, under 1% of within-person variance. The paper's "medium" label cites percentile benchmarks of the cross-lagged literature (.07 = the median published effect), not practical magnitude. Four may-nots: not causal (RI-CLPM does not identify causation), not sizeable, not social-media-specific (in the same team's earlier ABCD paper, social media was the weakest of six screen types, behind video chat), and not evidence for the girls' sensitive window (the paper ran no gender moderation) 【single-source verified + audited】.
Put every causal design on one axis and the magnitudes converge strikingly: platform introduction 0.06-0.16 (0.02 per semester), individual deactivation for 4-6 weeks 0.04-0.09, full-dose broadband 0.08, the reduction meta on depression g=0.25 (CI 0.10-0.41) — negative causal effects exist, in a band of roughly 0.04-0.11 SD (mostly adults and college students). For scale: US young adults' emotional state worsened ~0.37 SD over 2008-2022.
So the honest answer is three sentences: the strong-skeptic position ("zero effect, pure reverse causation") is falsified by deactivation experiments, broadband experiments, and within-person directionality together; the "single largest reason" magnitude has never been supported by any causal design; and the middle — a small real effect, multiplied by near-universal exposure, stacked on a multi-causal crisis — is where the evidence actually points. Note that "too small to matter" is equivalent to rejecting every study in this literature at once, because 0.085 is the literature's mode, not an outlier.
Abrahamsson's Norwegian middle-school phone-ban study is the policy camp's most-cited evidence. It cleared peer review in March 2026 at the Journal of Human Resources (calling it a "working paper" is obsolete) — and the published abstract's first sentence is 【single-source verified, R3 dual-seat; figures from the 2024 WP version】:
"banning smartphones in middle schools has no average effect on education or mental health but masks important gender differences."
The full-sample null is this study's most robust, most quotable result. The circulating positives each need conditions: "girls' psychologist visits fell nearly 60%" is girls × specialist care × number of visits × third post-ban year, an intensive margin (the extensive margin — whether treated at all — shows zero; the 29% figure is the GP margin at p=0.076, not significant; the two are different outcomes, not two computations of one number); girls' GPA +0.08 SD has p=0.064, and the paper itself states it "cannot reject the null hypothesis that the coefficients between girls and boys are equal" — "girls gained, boys didn't" outruns the statistics; bullying declines are p=0.067/0.094 in the full sample. The audit seat added identification-level wounds: no never-treated control group, ban years recalled by principals in a 2019 survey, responding schools systematically better-off on the outcome variables, key p-values clustered at 0.03-0.09 with no multiplicity correction. The most informative descriptive fact is elsewhere: of 477 "phone-ban" schools, 52% merely required silent mode in class, and only 4 made phones physically inaccessible — in "Norway proves phone bans work," the words "phone ban" are themselves mislabeled.
The policy evidence has the same construct ladder as §0: Norway's "ban" = silent mode, England's "restrictive" = in the bag, America's pouches, Australia's account age-gate — four "bans" are four different treatment intensities, mutually incomparable. Everything summed, the supported statement is: "School phone policies in their current forms have zero average effect on mental health; academics may gain slightly, more plausibly the stricter the ban."
The under-16 minimum-age regime took effect 2025-12-10 (what took effect is the platform-obligation provision): the first national-level, platform-obligation-based social media age minimum with no parental-consent route that has actually been enforced — France legislated a 15-year digital-majority threshold in 2023, but with a parental-consent route and reportedly without implementation, so "the world's first" holds only under those qualifiers. Meta self-reported closing 544,052 accounts (IG 330,639 / FB 173,497 / Threads 39,916) 【vendor account, unverifiable; "about 500k" was the regulator's prior estimate — the two circulate interchangeably】; industry-wide, ~4.7 million accounts removed or restricted (eSafety). The first independent evaluation at ~3 months (The BMJ 2026; baseline 436/follow-up 408, convenience sample): past-week use of restricted platforms fell 95%→86% (ages 12-13) and 98%→89% (14-15) — "over 85% still use them" and "down about 9 points" are two faces of the same data; quoting either alone is selective. The most load-bearing result: the preregistered RDD at the age-16 threshold detected no discontinuity — but the authors state the study was underpowered with a questionable continuity assumption. Write "an underpowered quasi-experiment found no discontinuity," never "the study proved the ban failed" 【single-source verified】.
Pew (identical items across four waves): TikTok "almost constantly" flat for three years (16/17/16%), jumping to 21% in 2025 — the only riser among five platforms; while the overall "almost constantly online" share fell 46%→40% in 2025. "Usage keeps climbing" fails as a thesis sentence; the accurate version is "platform-level daily use stable at a high plateau, TikTok up against the trend, the overall intensity metric down for the first time." The circulating "teens spend 8 hours 39 minutes a day on screens" is Common Sense's 2021 pandemic-period caliber (Common Sense is also an advocacy organization for child online-safety legislation, and its census is self-published, not peer-reviewed) — recreational, content-summed (multi-screen double-counted), never updated since — it cannot be written as "today" 【multi-source】.
The US Surgeon General operates a precise two-register structure 【multi-source】: the institutional advisory (2023) holds the safety judgment at "we do not yet have enough evidence to determine if social media is sufficiently safe" (the words causal/causation appear zero times); Murthy's personal argumentative register reaches contributory causation — "social media is an important driver of that crisis" (press release quote), "has emerged as an important contributor" (op-ed). He never said "proved to cause" — but "the wording never left associated" is also false. The gradient is real, just steeper than circulated. The federal warning label has not passed; Minnesota and California have legislated state versions (the former reportedly suspended at once by a First Amendment suit) — "America has no social media warning label" is no longer accurate.
The academy's thermometer: a Delphi collective review (preprint, not peer-reviewed; the word "consensus" was removed from its own title by the authors in v3) had 120+ researchers rate 26 claims drawn from Haidt's book. Its most-quoted number (99.2% agreeing mental health declined) is a text-accuracy rating of a hedged "there is evidence that… albeit with some heterogeneity" statement, with "don't know" answers removed. Its two genuinely informative readings sit elsewhere: asked the overall impact of smartphones and social media, 66.7% chose "it depends on the context," with only 26% clearly negative; and on Haidt's three policies, the finalized conclusions all read "the evidence is too preliminary to support or challenge" — with 93-95% agreeing with that insufficiency verdict. Citing this survey as "expert consensus backs the policies" inverts its finalized conclusions. The main critics (Odgers/Orben/Przybylski) declined to participate; Ferguson quit and withdrew his name; the authors concede critics are underrepresented 【single-source verified, R3 dual-seat】.
Odgers's Nature review itself, incidentally, is the audit's only intact group (1 HOLDS): "not supported by science" and "no evidence that using these platforms is rewiring children's brains or driving an epidemic" are verbatim — but her sentence is a pair, and Haidt quoted only half on X: "Two things can be independently true… First, that there is no evidence… Second, that considerable reforms to these platforms are required, given how much time young people spend on them." And one landmine defused: "the kids are not all right" is not Odgers — it is the title of a pro-Haidt review; misattributing it inverts the stance.
Answered in layers, each with its evidence grade:
Where he is right. The reality of the crisis (Anglosphere girls' self-report + US hard outcomes) — 【multi-source】, with the artifact reading boxed in by constant-wording surveys, invariance tests, and coding-independent data; the existence of a negative causal effect — 【multi-source】, four design families concur; and the principled point that small effects at population scale can matter (even though one of his two benchmarks was misattributed). His own 2022 qualifiers — conceding no experiment used middle schoolers, offering r=.10-.15 as an honest range — are more careful than most of his amplifiers.
Where he is unproven. "Single largest reason" — no causal design supports that magnitude; between the confirmed causal band (0.04-0.11 SD) and the crisis magnitude (0.37 SD) lies a gap nobody has filled. "Girls r=.15-.22" — no primary source. The strong version of international synchrony — Germany and the timing don't line up. Policy efficacy — the three policies he championed are rated "evidence too preliminary" even by the Delphi review he co-authored.
Where he is wrong or corrected. The "50-150%" range (YRBS runs below its floor); the Kelly figure's "including controls" caption; the lead-benchmark attribution; drawing the 2007-trough suicide series into a tidy 2012 story.
The other camp's correction list is just as long. The potato comparison's near-zero-denominator artifact; the nominal-N rhetoric; Currie's "flat self-harm" retracted nationally by the authors themselves; the Nordic scissors overturned by Swedish and Finnish young-female data; "Germany is improving" true only for parent-reported boys; Ferguson's corrected meta and bidirectional duration moderation; "the study proved the ban failed" over-read from an underpowered Australian RDD.
Which is this audit's central structural finding: the main casualty of this debate is not either camp's position — it is both camps' headline numbers. Twenty-three of 24 load-bearing groups needed caliber corrections, split roughly evenly. A field with heavyweight scholars on both sides, public ledger documents, and massive data got here because rhetoric outran measurement — and because a genuine adversarial collaboration has never once happened.
Ordered by evidence strength, each with its test.
Load-bearing claims in this essay went through three-vote adversarial verification plus dual-seat audits of single-source empirics; faithful paraphrase does not equal truth, and evidence grades are marked throughout. Related on this site: The Evidence Hierarchy of Learning Science (how to read effect sizes and control groups), The "95% of AI Pilots Fail" Physical (the yardstick ladder for survey numbers).