中文EN
← Deep Research
Deep Research · Deep dive

The '70% of Transformations Fail' Autopsy: Can the Last Wave's Post-Mortems Predict AI-Native? (Deep Dive)

This is the deep-dive edition · read the plain-language edition →
TL;DR
'70% of transformations fail' has a precise birth record: born in 1993 as an 'unscientific estimate', retracted by its authors in 1995, rewritten in 2000 as an unsourced 'brutal fact', and gifted a fabricated pedigree in 2009 when McKinsey cited 'research Kotter published in 1995' that does not exist; a 2011 academic autopsy found no valid empirical basis, and the number lived on into the AI era anyway. The honest measured answer is 'undetermined': published estimates span 7-90%, the failure rate is a function of where you set the pass bar, and the measured distribution is fat-tailed rather than majority-failure. The Agile/DevOps archives show the organizational machinery replaying in AI adoption (64% of CEOs invest before understanding value, only 21% redesigned workflows, Duolingo ran the full mandate-to-retreat cycle), while shadow bottom-up adoption and capability doubling every ~7 months are variables the last archive never saw. Eleven testable claims close the essay.
90 votes · 30/30 survived1993 'unscientific' → 2009 invented sourcebase rate: estimates span 7-90%11 testable claims

Empirical citations in this essay are graded. The 30 load-bearing claims cited in the body (the verbatim citation-chain originals behind the number, first-hand failure-rate measurements, the Agile/DevOps archives, change-theory evidence, AI adoption data) were each challenged by 3 independent verifiers (word-for-word checks against primary sources, counter-evidence search, wording audits): all 30 survived refutation, with 20 of them receiving 40+ wording calibrations per the verifiers' notes; citations that did not go through verification are marked [unverified; source]. Methodological caveats (self-reported data / vendor interest / forecast vs. measurement) are disclosed inline; a source index closes the essay.

0. Why a number deserves an autopsy

Open the first page of almost any AI transformation proposal and you will meet the same sentence: "Research shows that 70% of transformations fail — but we know how to make you one of the 30%." This number has served as the change-consulting industry's shop sign for thirty years, and it is now being carried into the AI era untouched: across 2024-2026, vendor articles titled Why 70% of AI Transformations Fail and The 70% Rule for AI Change Management appeared in bulk, their sentence structure identical, word for word, to thirty years ago — scare first, sell the antidote second. [unverified; source: novoslo.com, aiassemblylines.com, among others]

This site's When Code Becomes Cheap argued that the hard part of AI-native transformation is organizational, not technical. The natural next question along that thread: the last great transformation to sweep the software industry — Agile and DevOps — left a complete archive. How much predictive power does that archive hold for this round? And the first step in inspecting an archive is inspecting the loudest number on its cover.

This essay does three things: performs the autopsy on "70%" (where it came from, how it mutated, why it cannot be killed); asks measurement for an honest base rate (what the real transformation failure rate is — spoiler: the question itself is broken); and assesses the predictive power of the last wave's post-mortems for AI-native (which failure mechanisms are replaying verbatim, and which structural conditions have changed).

1. The life of a number: 1993-2026

"70%" has a precise birth record, and the birth certificate says "unscientific."

1993, birth. Hammer and Champy's Reengineering the Corporation: "Our unscientific estimate is that as many as 50 to 70 percent of the organizations that undertake a reengineering effort do not achieve the dramatic results they intended." [verified; editions vary slightly between "50 to 70 percent" and "50 percent to 70 percent"] Note the three qualifiers: the estimate is "unscientific"; the scope is reengineering alone, one species of change; and "failure" is defined as not achieving the intended dramatic results — not blowing up, just falling short of one's own announced ambitions.

1995, the author retracts. Hammer corrects the record himself in The Reengineering Revolution: "…this simple descriptive observation has been widely misrepresented and transmogrified and distorted into a normative statement. There is no inherent success or failure rate for reengineering." [verified] The retraction is almost never cited — the number survived; the disclaimer died.

1995, the wrongly accused "source." The paper later fingered by countless citations as the origin of 70% is Kotter's famous HBR article of the same year, "Leading Change: Why Transformation Efforts Fail." Yet the original contains no overall failure-rate figure whatsoever, let alone 70%. Kotter watched over 100 companies and offered an impressionistic verdict: "A few of these corporate change efforts have been very successful. A few have been utter failures. Most fall somewhere in between, with a distinct tilt toward the lower end of the scale." The only failure-adjacent percentage in the entire piece is stage-specific: far more than 50% of the companies he watched stumbled at the first step (establishing a sense of urgency). [verified]

2000, the critical mutation. Beer and Nohria write in HBR: "The brutal fact is that about 70% of all change initiatives fail." No footnote, no study, no source. [verified] Here the "unscientific estimate" completed two sheddings at once: scope shedding (reengineering → all change) and disclaimer shedding (unscientific estimate → brutal fact).

2008, Kotter finally says 70% out loud. In A Sense of Urgency: "From years of study, I estimate that today more than 70% of needed change either fails to be launched, even though some people clearly see the need, fails to be completed, even though some people exhaust themselves trying, or finishes over budget, late, and with initial aspirations unmet." [verified] Still a personal estimate, still no study cited — and the definition is now wide enough to count "should have changed but never did" as failure.

2008-2009, the third mutation: a source gets invented. McKinsey's Keller and Aiken, in "The Inconvenient Truth About Change Management": "In 1995, John Kotter published research that revealed only 30 percent of change programs are successful." [verified] As established above, that figure does not exist in Kotter's 1995 article — nor does any overall percentage. "Published research that revealed" is academic packaging conjured from nothing. The era's empirical fig leaf was McKinsey's own July 2008 global survey: 3,199 executives assessed their own organizations' transformations; of those, the roughly 2,663 who had been through one gave ratings, and only about one-third said their organizations had successfully achieved their goals. In the same dataset, only about 5%-6% rated their transformation explicitly as "completely unsuccessful" — "two-thirds didn't claim success" was sold as "70% fail." [verified]

2011, the academic autopsy. Hughes publishes "Do 70 per cent of all organizational change initiatives really fail?" in the Journal of Change Management, tracing five published 70% claims one by one — Hammer & Champy 1993, Beer & Nohria 2000, Kotter 2008, Bain's Senturia et al. 2008, McKinsey's Keller & Aiken 2009 — and concluding, verbatim: "whilst the existence of a popular narrative of 70 percent organizational change failure is acknowledged, there is no valid and reliable empirical evidence to support such a narrative." [verified]

After 2011, the zombie years. Being autopsied did nothing to circulation: McKinsey's 2015 "Changing Change Management" repeats, sourceless, that "70 percent of change programs fail to achieve their goals" [verified]; the 2021 "Losing from day one" goes further and footnotes Kotter (Leading Change 1996 and A Sense of Urgency 2008) as the authority for 70% [verified] — the consultancy cites the guru, while the guru's "research source" for the number is the very invention this consultancy produced years earlier. The citation chain has closed into a loop. BCG then rebooted the number for digital transformation in 2020 (70% fall short — dissected below). Then came the 2024-2026 AI-flavored hosts.

This transmission chain has a ready-made taxonomy in medical bibliometrics: Greenberg's study of citation networks named these mechanisms citation bias, amplification and invention, with the memorable warning that "citation can be used to generate information cascades resulting in unfounded authority of claims." [unverified; source: BMJ 2009;339:b2680] The thirty-year journey of "70%" is a textbook specimen of all three: scope shedding is bias, serial retelling is amplification, and "Kotter 1995 published research" is invention.

The life of '70%': circulation above, debunking below ● top: the number cited / strengthened ● bottom: retraction & autopsy (rarely cited) 1993 2009 2026 1993: born 'unscientific', reengineering only 2000: upgraded to 'brutal fact', unsourced 2008-09: Kotter 'I estimate' + McKinsey invents source 2024: new host, AI transformation 1995: authors retract, 'no inherent failure rate' 2011: Hughes autopsy, 'no valid evidence'
Schematic: every mutation got louder (top track) while both corrections went uncited (bottom) — section 1 unpacks each dot

2. So what is the real failure rate? Measurement's honest answer

Archaeology can only prove that "70%" has no source; it cannot prove the failure rate isn't high. To answer "how much, really," you have to look at what the people who measured seriously actually measured — and what measurement itself ran into.

2.1 The most-cited "measurement": Standish CHAOS and its three dissections

The software industry's own "70%" is called the CHAOS report. The Standish Group's first report, 1994: 16.2% successful, 52.7% "challenged," 31.1% impaired/cancelled (365 respondents covering 8,380 application projects). [verified] Note the definition: "success" = on time, on budget, with all specified features delivered — clear all three bars or you lose. This report fed thirty years of "most software projects fail" narrative, but the three independent dissections it underwent deserve the record, one by one:

2.2 What careful measurers found: fat tails, not "most fail"

The Flyvbjerg group's project database is the largest-sample measurement in the field:

Read side by side, the lesson of these two datasets is not "most projects fail" but that the distribution is fat-tailed: the typical project's overrun is manageable; what devastates is the tail — that one-in-six eats a disproportionate share of the losses. And note the structure of the 0.5% figure: it is the same "clear every bar or you lose" definitional device as Standish's 16.2% — the more bars you add, the rarer "success" becomes. Declaring "99.5% fail" on a three-bar measure and declaring "only 19% fail" on cancellation rates uses the very same projects.

The measured danger is the fat tail, not 'most fail' (Flyvbjerg database) Mean IT cost overrun (~16,000-project database)+73%Mean overrun of the 1-in-6 'Black Swan' projects+200%Mean overrun of the 18% of IT projects >50% over+447%
Schematic: the mean is survivable, the tail is lethal — decision-to-build baseline, real terms; bars proportional to values, caveats in section 2.2

2.3 The academic base rate for organizational change: the answer is "no answer"

Step outside software projects, and the measurement record for organizational change at large is thinner still:

That is academia's honest answer: the real failure rate is "undetermined" — not because nobody has measured, but because "failure" has no shared definition, and the choice of definition determines the number.

2.4 The definition machine: how consulting numbers are manufactured

Once you see the leverage in definitions, today's circulating transformation statistics disassemble on sight:

Self-reported measures carry one more systemic problem: the executives doing the scoring are simultaneously the transformation's owners and its judges, and consulting surveys naturally use "did you achieve your original objectives" as the yardstick — objectives that were set during the sales phase. The institutions that sell transformation services also hold a monopoly on scoring whether transformations succeed. No conspiracy theory required; just notice that the final chapter of every such report is a services overview.

The definition machine: one dataset, two report cards BCG 2020: digital transformations (825 executives, self-graded) 30% met targets 44% created value, missed targets 26% limited value → packaged as: '70% fall short' Bain 2024: business transformations (400+ executives, self-graded) 12% ~75% got at least halfway, short of full ambition ~13% → packaged as: '88% fail' (same data: ~87% achieved at least half)
Schematic: set the pass bar at 'perfect' and the failure rate is whatever you need; segments are each report's own breakdown — section 2.4 unpacks both

3. The Agile archive: a transformation movement's own records

Agile was the software industry's last industry-wide transformation, and the archive it left has a unique asset: a seventeen-year questionnaire asking the same questions, plus a measurement promise that was never kept.

3.1 Seventeen years of surveys: the same challenges top the chart every year

The State of Agile survey (VersionOne → CollabNet → Digital.ai) is the Agile movement's mirror of itself — vendor-run, respondents self-selected; enter both facts into the record first. What it measured is interesting precisely because of that:

The academic systematic review corroborates from the side: Dikert, Paasivaara and Lassenius (2016), reviewing 52 publications on large-scale agile transformations (42 cases), found nearly 90% were experience reports rather than rigorous research; the single most frequent challenge was other functions' unwillingness to change (about 31% of cases); the authors conclude, verbatim: "large-scale agile cannot be just taken into use off-the-shelf" — it must be carefully customized. [verified]

3.2 The counter-evidence honesty demands: agile practices themselves may well work

Performing an autopsy on Agile is not a verdict that agile doesn't work. Jørgensen (IEEE Software 2019), analyzing 196 Norwegian software projects: projects using agile methods had better outcomes than non-agile across all size bands (small projects p<0.01, medium p≈0.03, large not significant; success self-rated, design correlational). [verified] This does not contradict the archive above — it completes the movement's most important lesson: practices and transformations are two different things. Iterative delivery and continuous integration correlate with better outcomes; the organizational ritual called "agile transformation" wholesales the name of the practices, not the practices themselves.

3.3 Scaling frameworks: an industry with almost zero evidence

If agile practices have evidence and agile transformations have thin evidence, then "scaled agile frameworks" are an evidence vacuum:

3.4 Two specimens: an imitated fiction, and a named case in retreat

4. The DevOps archive: the most seriously measured wave, and its ceiling

The DevOps generation took a genuine step forward in evidence over the Agile generation — and precisely because of that, its archive exposes where this kind of measurement tops out.

The progress is real. DORA/State of DevOps measures outcome metrics (deployment frequency, change lead time, change failure rate, time to restore), not satisfaction; the four key metrics can be measured directly from system data (Google open-sourced the Four Keys tooling), so others can recompute them on their own data. Compared with "do you feel agile," that is a paradigm upgrade.

But its longitudinal narrative cannot carry its own method. In the 2022 report, the "elite" performance cluster vanished outright — official wording: "Unlike in years past, there was no evidence of an 'Elite' cluster." The low-performer group jumped from 7% in 2021 to 19%, and DORA's offered explanation was an untested pandemic hypothesis. String the elite shares across the years into a line (7% in 2018 → 20% in 2019 → 26% in 2021 → vanished in 2022 → about 18-19% in 2023-24) and the "industry is improving" curve turns out not to be comparable: DORA's own FAQ concedes that the clusters re-emerge each year from that year's different respondents — it is not a calibrated industry index. Add self-reported questionnaires, possible self-selection bias (people who consider themselves elite are more willing to answer; The Register called this out in 2021), unpublished raw data, and no independent replication of the capabilities→performance path model — the most seriously measured wave could only be this serious. [verified]

A forecast got treated as a measurement. The widely circulated "75% of DevOps initiatives will fall short of expectations due to organizational learning and change issues" (Gartner, 2019, forecast horizon 2022; a 90% version ran to 2023) is an analyst prediction, never revisited for validation. [unverified; source: Gartner 2019 and its social-media posts] It entered countless slide decks as "research shows" anyway.

The named mega-case: GE Digital. The loudest case in the transformation-narrative economy delivered the hardest closing numbers: GE announced in 2015 the target of over $15 billion in software and solutions revenue by 2020 (from about $5 billion at the time); reportedly over $4 billion went into GE Digital in 2016 alone (third-party estimate; GE never formally disclosed a figure on that basis); in December 2018, GE announced it was reorganizing the digital business as an independently operated but still wholly owned company, at which point annual software revenue was about $1.2 billion — 8% of the target — and it sold the majority stake in ServiceMax (bought in 2016 for $915 million) to Silver Lake (the spin-off itself was never completed either; GE Digital ultimately folded into GE Vernova). [verified; note that GE's 2018 goodwill write-down of $22 billion belonged to GE Power, not Digital — the two are frequently conflated] All the while, GE-family transformation stories were still touring industry summits as success cases. [unverified; source: DOES agenda archives]

A survivor economy. The DevOps case library is structurally tilted toward success: enterprise summit stages are built from self-nominated success stories, and the movement's founding text (The Phoenix Project) is itself a novel; there is no breakout session for failures. [unverified; source: DOES 2016 press release, IT Revolution case library] Nobody is cheating; this is how such archives get generated — remember the mechanism, because it applies unedited when you read AI case collections in Section 7.

And this wave's archive already contains AI's first line of record. DORA 2024: for every 25% increase in AI adoption, delivery throughput was estimated to drop 1.5% and delivery stability 7.2%; in the 2025 report throughput turned positive, stability remained negative, and the official framing shifted to amplifier: "AI doesn't fix a team; it amplifies what's already there." verified; wording carried over from [When Code Becomes Cheap] The measurement infrastructure the last transformation built is already issuing report cards to the next one — the most valuable continuity between the two archives.

Two waves, two archives, side by side The Agile archive (2001-2025) The DevOps archive (2014-2025) MethodSentiment surveys, vendor-run, self-selectedOutcome metrics (four keys), recomputableThe tell'Culture/leadership' top challenge for 17 years2022: elite cluster vanished; years not comparableFrameworksSAFe: zero independent controlled studies in 9 yrsPath model never independently replicatedRetreatCapital One cut ~1,100 agile rolesGE: $15B target → $1.2B revenue at carve-out
Schematic: each wave measured harder than the last, and neither could produce an industry failure rate — sections 3-4 unpack each cell

5. The people selling the cure: change management's own evidence check-up

"70% will fail" is the scare; "follow our framework and you'll succeed" is the antidote. The scare has had its autopsy; now for the antidote.

This section's combined verdict: the change industry's scare number has no empirical source, its antidote frameworks have never been tested whole, and the buying behavior is driven by legitimacy — thirty years of selling smoke, then? No. What was sold really did do something; the mechanism was just Staw-Epstein: adopters got legitimacy, prestige and pay, and the consultants got revenue. The only thing never demonstrated is any effect on that denominator behind "70%."

6. The "frozen middle" retried: who actually kills transformations

Every transformation story has an official villain: the middle manager — "the top wants change, the bottom wants change, the middle is frozen." This "frozen middle" deserves its own hearing, because it is being carried, sentence intact, into the AI narrative.

The pedigree is suspect from birth. The phrase is usually attributed to 1980s GM CEO Roger Smith, but all that can be located is consulting blogs citing each other (even Dartmouth Tuck's case page does it) — no primary 1980s source whatsoever (speech, interview, contemporary reporting) can be pinned down. [verified (as an absence)] A concept used to explain transformation failure whose own provenance is a chain of retellings — the same species as "70%."

The academic archive exonerates the middle — conditionally. Wooldridge and Floyd (1990, a quantitative study of 20 organizations): middle-manager involvement in strategy formation correlates positively with organizational performance. Huy (2002, three-year field study, ASQ): in radical change, middle managers' "emotional balancing" work — pushing projects with one hand while catching subordinates' emotions with the other — is the mechanism by which adaptation happens at all. Balogun and Johnson (2004, AMJ): change derails mainly because, after top management withdraws, middle managers are left to their own sensemaking in a communication vacuum — not because of deliberate resistance. The nail on the other side is academic too: Guth and MacMillan (1986) confirmed that when middle managers judge a change to harm their own interests, they can slow a strategy, degrade it, or "totally sabotage" it. [verified] Taken together: the middle layer is a conditional reactor, not permafrost — the conditions being incentive alignment, trust and participation, and all three of those dials sit in the C-suite's hands.

The largest practitioner dataset points the finger upstairs. Prosci's biennial survey, running since 1998 (vendor self-published, practitioner sample — both into the record): "active and visible executive sponsorship" has ranked as the number-one contributor to change success in every single edition, with a mention rate roughly three times the runner-up; "engagement with middle managers" ranks seventh of seven. With extremely effective sponsorship, projects meet objectives at roughly 3.5 times the rate seen under extremely ineffective sponsorship (that multiplier is the current edition's; other editions run 2.5-2.9x). The same material also honestly records that middle managers are the most resistant group in its surveys — consistent with the "conditional reactor" reading: the resistance is measurable, but it is the dependent variable. [verified] A converging line of evidence comes from the negative side: MIT Sloan's Johnson, summarizing the sampling bias of change research — most studies cover only the first few months and interview only the leadership; the evidence for "change failed, blame the lazy managers" is manufactured exactly that way. [unverified; source: strategy+business 2020]

Even "participation," the movement's proudest brand, is conditional. Change management's genesis experiment — Coch and French (1948, the Harwood pajama factory) — is taught in textbooks as "participation removes resistance": the no-participation group's output fell from about 60 units/hour to about 50 (roughly 17-20%) with no recovery over 32 days; the full-participation groups recovered and exceeded pre-change levels by about 14%. But the original experiment's control group had only 18 people, the experimental groups 13/8/7; 17% of the control group quit within the first 40 days; Bartlem and Locke (1981) pointed out that explanation, training, job availability and piece-rate fairness were all confounds; and the Norwegian replication (French, Israel and Ås, 1960) detected no output effect. [verified] Participation→commitment is a conditional heuristic, not a law — a direct calibration for the AI mandate-versus-bottom-up dispute below.

The frozen-middle narrative's new job in the AI era — and what the data says. "AI frozen middle" articles appeared in bulk across 2025-2026, but the reports they cite frequently say the opposite: McKinsey's Superagency (2025) concludes that employees are ready and leadership is the biggest bottleneck; Kyndryl's 2025 survey finding that "45% of CEOs think their employees are resistant to AI" is CEOs' perception of employees at large, retrofitted by blogs into "the middle is blocking AI." [unverified; source: the original reports and their re-citations] And the data that directly measures usage puts the gradient the opposite way entirely — leaders > middle managers > frontline (Gallup, Q4 2025: using AI a few times a week or more, leaders about 44%, managers about 30%, frontline about 23%). [verified] The middle is not AI's permafrost; if there is permafrost, it is elsewhere (next section).

Change fatigue is real; its numbers are a mess. The employee-side archive: the average employee experienced 10 planned enterprise changes in 2022, versus just 2 in 2016 (HBR 2023, Gartner authors). [verified] The fill-in-the-blank — "willingness to support change fell from 74% in 2016 to __% in 2022" — has been given two answers by Gartner's own publications: 43% (October 2022 release materials) and 38% (a Q1 2023 periodical). This essay initially ruled the 43% a transmission error; the verifiers corrected us: the discrepancy lives inside Gartner itself, coexisting with a coincidentally same-numbered "intent to stay 43%/74%" pair of statistics. [verified] A statistic about change wearing people out has worn itself out into two versions — that is not a quip; it is a free teaching aid for this essay's methodology.

7. Assessing predictive power: what this autopsy says about AI-native

Now close the files and answer the title question. List the failure mechanisms of the first six sections and check them against the 2024-2026 AI adoption record, and you get a sheet that is mostly checkmarks, with three fractures.

7.1 Mechanisms replaying, unmodified

7.2 Three fractures: where this round is different

7.3 Closing the book

The last wave's post-mortems' predictive power for AI-native compresses into three sentences:

  1. At the level of organizational mechanisms, predictive power is high. Mimetic adoption, definition-driven scare numbers, ritual substituting for substance, the Goodhart backlash to mandates, the misdirected blame-the-middle narrative — all five have already recurred in the AI adoption record, most of them with quantitative evidence. The shape of failure will rhyme.
  2. At the level of the failure rate, predictive power is zero — because that number never existed. "70% of AI transformations fail" was circulating before anyone had measured it, and that is the whole point of this essay's archaeology: it is a rhetorical device, not a measurement. On the same grounds, be pre-armed against the next generation of candidate zombie numbers, "95% of AI pilots fail" and kin (that's another essay's topic).
  3. At the level of technical dynamics, the analogy fractures, direction unresolved. Bottom-up shadow adoption and exponentially improving tools are variables absent from the last archive; they could make this round genuinely different, or merely relocate the same organizational diseases to a new lesion. The instrument for telling the two apart happens to be the last wave's finest legacy: outcome measurement. When Code Becomes Cheap's prescription closes its loop here — keep your transformation's books with delivery data, not gut feel and questionnaires; DORA-style dashboards are already grading AI. Don't let the next thirty years navigate by a made-up percentage.
Last wave's script vs this wave's reality Mechanisms replaying Conditions that broke Herd: 64% of CEOs invest before seeing valueMandates: Duolingo, mandate→retreat in 12 monthsRitual: only 21% redesigned workflowsCert boom: AI certs near 30%, ~20x pre-ChatGPTInverted: 78% BYOAI, 57% hide their useCapability doubles ~every 7 monthsSeat-based SaaS layer (half a break)
Schematic: the organizational machinery replays (left) while three structural conditions have no precedent in the last archive (right) — section 7 unpacks each

8. Coda: eleven testable claims

Ordered by evidence strength:

  1. "70% of transformations fail" has no empirical source: it was born in 1993 as an "unscientific estimate," retracted by its author in 1995, rewritten in 2000 as a sourceless "brutal fact," and fitted in 2008-09 with the invented provenance of "Kotter's 1995 research." (strong: full chain checked word for word, adversarially verified)
  2. Kotter's famous 1995 article contains no overall failure rate of any kind; he first personally gave 70% in 2008, and called it an estimate. (strong: primary-text check; the only percentage is the stage-level observation that far more than 50% fail at step one)
  3. Hughes's 2011 academic trace holds — "no valid and reliable empirical evidence" behind the five published instances — and the number kept circulating after its autopsy, including McKinsey 2015/2021 and the 2024-26 AI variants. (strong)
  4. The measurable true base rate is "undetermined": the academic range is 7-90% with a mean around 50%, on evidence that is "outdated, fragmentary, fragile or just absent"; the failure rate is a function of the definition — the same dataset can produce "88% fail" and "87% achieved at least half." (strong: Cândido & Santos; BCG's and Bain's own splits)
  5. The measured distribution of software projects is fat-tailed, not "mostly failing": average overruns of 27-73%, but the one-in-six black swans average +200%, and the IT projects overrunning by more than 50% average +447%. (strong: two generations of Flyvbjerg data; the baseline definition is academically contested)
  6. The CHAOS report's figures are an artifact of forecasting bias: flip the direction of an organization's estimation bias and its "success rate" swings 5.8% ↔ 94.2%; Standish's chairman himself calls the reports "opinion." (strong: the Eveleens & Verhoef recomputation plus the on-the-record exchange)
  7. The Agile archive's self-diagnosis went unchanged for seventeen years — culture and leadership were always the top self-reported obstacles — while the scaling-framework industry produced not one independent controlled outcome study across 2016-2025; the largest independent comparison measured the practical difference of framework choice as negligible. (medium-strong: negative signals from vendor surveys + Dikert + Verwijs & Russo, all self-reported)
  8. Agile practices themselves correlate with better project outcomes (all size bands; significant for small and medium) — the failure archive belongs to the "transformation ritual," not to the practices. (medium: single-country sample, self-rated, correlational)
  9. The change industry's antidote frameworks have never been tested whole: Kotter's eight steps have no whole-model controlled study, Lewin's three-step model is a posthumous construction (a claim with an academic opposing party), the one systematic reconciliation (Stouten 2018) judges the foundations to be expert opinion; and the empirical driver of adoption is legitimacy and CEO pay, not performance. (medium-strong: Appelbaum / Cummings [contested by Burnes] / Stouten / Staw & Epstein)
  10. The "frozen middle" has no traceable source, and the empirical direction runs against it: middle-manager involvement correlates positively with performance, blocking behavior is a conditional response to misaligned incentives; the largest practitioner dataset gives the top factor to executive sponsorship (middle-manager engagement ranks seventh); and the measured AI-usage gradient — leaders 44% > managers 30% > frontline 23% — points the opposite way from the frozen-middle narrative. (medium: conditional academic evidence + Prosci vendor self-report + Gallup)
  11. AI transformation is replaying the last round's organizational mechanisms (64% of CEOs self-report investing before understanding, only 21% have redesigned workflows, Duolingo has run the full mandate→withdrawal cycle, and the consulting/certification layer is reorganizing at multi-billion-dollar scale), while diverging structurally in two places: shadow bottom-up adoption (78% BYOAI, 57% concealment) and tool capability doubling every 7 months. (medium-strong: the first-hand vendor/survey numbers are verified, but all are fast-moving variables, as of July 2026)

Signals worth watching: whether Meta's AI performance metric survives 2026 (the Duolingo script predicts it retreats); where McKinsey's "21% redesigned workflows" moves in the 2026-27 surveys (up = substance beginning, flat = the ritual phase extended); the sign of DORA 2026's AI-stability reading; whether the scaled-agile-framework industry produces its first independent controlled study in the AI era (base rate: nine years, zero); and whether the citation chain of the next zombie number — "95% of AI pilots fail" — re-runs Section 1 of this essay (we intend to give it an autopsy too, when the time comes). Thirty years ago, a self-described unscientific estimate got promoted to an industry's navigation system; this round's organizations have something that era lacked — their own delivery data. Use it.


Appendix: Primary Sources

Number archaeology: Hammer & Champy, Reengineering the Corporation (1993) · Hammer & Stanton, The Reengineering Revolution (1995) · Kotter, "Leading Change: Why Transformation Efforts Fail" (HBR, 1995) · Beer & Nohria, "Cracking the Code of Change" (HBR, 2000) · Kotter, A Sense of Urgency (2008) · Keller & Aiken, "The Inconvenient Truth About Change Management" (McKinsey, 2009) · McKinsey Global Survey, "Creating Organizational Transformations" (2008-07) · Hughes, "Do 70 per cent of all organizational change initiatives really fail?" (J. Change Management 11(4), 2011, DOI 10.1080/14697017.2011.630506) · Ewenstein, Smith & Sologar, "Changing Change Management" (McKinsey, 2015) · McKinsey, "Losing from day one" (2021-12) · Greenberg, "How citation distortions create unfounded authority" (BMJ 2009;339:b2680) · Tourish, Management Studies in Crisis (CUP, 2019)

Base-rate measurement: Standish CHAOS Reports (1994/2015/2020) · Eveleens & Verhoef, "The Rise and Fall of the Chaos Report Figures" (IEEE Software 27(1), 2010) · Jørgensen & Moløkken, "How large are software cost overruns?" (IST 48(4), 2006) · Flyvbjerg & Budzier, "Why Your IT Project May Be Riskier Than You Think" (HBR, 2011; arXiv:1304.0265) · Flyvbjerg & Gardner, How Big Things Get Done (2023) · Smith, "Success rates for different types of organizational change" (Performance Improvement 41(1), 2002) · Cândido & Santos, "Strategy implementation: What is the failure rate?" (JMO 21(2), 2015) · BCG, "Flipping the Odds of Digital Transformation Success" (2020-10) · Bain, 88% press release (2024-04) and Mankins & Litre, "Transformations That Work" (HBR, 2024) · Sauer, Gemino & Reich (CACM 50(11), 2007) · Loureiro et al. (Heliyon, 2024) · Ika & Pinto, "The re-meaning of project success" (IJPM 40(7), 2022)

The Agile archive: Digital.ai/VersionOne, State of Agile Reports, editions 14-18 · Dikert, Paasivaara & Lassenius (JSS 119, 2016) · Jørgensen, "Relationships Between Project Size, Agile Practices, and Successful Software Development" (IEEE Software 36(2), 2019) · Scaled Agile, framework.scaledagile.com/about and scaledagile.com marketing pages · Putta, Paasivaara & Lassenius (XP/PROFES 2018) · Verwijs & Russo, "Do Agile Scaling Approaches Make A Difference?" (EMSE, 2024; arXiv:2310.06599) · USAF CSO Chaillan, Memorandum for Record on Agile Frameworks (2019-12-28) · Jeremiah Lee, "Spotify's Failed #SquadGoals" (2020) · Kniberg (blog.crisp.se, 2015) · Banking Dive, Capital One coverage (2023-01) · Fowler, "FlaccidScrum" (2009) · Dave Thomas, "Agile is Dead (Long Live Agility)" (2014)

The DevOps archive: DORA/Google Cloud, State of DevOps / DORA Reports (2018-2025) and dora.dev/faq · Forsgren, Humble & Kim, Accelerate (2018) · Sallin et al. (XP 2021) · The Register (2021-09) · Gartner, "The Secret to DevOps Success" (2019) · GE press releases (2015-09-29, 2018-12-13) · The Conversation, GE Digital analysis (2018) · Keunwoo Lee, review of Accelerate and Humble's response

Change theory and the middle layer: Appelbaum, Habashy, Malo & Shafiq (JMD 31(8), 2012) · Cummings, Bridgman & Brown (Human Relations 69(1), 2016) · Burnes (JABS 56(1), 2020) · Stouten, Rousseau & De Cremer (AMA 12(2), 2018) · Pollack & Pollack (SPAR 28, 2015) · DiMaggio & Powell (ASR 48(2), 1983) · Staw & Epstein (ASQ 45(3), 2000) · Westphal, Gulati & Shortell (ASQ 42(2), 1997) · Abrahamson (AMR 21(1), 1996) · Abrahamson & Fairchild (ASQ 44(4), 1999) · Wooldridge & Floyd (SMJ 11(3), 1990) · Huy (ASQ 47(1), 2002) · Balogun & Johnson (AMJ 47(4), 2004) · Guth & MacMillan (SMJ 7(4), 1986) · Oreg, Vakola & Armenakis (JABS 47(4), 2011) · Coch & French (Human Relations 1(4), 1948) · Bartlem & Locke (1981) · French, Israel & Ås (1960) · Prosci, "Top Contributors to Success" · Gartner change-fatigue series (2022-2023) · O Morain & Aykens (HBR, 2023-05) · Reichers, Wanous & Austin (AME 11(1), 1997)

The AI adoption record: Lütke memo (X, 2025-04-07) · Fortune, coverage of von Ahn's podcast remarks (2026-04-13) · Business Insider/Entrepreneur, coverage of the Liuson internal message (2025-06) · HR Grapevine, Meta performance policy (2025-11) · TechCrunch, Coinbase (2025-08) · McKinsey, "The State of AI" (2025-03) · IBM/Oxford Economics CEO survey (2025-05) · Accenture 8-K (SEC, FY2024/FY2025) · BCG annual results press release (2025) · WSJ/Alex Singla, QuantumBlack (2025) · Revelio Labs, AI certification analysis (2026) · Microsoft/LinkedIn Work Trend Index (2024) · KPMG × University of Melbourne, "Trust, attitudes and use of AI" (2025) · Gallup, Q4 2025 workplace AI use (gallup.com/workplace/701195) · METR, "Measuring AI Ability to Complete Long Tasks" (arXiv:2503.14499) · DORA 2024/2025 (wording carried over from this site's When Code Becomes Cheap verified record)

The research materials and all verification rulings live in the research base (7 research threads; 30 load-bearing claims × 3 adversarial verification votes: 30/30 survived, 20 with wording calibrations, 0 overturned).