'70% of transformations fail' has a precise birth record: born in 1993 as an 'unscientific estimate', retracted by its authors in 1995, rewritten in 2000 as an unsourced 'brutal fact', and gifted a fabricated pedigree in 2009 when McKinsey cited 'research Kotter published in 1995' that does not exist; a 2011 academic autopsy found no valid empirical basis, and the number lived on into the AI era anyway. The honest measured answer is 'undetermined': published estimates span 7-90%, the failure rate is a function of where you set the pass bar, and the measured distribution is fat-tailed rather than majority-failure. The Agile/DevOps archives show the organizational machinery replaying in AI adoption (64% of CEOs invest before understanding value, only 21% redesigned workflows, Duolingo ran the full mandate-to-retreat cycle), while shadow bottom-up adoption and capability doubling every ~7 months are variables the last archive never saw. Eleven testable claims close the essay.
Empirical citations in this essay are graded. The 30 load-bearing claims cited in the body (the verbatim citation-chain originals behind the number, first-hand failure-rate measurements, the Agile/DevOps archives, change-theory evidence, AI adoption data) were each challenged by 3 independent verifiers (word-for-word checks against primary sources, counter-evidence search, wording audits): all 30 survived refutation, with 20 of them receiving 40+ wording calibrations per the verifiers' notes; citations that did not go through verification are marked [unverified; source]. Methodological caveats (self-reported data / vendor interest / forecast vs. measurement) are disclosed inline; a source index closes the essay.
0. Why a number deserves an autopsy
Open the first page of almost any AI transformation proposal and you will meet the same sentence: "Research shows that 70% of transformations fail — but we know how to make you one of the 30%." This number has served as the change-consulting industry's shop sign for thirty years, and it is now being carried into the AI era untouched: across 2024-2026, vendor articles titled Why 70% of AI Transformations Fail and The 70% Rule for AI Change Management appeared in bulk, their sentence structure identical, word for word, to thirty years ago — scare first, sell the antidote second. [unverified; source: novoslo.com, aiassemblylines.com, among others]
This site's When Code Becomes Cheap argued that the hard part of AI-native transformation is organizational, not technical. The natural next question along that thread: the last great transformation to sweep the software industry — Agile and DevOps — left a complete archive. How much predictive power does that archive hold for this round? And the first step in inspecting an archive is inspecting the loudest number on its cover.
This essay does three things: performs the autopsy on "70%" (where it came from, how it mutated, why it cannot be killed); asks measurement for an honest base rate (what the real transformation failure rate is — spoiler: the question itself is broken); and assesses the predictive power of the last wave's post-mortems for AI-native (which failure mechanisms are replaying verbatim, and which structural conditions have changed).
1. The life of a number: 1993-2026
"70%" has a precise birth record, and the birth certificate says "unscientific."
1993, birth. Hammer and Champy's Reengineering the Corporation: "Our unscientific estimate is that as many as 50 to 70 percent of the organizations that undertake a reengineering effort do not achieve the dramatic results they intended." [verified; editions vary slightly between "50 to 70 percent" and "50 percent to 70 percent"] Note the three qualifiers: the estimate is "unscientific"; the scope is reengineering alone, one species of change; and "failure" is defined as not achieving the intended dramatic results — not blowing up, just falling short of one's own announced ambitions.
1995, the author retracts. Hammer corrects the record himself in The Reengineering Revolution: "…this simple descriptive observation has been widely misrepresented and transmogrified and distorted into a normative statement. There is no inherent success or failure rate for reengineering." [verified] The retraction is almost never cited — the number survived; the disclaimer died.
1995, the wrongly accused "source." The paper later fingered by countless citations as the origin of 70% is Kotter's famous HBR article of the same year, "Leading Change: Why Transformation Efforts Fail." Yet the original contains no overall failure-rate figure whatsoever, let alone 70%. Kotter watched over 100 companies and offered an impressionistic verdict: "A few of these corporate change efforts have been very successful. A few have been utter failures. Most fall somewhere in between, with a distinct tilt toward the lower end of the scale." The only failure-adjacent percentage in the entire piece is stage-specific: far more than 50% of the companies he watched stumbled at the first step (establishing a sense of urgency). [verified]
2000, the critical mutation. Beer and Nohria write in HBR: "The brutal fact is that about 70% of all change initiatives fail." No footnote, no study, no source. [verified] Here the "unscientific estimate" completed two sheddings at once: scope shedding (reengineering → all change) and disclaimer shedding (unscientific estimate → brutal fact).
2008, Kotter finally says 70% out loud. In A Sense of Urgency: "From years of study, I estimate that today more than 70% of needed change either fails to be launched, even though some people clearly see the need, fails to be completed, even though some people exhaust themselves trying, or finishes over budget, late, and with initial aspirations unmet." [verified] Still a personal estimate, still no study cited — and the definition is now wide enough to count "should have changed but never did" as failure.
2008-2009, the third mutation: a source gets invented. McKinsey's Keller and Aiken, in "The Inconvenient Truth About Change Management": "In 1995, John Kotter published research that revealed only 30 percent of change programs are successful." [verified] As established above, that figure does not exist in Kotter's 1995 article — nor does any overall percentage. "Published research that revealed" is academic packaging conjured from nothing. The era's empirical fig leaf was McKinsey's own July 2008 global survey: 3,199 executives assessed their own organizations' transformations; of those, the roughly 2,663 who had been through one gave ratings, and only about one-third said their organizations had successfully achieved their goals. In the same dataset, only about 5%-6% rated their transformation explicitly as "completely unsuccessful" — "two-thirds didn't claim success" was sold as "70% fail." [verified]
2011, the academic autopsy. Hughes publishes "Do 70 per cent of all organizational change initiatives really fail?" in the Journal of Change Management, tracing five published 70% claims one by one — Hammer & Champy 1993, Beer & Nohria 2000, Kotter 2008, Bain's Senturia et al. 2008, McKinsey's Keller & Aiken 2009 — and concluding, verbatim: "whilst the existence of a popular narrative of 70 percent organizational change failure is acknowledged, there is no valid and reliable empirical evidence to support such a narrative." [verified]
After 2011, the zombie years. Being autopsied did nothing to circulation: McKinsey's 2015 "Changing Change Management" repeats, sourceless, that "70 percent of change programs fail to achieve their goals" [verified]; the 2021 "Losing from day one" goes further and footnotes Kotter (Leading Change 1996 and A Sense of Urgency 2008) as the authority for 70% [verified] — the consultancy cites the guru, while the guru's "research source" for the number is the very invention this consultancy produced years earlier. The citation chain has closed into a loop. BCG then rebooted the number for digital transformation in 2020 (70% fall short — dissected below). Then came the 2024-2026 AI-flavored hosts.
This transmission chain has a ready-made taxonomy in medical bibliometrics: Greenberg's study of citation networks named these mechanisms citation bias, amplification and invention, with the memorable warning that "citation can be used to generate information cascades resulting in unfounded authority of claims." [unverified; source: BMJ 2009;339:b2680] The thirty-year journey of "70%" is a textbook specimen of all three: scope shedding is bias, serial retelling is amplification, and "Kotter 1995 published research" is invention.
Schematic: every mutation got louder (top track) while both corrections went uncited (bottom) — section 1 unpacks each dot
2. So what is the real failure rate? Measurement's honest answer
Archaeology can only prove that "70%" has no source; it cannot prove the failure rate isn't high. To answer "how much, really," you have to look at what the people who measured seriously actually measured — and what measurement itself ran into.
2.1 The most-cited "measurement": Standish CHAOS and its three dissections
The software industry's own "70%" is called the CHAOS report. The Standish Group's first report, 1994: 16.2% successful, 52.7% "challenged," 31.1% impaired/cancelled (365 respondents covering 8,380 application projects). [verified] Note the definition: "success" = on time, on budget, with all specified features delivered — clear all three bars or you lose. This report fed thirty years of "most software projects fail" narrative, but the three independent dissections it underwent deserve the record, one by one:
Jørgensen and Moløkken (2006) pointed out that Standish's 189% average cost overrun is roughly 5-6 times the 33-34% measured by three earlier peer-reviewed surveys (Jenkins 1984, Phan 1988, Bergeron 1992) — the authors note the two sides' measures are not fully comparable (Standish counts only "challenged" projects), but the gap is too large for that to explain — and stated bluntly that there may be serious problems in survey design and analysis method, for instance sampling possibly biased strongly toward "failure projects," while Standish never publishes its methodology. [verified]
Eveleens and Verhoef (2010, IEEE Software) landed the most thorough blow: they applied Standish's definitions to 5,457 budget forecasts across 1,211 real projects and found the CHAOS figures are an artifact of forecasting bias — an unbiased, best-in-class organization scores only about 35% "success" by the Standish definition (on the cost-plus-functionality measure). They then constructed a mirror organization (a counterfactual with identical bias magnitude, opposite direction), whose Standish "success rate" flipped from 5.8% to 94.2%. The verdict, verbatim: Standish's success/challenged definitions suffer four problems — they are "misleading, one-sided, pervert the estimation practice, and result in meaningless figures." Confronted directly, Standish chairman Jim Johnson gave an answer worth framing: "All data and information in the Chaos reports and all Standish reports should be considered Standish opinion and the reader bears all risk in the use of this opinion." [verified]
Standish's own numbers fight each other. Its success rate oscillated between 16-35% across 1994-2009 and its failure rate between 18-40%, non-monotonically; in the 2015 report, the same projects score 36-41% successful under the traditional triple-constraint definition and only 27-31% under the new "on time, on budget, with a satisfactory result" definition — change the definition and the same projects' report card moves 10 percentage points. [verified] For the record: even under Standish's harshest measure, outright failure (cancellation) never reached 70%; the 2020 report ran roughly 31% successful / 50% challenged / 19% failed. [verified]
2.2 What careful measurers found: fat tails, not "most fail"
The Flyvbjerg group's project database is the largest-sample measurement in the field:
1,471 IT projects (2011, HBR): average cost overrun only 27% — but one in six projects is a "black swan," averaging roughly 200% cost overrun and 70% schedule overrun. [verified]
A database of roughly 16,000 projects (2023, How Big Things Get Done): 47.9% of projects meet or beat budget; only 8.5% meet both budget and schedule; only 0.5% meet budget, schedule and benefits. IT projects average +73% overrun (in real prices, against the decision-to-build baseline); and the 18% of IT projects that overrun by more than 50% average +447%. [verified; the "decision-to-build" baseline is academically contested — Love & Ahiaga-Dagbui and Ika argue it inflates overruns relative to a contract baseline [unverified; source: the associated commentary literature]]
Read side by side, the lesson of these two datasets is not "most projects fail" but that the distribution is fat-tailed: the typical project's overrun is manageable; what devastates is the tail — that one-in-six eats a disproportionate share of the losses. And note the structure of the 0.5% figure: it is the same "clear every bar or you lose" definitional device as Standish's 16.2% — the more bars you add, the rarer "success" becomes. Declaring "99.5% fail" on a three-bar measure and declaring "only 19% fail" on cancellation rates uses the very same projects.
Schematic: the mean is survivable, the tail is lethal — decision-to-build baseline, real terms; bars proportional to values, caveats in section 2.2
2.3 The academic base rate for organizational change: the answer is "no answer"
Step outside software projects, and the measurement record for organizational change at large is thinner still:
Smith (2002) compiled 49 published studies (covering more than 40,000 organizations) and reported median reported success rates by change type: restructuring/downsizing about 46%, merger integration about 33%, software system installation about 26%, culture change about 19%. [verified] Two qualifiers must travel with these numbers: they are medians of each study's self-reported measures (mostly practitioner surveys and self-assessments), and the study counts under each category are tiny (culture change: about 3 studies) — fragile medians, not stable base rates.
Cândido and Santos (2015) systematically reviewed strategy-implementation failure rates: published estimates run from 7% to 90% (28-90% for general business strategy, dipping to 7% once studies of specific strategies are counted in), with a mean of about 50%; the verdict, verbatim: "Most of the estimates presented in the literature are based on evidence that is outdated, fragmentary, fragile or just absent." [verified]
That is academia's honest answer: the real failure rate is "undetermined" — not because nobody has measured, but because "failure" has no shared definition, and the choice of definition determines the number.
2.4 The definition machine: how consulting numbers are manufactured
Once you see the leverage in definitions, today's circulating transformation statistics disassemble on sight:
BCG 2020: 70% of digital transformations "fall short of their objectives" (825 executive self-assessments plus about 70 BCG client cases). Look at its own three-way split: 30% fully met targets, 44% created some value but missed targets, 26% created limited value. [verified] Of the 70, 44 percentage points are "partial success" — booked as failure. The same maneuver as 1993's "do not achieve the dramatic results they intended."
Bain 2024: "88% of business transformations fail to achieve their original ambitions" (400+ executives). The companion article to the same research states: define failure as achieving less than half the targets, and only about 13% qualify. [verified] The same dataset can be "88% fail" or "roughly 87% achieved at least half" — depending entirely on whether you set the passing line at "ambition" or at "halfway."
The earlier classics include a unit slippage: KPMG New Zealand's 2010 survey finding — 70% of organizations had at least one project failure in the prior 12 months — circulated as "KPMG: 70% of projects fail." [unverified; source: the Calleam compilation]
Self-reported measures carry one more systemic problem: the executives doing the scoring are simultaneously the transformation's owners and its judges, and consulting surveys naturally use "did you achieve your original objectives" as the yardstick — objectives that were set during the sales phase. The institutions that sell transformation services also hold a monopoly on scoring whether transformations succeed. No conspiracy theory required; just notice that the final chapter of every such report is a services overview.
Schematic: set the pass bar at 'perfect' and the failure rate is whatever you need; segments are each report's own breakdown — section 2.4 unpacks both
3. The Agile archive: a transformation movement's own records
Agile was the software industry's last industry-wide transformation, and the archive it left has a unique asset: a seventeen-year questionnaire asking the same questions, plus a measurement promise that was never kept.
3.1 Seventeen years of surveys: the same challenges top the chart every year
The State of Agile survey (VersionOne → CollabNet → Digital.ai) is the Agile movement's mirror of itself — vendor-run, respondents self-selected; enter both facts into the record first. What it measured is interesting precisely because of that:
The challenge chart hasn't changed its cast in a decade. The 14th edition (2019, 1,121 respondents), top five: general organizational resistance to change 48%, insufficient leadership participation 46%, inconsistent processes and practices across teams 45%, organizational culture at odds with agile values 44%, inadequate management support and sponsorship 43% — four of the five form a culture-and-leadership cluster (in the 15th edition the single top item was process inconsistency, with the culture items at positions 2-3). By the 17th edition (2023), the top spot was still "general resistance to change / culture clash" (47%, as the biggest barrier to business-side agile adoption). [verified] A movement that sold culture change spent a dozen-plus consecutive years reporting, in its own survey, "culture didn't change" as its biggest obstacle.
Satisfaction collapse and ebb tide. 17th edition: satisfaction with the organization's agile practices fell from 71% the prior year to 59%. 18th edition (surveyed July-August 2025, sample down to just 349 people — the sample size is itself an ebb-tide signal): 74% use hybrid/homegrown methods (16th edition: about 50%, question wording not fully comparable); only 13% say agile is deeply rooted in their organization; 42% rate it "better than nothing"; 24% cut their agile investment. [verified]
Its positive numbers are unusable. The 15th edition claimed software-team agile adoption jumped from 37% to 86% within one year — swap in a different batch of people in a self-selected sample and you can manufacture that kind of "growth." [verified] Within the same questionnaire, the negative signals (satisfaction, the challenge chart) are more credible than the positive ones (adoption rates), because the vendor has no incentive to inflate the former.
The academic systematic review corroborates from the side: Dikert, Paasivaara and Lassenius (2016), reviewing 52 publications on large-scale agile transformations (42 cases), found nearly 90% were experience reports rather than rigorous research; the single most frequent challenge was other functions' unwillingness to change (about 31% of cases); the authors conclude, verbatim: "large-scale agile cannot be just taken into use off-the-shelf" — it must be carefully customized. [verified]
3.2 The counter-evidence honesty demands: agile practices themselves may well work
Performing an autopsy on Agile is not a verdict that agile doesn't work. Jørgensen (IEEE Software 2019), analyzing 196 Norwegian software projects: projects using agile methods had better outcomes than non-agile across all size bands (small projects p<0.01, medium p≈0.03, large not significant; success self-rated, design correlational). [verified] This does not contradict the archive above — it completes the movement's most important lesson: practices and transformations are two different things. Iterative delivery and continuous integration correlate with better outcomes; the organizational ritual called "agile transformation" wholesales the name of the practices, not the practices themselves.
3.3 Scaling frameworks: an industry with almost zero evidence
If agile practices have evidence and agile transformations have thin evidence, then "scaled agile frameworks" are an evidence vacuum:
SAFe's evidence is entirely home-grown. The Scaled Agile site: "Seventy percent of Fortune 100 companies... have certified SAFe professionals" — mind the measure: that is having certified employees, not using the framework; add the self-reported "20,000+ enterprises" and "30-75% faster time-to-market." Putta et al.'s multivocal review (XP 2018) found SAFe's evidence base to be overwhelmingly grey literature, much of it published on Scaled Agile's own website, with business benefits appearing only in vendor case studies. Across 2016-2025, neither we nor the verifiers found a single independent controlled outcome study — the void is itself the finding. [verified]
The largest independent comparison measured "no practical difference." Verwijs and Russo (EMSE 2024; about 15,000 agile team members, 4,013 teams, plus 1,841 stakeholders) compared the scaling routes (SAFe, LeSS, Scrum of Scrums, homegrown, no scaling): statistically significant differences, but effect sizes too small to have practical meaning — the paper's own wording is that framework choice "does not markedly influence" team effectiveness; the strongest predictor of effectiveness was the team's agile experience, not which framework it used. [verified; effectiveness self-reported] This single shot passes through both camps: the vendors' "30-75% faster" and the critics' "SAFe is uniquely harmful."
A buyer voting with its feet. US Air Force Chief Software Officer Chaillan's December 2019 memorandum, in black and white: "Programs are highly discouraged from using rigid, prescriptive frameworks such as the Scaled Agile Framework (SAFe)". [verified; note the USAF had renewed engagement with Scaled Agile in 2020 — this document is not a standing ban [unverified; source: subsequent coverage]]
3.4 Two specimens: an imitated fiction, and a named case in retreat
The Spotify model is, by its own authors' admission, an aspiration paper. Former Spotify agile coach Joakim Sundén: "Even at the time we wrote it, we weren't doing it. It was part ambition, part approximation." Co-author Ivarsson: "It worries me when people look at what we do and think it's a framework they can just copy and implement." Kniberg: it was never a general framework at all, just an example of one company's way of working. [verified] The org chart an entire industry copied was, at its point of origin, an ideal draft that was never implemented — a specimen of rare purity for Section 5's mimetic isomorphism.
Capital One, January 2023: roughly 1,100 agile roles eliminated (agile coaches, delivery leads, portfolio leads), spokesperson verbatim: "The agile role in our tech organization was critical to our earlier transformation phases but as our organization matured, the natural next step is to integrate agile delivery processes directly into our core engineering practices." [verified] The same sentence supports two readings — "agile graduated" or "the agile layer got cut" — and both readings are consistent with the archive. What is certain is that the dedicated agile caste is exiting the stage, with practitioner-side training enrollment and job-posting data pointing the same direction [unverified; source: Age of Product, BirJob compilations].
4. The DevOps archive: the most seriously measured wave, and its ceiling
The DevOps generation took a genuine step forward in evidence over the Agile generation — and precisely because of that, its archive exposes where this kind of measurement tops out.
The progress is real. DORA/State of DevOps measures outcome metrics (deployment frequency, change lead time, change failure rate, time to restore), not satisfaction; the four key metrics can be measured directly from system data (Google open-sourced the Four Keys tooling), so others can recompute them on their own data. Compared with "do you feel agile," that is a paradigm upgrade.
But its longitudinal narrative cannot carry its own method. In the 2022 report, the "elite" performance cluster vanished outright — official wording: "Unlike in years past, there was no evidence of an 'Elite' cluster." The low-performer group jumped from 7% in 2021 to 19%, and DORA's offered explanation was an untested pandemic hypothesis. String the elite shares across the years into a line (7% in 2018 → 20% in 2019 → 26% in 2021 → vanished in 2022 → about 18-19% in 2023-24) and the "industry is improving" curve turns out not to be comparable: DORA's own FAQ concedes that the clusters re-emerge each year from that year's different respondents — it is not a calibrated industry index. Add self-reported questionnaires, possible self-selection bias (people who consider themselves elite are more willing to answer; The Register called this out in 2021), unpublished raw data, and no independent replication of the capabilities→performance path model — the most seriously measured wave could only be this serious. [verified]
A forecast got treated as a measurement. The widely circulated "75% of DevOps initiatives will fall short of expectations due to organizational learning and change issues" (Gartner, 2019, forecast horizon 2022; a 90% version ran to 2023) is an analyst prediction, never revisited for validation. [unverified; source: Gartner 2019 and its social-media posts] It entered countless slide decks as "research shows" anyway.
The named mega-case: GE Digital. The loudest case in the transformation-narrative economy delivered the hardest closing numbers: GE announced in 2015 the target of over $15 billion in software and solutions revenue by 2020 (from about $5 billion at the time); reportedly over $4 billion went into GE Digital in 2016 alone (third-party estimate; GE never formally disclosed a figure on that basis); in December 2018, GE announced it was reorganizing the digital business as an independently operated but still wholly owned company, at which point annual software revenue was about $1.2 billion — 8% of the target — and it sold the majority stake in ServiceMax (bought in 2016 for $915 million) to Silver Lake (the spin-off itself was never completed either; GE Digital ultimately folded into GE Vernova). [verified; note that GE's 2018 goodwill write-down of $22 billion belonged to GE Power, not Digital — the two are frequently conflated] All the while, GE-family transformation stories were still touring industry summits as success cases. [unverified; source: DOES agenda archives]
A survivor economy. The DevOps case library is structurally tilted toward success: enterprise summit stages are built from self-nominated success stories, and the movement's founding text (The Phoenix Project) is itself a novel; there is no breakout session for failures. [unverified; source: DOES 2016 press release, IT Revolution case library] Nobody is cheating; this is how such archives get generated — remember the mechanism, because it applies unedited when you read AI case collections in Section 7.
And this wave's archive already contains AI's first line of record. DORA 2024: for every 25% increase in AI adoption, delivery throughput was estimated to drop 1.5% and delivery stability 7.2%; in the 2025 report throughput turned positive, stability remained negative, and the official framing shifted to amplifier: "AI doesn't fix a team; it amplifies what's already there." verified; wording carried over from [When Code Becomes Cheap] The measurement infrastructure the last transformation built is already issuing report cards to the next one — the most valuable continuity between the two archives.
Schematic: each wave measured harder than the last, and neither could produce an industry failure rate — sections 3-4 unpack each cell
5. The people selling the cure: change management's own evidence check-up
"70% will fail" is the scare; "follow our framework and you'll succeed" is the antidote. The scare has had its autopsy; now for the antidote.
Kotter's eight steps: never tested as a whole. Appelbaum et al. (2012) did the first systematic reconciliation 15 years after the model appeared: most individual steps find scattered support, but no formal study has ever tested the whole model; the model rests on Kotter's personal business and research experience and cites no external research; conclusion, verbatim: "Kotter's change management model appears to derive its popularity more from its direct and usable format than from any scientific consensus on the results." [verified] From then to now, neither we nor the verifiers have found a controlled test of the eight steps as a whole; the main empirical application (Pollack & Pollack 2015) found instead that in practice the steps run in parallel and loop back, defying the linear narrative [unverified; source: Systemic Practice and Action Research 28]. Kotter's 2014 response (Accelerate, the dual operating system) recast the eight steps as eight concurrent "accelerators" — conceding the linearity critique, still bringing no new evidence, and Kotter Inc. sells it as a product. [unverified; source: HBS Press]
Lewin's "unfreeze-change-refreeze": a model constructed after his death. Cummings, Bridgman and Brown (2016, Human Relations), comparing Lewin's original writings with later textbooks, argue verbatim: "we argue that he never developed such a model and it took form after his death" — assembled layer by layer, after his death in 1947, by Lippitt, Schein and the textbooks. [verified] Full disclosure: this claim has an opposing party — Burnes (2020) wrote a rebuttal arguing the three-step account is deeply rooted in Lewin's field theory and is no simplistic fabrication [verified (existence of the rebuttal)]. For this essay's argument, a draw is enough: lesson one of the change-management textbook is itself an active scholarly dispute, not a validated method.
There is exactly one systematic evidence reconciliation, and its conclusion is modest. Stouten, Rousseau and De Cremer (2018, Academy of Management Annals) checked seven popular models (Lewin, Beer, Judson, Kanter/Stein/Jick, Kotter, ADKAR, Appreciative Inquiry) against the academic evidence item by item: the models rest more on expert opinion than on scientific evidence; some signature prescriptions (such as "create urgency first" without diagnosis) lack support, while others (vision, communication, participation) partly align with the evidence; the authors closed by distilling their own evidence-based checklist of roughly ten steps. [verified]
Why it sells without evidence: legitimacy, not effectiveness. DiMaggio and Powell (1983) supply the theoretical mechanism, mimetic isomorphism: "When organizational technologies are poorly understood, when goals are ambiguous, or when the environment creates symbolic uncertainty, organizations may model themselves on other organizations." And: "The ubiquity of certain kinds of structural arrangements can more likely be credited to the universality of mimetic processes than to any concrete evidence that the adopted models enhance efficiency." [verified] The empirics caught up with the theory: Staw and Epstein (2000), testing the hundred largest US corporations — association with popular management techniques (TQM and kin) brought no economic performance gain, but brought higher reputation ratings and higher CEO pay; Westphal, Gulati and Shortell (1997; 2,700+ hospitals) — early adopters customized TQM for efficiency, late adopters copied the standard form for legitimacy, and the degree of copying correlated negatively with efficiency gains. [verified] (One boundary note, in fairness: later research shows earnest implementers do see returns — quality-award winners show better financial performance, and the meta-analytic average effect of isomorphism is not negative [unverified; source: Hendricks & Singhal 1997/2001, Heugens & Lander 2009]. What stands condemned is not the practice but the ritual adopted in order to look like one's peers.)
Management fashions have a measurable life cycle. Abrahamson's management-fashion theory and its empirical test with Fairchild: discourse volume for waves like quality circles surges and crashes in a bell curve; upswing-phase discourse is emotional, zealous and unqualified, with sober hedging returning only on the way down — a pattern they call "superstitious" collective learning. [unverified; source: AMR 1996, ASQ 1999] Keep that curve in mind: Section 7 lays it against AI.
This section's combined verdict: the change industry's scare number has no empirical source, its antidote frameworks have never been tested whole, and the buying behavior is driven by legitimacy — thirty years of selling smoke, then? No. What was sold really did do something; the mechanism was just Staw-Epstein: adopters got legitimacy, prestige and pay, and the consultants got revenue. The only thing never demonstrated is any effect on that denominator behind "70%."
6. The "frozen middle" retried: who actually kills transformations
Every transformation story has an official villain: the middle manager — "the top wants change, the bottom wants change, the middle is frozen." This "frozen middle" deserves its own hearing, because it is being carried, sentence intact, into the AI narrative.
The pedigree is suspect from birth. The phrase is usually attributed to 1980s GM CEO Roger Smith, but all that can be located is consulting blogs citing each other (even Dartmouth Tuck's case page does it) — no primary 1980s source whatsoever (speech, interview, contemporary reporting) can be pinned down. [verified (as an absence)] A concept used to explain transformation failure whose own provenance is a chain of retellings — the same species as "70%."
The academic archive exonerates the middle — conditionally. Wooldridge and Floyd (1990, a quantitative study of 20 organizations): middle-manager involvement in strategy formation correlates positively with organizational performance. Huy (2002, three-year field study, ASQ): in radical change, middle managers' "emotional balancing" work — pushing projects with one hand while catching subordinates' emotions with the other — is the mechanism by which adaptation happens at all. Balogun and Johnson (2004, AMJ): change derails mainly because, after top management withdraws, middle managers are left to their own sensemaking in a communication vacuum — not because of deliberate resistance. The nail on the other side is academic too: Guth and MacMillan (1986) confirmed that when middle managers judge a change to harm their own interests, they can slow a strategy, degrade it, or "totally sabotage" it. [verified] Taken together: the middle layer is a conditional reactor, not permafrost — the conditions being incentive alignment, trust and participation, and all three of those dials sit in the C-suite's hands.
The largest practitioner dataset points the finger upstairs. Prosci's biennial survey, running since 1998 (vendor self-published, practitioner sample — both into the record): "active and visible executive sponsorship" has ranked as the number-one contributor to change success in every single edition, with a mention rate roughly three times the runner-up; "engagement with middle managers" ranks seventh of seven. With extremely effective sponsorship, projects meet objectives at roughly 3.5 times the rate seen under extremely ineffective sponsorship (that multiplier is the current edition's; other editions run 2.5-2.9x). The same material also honestly records that middle managers are the most resistant group in its surveys — consistent with the "conditional reactor" reading: the resistance is measurable, but it is the dependent variable. [verified] A converging line of evidence comes from the negative side: MIT Sloan's Johnson, summarizing the sampling bias of change research — most studies cover only the first few months and interview only the leadership; the evidence for "change failed, blame the lazy managers" is manufactured exactly that way. [unverified; source: strategy+business 2020]
Even "participation," the movement's proudest brand, is conditional. Change management's genesis experiment — Coch and French (1948, the Harwood pajama factory) — is taught in textbooks as "participation removes resistance": the no-participation group's output fell from about 60 units/hour to about 50 (roughly 17-20%) with no recovery over 32 days; the full-participation groups recovered and exceeded pre-change levels by about 14%. But the original experiment's control group had only 18 people, the experimental groups 13/8/7; 17% of the control group quit within the first 40 days; Bartlem and Locke (1981) pointed out that explanation, training, job availability and piece-rate fairness were all confounds; and the Norwegian replication (French, Israel and Ås, 1960) detected no output effect. [verified] Participation→commitment is a conditional heuristic, not a law — a direct calibration for the AI mandate-versus-bottom-up dispute below.
The frozen-middle narrative's new job in the AI era — and what the data says. "AI frozen middle" articles appeared in bulk across 2025-2026, but the reports they cite frequently say the opposite: McKinsey's Superagency (2025) concludes that employees are ready and leadership is the biggest bottleneck; Kyndryl's 2025 survey finding that "45% of CEOs think their employees are resistant to AI" is CEOs' perception of employees at large, retrofitted by blogs into "the middle is blocking AI." [unverified; source: the original reports and their re-citations] And the data that directly measures usage puts the gradient the opposite way entirely — leaders > middle managers > frontline (Gallup, Q4 2025: using AI a few times a week or more, leaders about 44%, managers about 30%, frontline about 23%). [verified] The middle is not AI's permafrost; if there is permafrost, it is elsewhere (next section).
Change fatigue is real; its numbers are a mess. The employee-side archive: the average employee experienced 10 planned enterprise changes in 2022, versus just 2 in 2016 (HBR 2023, Gartner authors). [verified] The fill-in-the-blank — "willingness to support change fell from 74% in 2016 to __% in 2022" — has been given two answers by Gartner's own publications: 43% (October 2022 release materials) and 38% (a Q1 2023 periodical). This essay initially ruled the 43% a transmission error; the verifiers corrected us: the discrepancy lives inside Gartner itself, coexisting with a coincidentally same-numbered "intent to stay 43%/74%" pair of statistics. [verified] A statistic about change wearing people out has worn itself out into two versions — that is not a quip; it is a free teaching aid for this essay's methodology.
7. Assessing predictive power: what this autopsy says about AI-native
Now close the files and answer the title question. List the failure mechanisms of the first six sections and check them against the 2024-2026 AI adoption record, and you get a sheet that is mostly checkmarks, with three fractures.
7.1 Mechanisms replaying, unmodified
Mimetic adoption — this time with measured numbers. The IBM/Oxford Economics survey of 2,000 CEOs (2025): 64% admit that the risk of falling behind drives them to invest in some technologies before they have worked out their value; CEOs self-report that only about 25% of AI initiatives have delivered the expected ROI. [verified] This is DiMaggio-Powell's 1983 sentence coming back as a survey echo — "ambiguous goals + symbolic uncertainty → imitation." The accompanying legitimacy theater is complete too: S&P 500 earnings-call AI mentions setting fresh ten-year records [unverified; source: FactSet], board-pressure surveys arriving in dense formation [unverified; source: Dataiku/Harris, BCG press releases].
Mandates and metrics: the Goodhart script has already finished its premiere. Shopify (April 2025, the Lütke memo): "Reflexive AI usage is now a baseline expectation" — AI use enters performance and peer reviews, and "Before asking for more headcount and resources, teams must demonstrate why they cannot get what they want done using AI". Microsoft's developer division (June 2025, the Liuson internal message, as reported by Business Insider): AI use is "no longer optional" and enters performance reflections. Meta (announced November 2025): from 2026, "AI-driven impact" becomes a core performance expectation for all employees. Coinbase's Armstrong, by his own account, fired engineers who had not gotten hands-on with AI coding tools by the deadline without a legitimate reason. And Duolingo ran the entire cycle in roughly 12 months: April 2025 "AI-first" memo, AI use into performance reviews → public softening within weeks → April 2026, von Ahn confirms the AI-use performance requirement has been withdrawn, verbatim (Silicon Valley Girl podcast, as reported by Fortune): "At the end, we backtracked... the most important thing in your performance is that you are doing whatever your job is as well as possible." [verified] In the same year one company retreated from "grade AI usage" back to "grade the work itself," another institutionalized the former — Section 3's ritual adoption and Section 6's mandate-versus-participation lesson playing out in a single news cycle. Engineers gaming token dashboards and clicking accept before rewriting: practitioners have already documented the specific coping moves. [unverified; source: Patrick God 2026, among others]
The ratio of ritual adoption to substantive redesign has been measured. McKinsey's March 2025 survey: among organizations using generative AI, only 21% have fundamentally redesigned even some workflows — while of the 25 attributes it tested, workflow redesign showed the largest correlation with EBIT impact (self-reported, correlational design). [verified] 79% of "adoption" is a new tool wedged under old processes — exactly Westphal-style ritual adoption, and DORA's 2024-2025 amplifier readings (throughput turned positive, stability persistently negative) annotate it from the other side.
The antidote-selling layer has already reorganized, at unprecedented scale. Accenture's GenAI new bookings: $3.0 billion in FY2024, $5.9 billion in FY2025 (SEC filings). BCG: AI-related business at about 20% of a record $13.5 billion in 2024 revenue. McKinsey leadership (as reported by WSJ and others): AI-related work at about 40% of the firm's business, led by the roughly thousand-person QuantumBlack. The certification industry is re-tracing the SAFe curve on schedule: AI certifications now account for nearly 30% of certifications listed on professional profiles, about 20 times the pre-ChatGPT level (Revelio Labs). [verified] Add the Chief AI Officer boom and the "AI Center of Excellence" returning as a standard commodity [unverified; source: IEEE-USA, various consulting sites]. Same firms, same product shape (transformation program + certification + CoE), new noun.
7.2 Three fractures: where this round is different
The direction of adoption has reversed: from "ritual pressed down" to "shadow welling up." Agile's classic pathology: the top bought the ritual, and the front line performed it for show. AI's first large-sample datasets draw the inverse picture: 75% of knowledge workers already use AI at work, and of those, 78% bring their own tools (BYOAI; Microsoft/LinkedIn 2024, 31,000 people across 31 countries); 57% of employed AI users admit to non-transparent use — including passing off AI output as their own or avoiding disclosure of AI use (KPMG × University of Melbourne 2025, 48,000+ people across 47 countries; combined measure of the two behaviors). [verified] Agile-era employees performed usage; AI-era employees conceal usage — the organizational disease is the same (ritual decoupled from substance), but the symptom has flipped sign. That changes the prescription directly: agile transformation had to solve "how do we get people to actually use it"; AI transformation has to solve "how do we get the people already using it to say so — and to use it well."
The thing being adopted is improving — measurably, exponentially. Agile's rituals went twenty years without changing; the capability of the AI tools being mandated is climbing a measured curve — METR: the length of task frontier models can complete autonomously at 50% reliability has doubled roughly every 7 months over the past six years. [verified] That breaks a premise implicit in the post-mortems: last round, "transformation failure" was essentially "organizational failure," because the method itself was a constant; this round, the organization can bungle its transformation while the technology grows good enough on its own — or do everything right and have its premises rewritten by the next model generation. Management-fashion theory's bell-shaped decay curve (Section 5) is meeting, for the first time, a fashion whose host's intrinsic capability keeps rising — the old curve may not apply.
The supply-side economics have partly changed. Agile's money sat mostly in hourly-billed transformation services and certification; AI's tool layer is per-seat SaaS, self-serve, with vendor revenue tied to usage rather than to "transformation timelines" — OpenAI reached 3 million paying business users by mid-2025 [unverified; source: CNBC]. But this only counts as half a fracture: Accenture's $5.9 billion above shows the hourly-billed transformation layer is booming in parallel; the two layers coexist.
7.3 Closing the book
The last wave's post-mortems' predictive power for AI-native compresses into three sentences:
At the level of organizational mechanisms, predictive power is high. Mimetic adoption, definition-driven scare numbers, ritual substituting for substance, the Goodhart backlash to mandates, the misdirected blame-the-middle narrative — all five have already recurred in the AI adoption record, most of them with quantitative evidence. The shape of failure will rhyme.
At the level of the failure rate, predictive power is zero — because that number never existed. "70% of AI transformations fail" was circulating before anyone had measured it, and that is the whole point of this essay's archaeology: it is a rhetorical device, not a measurement. On the same grounds, be pre-armed against the next generation of candidate zombie numbers, "95% of AI pilots fail" and kin (that's another essay's topic).
At the level of technical dynamics, the analogy fractures, direction unresolved. Bottom-up shadow adoption and exponentially improving tools are variables absent from the last archive; they could make this round genuinely different, or merely relocate the same organizational diseases to a new lesion. The instrument for telling the two apart happens to be the last wave's finest legacy: outcome measurement. When Code Becomes Cheap's prescription closes its loop here — keep your transformation's books with delivery data, not gut feel and questionnaires; DORA-style dashboards are already grading AI. Don't let the next thirty years navigate by a made-up percentage.
Schematic: the organizational machinery replays (left) while three structural conditions have no precedent in the last archive (right) — section 7 unpacks each
8. Coda: eleven testable claims
Ordered by evidence strength:
"70% of transformations fail" has no empirical source: it was born in 1993 as an "unscientific estimate," retracted by its author in 1995, rewritten in 2000 as a sourceless "brutal fact," and fitted in 2008-09 with the invented provenance of "Kotter's 1995 research." (strong: full chain checked word for word, adversarially verified)
Kotter's famous 1995 article contains no overall failure rate of any kind; he first personally gave 70% in 2008, and called it an estimate. (strong: primary-text check; the only percentage is the stage-level observation that far more than 50% fail at step one)
Hughes's 2011 academic trace holds — "no valid and reliable empirical evidence" behind the five published instances — and the number kept circulating after its autopsy, including McKinsey 2015/2021 and the 2024-26 AI variants. (strong)
The measurable true base rate is "undetermined": the academic range is 7-90% with a mean around 50%, on evidence that is "outdated, fragmentary, fragile or just absent"; the failure rate is a function of the definition — the same dataset can produce "88% fail" and "87% achieved at least half." (strong: Cândido & Santos; BCG's and Bain's own splits)
The measured distribution of software projects is fat-tailed, not "mostly failing": average overruns of 27-73%, but the one-in-six black swans average +200%, and the IT projects overrunning by more than 50% average +447%. (strong: two generations of Flyvbjerg data; the baseline definition is academically contested)
The CHAOS report's figures are an artifact of forecasting bias: flip the direction of an organization's estimation bias and its "success rate" swings 5.8% ↔ 94.2%; Standish's chairman himself calls the reports "opinion." (strong: the Eveleens & Verhoef recomputation plus the on-the-record exchange)
The Agile archive's self-diagnosis went unchanged for seventeen years — culture and leadership were always the top self-reported obstacles — while the scaling-framework industry produced not one independent controlled outcome study across 2016-2025; the largest independent comparison measured the practical difference of framework choice as negligible. (medium-strong: negative signals from vendor surveys + Dikert + Verwijs & Russo, all self-reported)
Agile practices themselves correlate with better project outcomes (all size bands; significant for small and medium) — the failure archive belongs to the "transformation ritual," not to the practices. (medium: single-country sample, self-rated, correlational)
The change industry's antidote frameworks have never been tested whole: Kotter's eight steps have no whole-model controlled study, Lewin's three-step model is a posthumous construction (a claim with an academic opposing party), the one systematic reconciliation (Stouten 2018) judges the foundations to be expert opinion; and the empirical driver of adoption is legitimacy and CEO pay, not performance. (medium-strong: Appelbaum / Cummings [contested by Burnes] / Stouten / Staw & Epstein)
The "frozen middle" has no traceable source, and the empirical direction runs against it: middle-manager involvement correlates positively with performance, blocking behavior is a conditional response to misaligned incentives; the largest practitioner dataset gives the top factor to executive sponsorship (middle-manager engagement ranks seventh); and the measured AI-usage gradient — leaders 44% > managers 30% > frontline 23% — points the opposite way from the frozen-middle narrative. (medium: conditional academic evidence + Prosci vendor self-report + Gallup)
AI transformation is replaying the last round's organizational mechanisms (64% of CEOs self-report investing before understanding, only 21% have redesigned workflows, Duolingo has run the full mandate→withdrawal cycle, and the consulting/certification layer is reorganizing at multi-billion-dollar scale), while diverging structurally in two places: shadow bottom-up adoption (78% BYOAI, 57% concealment) and tool capability doubling every 7 months. (medium-strong: the first-hand vendor/survey numbers are verified, but all are fast-moving variables, as of July 2026)
Signals worth watching: whether Meta's AI performance metric survives 2026 (the Duolingo script predicts it retreats); where McKinsey's "21% redesigned workflows" moves in the 2026-27 surveys (up = substance beginning, flat = the ritual phase extended); the sign of DORA 2026's AI-stability reading; whether the scaled-agile-framework industry produces its first independent controlled study in the AI era (base rate: nine years, zero); and whether the citation chain of the next zombie number — "95% of AI pilots fail" — re-runs Section 1 of this essay (we intend to give it an autopsy too, when the time comes). Thirty years ago, a self-described unscientific estimate got promoted to an industry's navigation system; this round's organizations have something that era lacked — their own delivery data. Use it.
Appendix: Primary Sources
Number archaeology: Hammer & Champy, Reengineering the Corporation (1993) · Hammer & Stanton, The Reengineering Revolution (1995) · Kotter, "Leading Change: Why Transformation Efforts Fail" (HBR, 1995) · Beer & Nohria, "Cracking the Code of Change" (HBR, 2000) · Kotter, A Sense of Urgency (2008) · Keller & Aiken, "The Inconvenient Truth About Change Management" (McKinsey, 2009) · McKinsey Global Survey, "Creating Organizational Transformations" (2008-07) · Hughes, "Do 70 per cent of all organizational change initiatives really fail?" (J. Change Management 11(4), 2011, DOI 10.1080/14697017.2011.630506) · Ewenstein, Smith & Sologar, "Changing Change Management" (McKinsey, 2015) · McKinsey, "Losing from day one" (2021-12) · Greenberg, "How citation distortions create unfounded authority" (BMJ 2009;339:b2680) · Tourish, Management Studies in Crisis (CUP, 2019)
Base-rate measurement: Standish CHAOS Reports (1994/2015/2020) · Eveleens & Verhoef, "The Rise and Fall of the Chaos Report Figures" (IEEE Software 27(1), 2010) · Jørgensen & Moløkken, "How large are software cost overruns?" (IST 48(4), 2006) · Flyvbjerg & Budzier, "Why Your IT Project May Be Riskier Than You Think" (HBR, 2011; arXiv:1304.0265) · Flyvbjerg & Gardner, How Big Things Get Done (2023) · Smith, "Success rates for different types of organizational change" (Performance Improvement 41(1), 2002) · Cândido & Santos, "Strategy implementation: What is the failure rate?" (JMO 21(2), 2015) · BCG, "Flipping the Odds of Digital Transformation Success" (2020-10) · Bain, 88% press release (2024-04) and Mankins & Litre, "Transformations That Work" (HBR, 2024) · Sauer, Gemino & Reich (CACM 50(11), 2007) · Loureiro et al. (Heliyon, 2024) · Ika & Pinto, "The re-meaning of project success" (IJPM 40(7), 2022)
The Agile archive: Digital.ai/VersionOne, State of Agile Reports, editions 14-18 · Dikert, Paasivaara & Lassenius (JSS 119, 2016) · Jørgensen, "Relationships Between Project Size, Agile Practices, and Successful Software Development" (IEEE Software 36(2), 2019) · Scaled Agile, framework.scaledagile.com/about and scaledagile.com marketing pages · Putta, Paasivaara & Lassenius (XP/PROFES 2018) · Verwijs & Russo, "Do Agile Scaling Approaches Make A Difference?" (EMSE, 2024; arXiv:2310.06599) · USAF CSO Chaillan, Memorandum for Record on Agile Frameworks (2019-12-28) · Jeremiah Lee, "Spotify's Failed #SquadGoals" (2020) · Kniberg (blog.crisp.se, 2015) · Banking Dive, Capital One coverage (2023-01) · Fowler, "FlaccidScrum" (2009) · Dave Thomas, "Agile is Dead (Long Live Agility)" (2014)
The DevOps archive: DORA/Google Cloud, State of DevOps / DORA Reports (2018-2025) and dora.dev/faq · Forsgren, Humble & Kim, Accelerate (2018) · Sallin et al. (XP 2021) · The Register (2021-09) · Gartner, "The Secret to DevOps Success" (2019) · GE press releases (2015-09-29, 2018-12-13) · The Conversation, GE Digital analysis (2018) · Keunwoo Lee, review of Accelerate and Humble's response
The AI adoption record: Lütke memo (X, 2025-04-07) · Fortune, coverage of von Ahn's podcast remarks (2026-04-13) · Business Insider/Entrepreneur, coverage of the Liuson internal message (2025-06) · HR Grapevine, Meta performance policy (2025-11) · TechCrunch, Coinbase (2025-08) · McKinsey, "The State of AI" (2025-03) · IBM/Oxford Economics CEO survey (2025-05) · Accenture 8-K (SEC, FY2024/FY2025) · BCG annual results press release (2025) · WSJ/Alex Singla, QuantumBlack (2025) · Revelio Labs, AI certification analysis (2026) · Microsoft/LinkedIn Work Trend Index (2024) · KPMG × University of Melbourne, "Trust, attitudes and use of AI" (2025) · Gallup, Q4 2025 workplace AI use (gallup.com/workplace/701195) · METR, "Measuring AI Ability to Complete Long Tasks" (arXiv:2503.14499) · DORA 2024/2025 (wording carried over from this site's When Code Becomes Cheap verified record)
The research materials and all verification rulings live in the research base (7 research threads; 30 load-bearing claims × 3 adversarial verification votes: 30/30 survived, 20 with wording calibrations, 0 overturned).