中文EN
← Deep Research
Deep Research · Plain-language

'70% of Transformations Fail'? That Number Was Made Up (Plain-Language Edition)

This is the plain-language edition · read the deep dive (full arguments & sources) →
TL;DR
Nobody ever measured '70% of transformations fail': it was born an 'unscientific estimate', retracted by its authors, upgraded to a 'fact', given a fake pedigree, academically debunked in 2011 — and kept circulating right into the AI era. The real failure rate depends on where you draw the pass line: the same data can read '88% fail' or '87% got at least halfway'; the measured danger is a fat tail of disasters, not majority failure. The last wave's lesson: buy practices, not rituals. This wave's new variables: employees using AI in secret, and tools that keep getting stronger.
Kotter 1995: no 70% in the textsame data: 88% fail or 87% halfwayonly 21% redesigned workflows78% bring their own AI

This is the condensed edition of the deep dive of the same name. All key figures were independently verified; for the full argument and sources, read the deep dive.

A number that's been a calling card for thirty years

You've probably heard the line: "Research shows that 70% of transformations fail." It shows up in management books, in consulting proposals, in leaders' rally-the-troops speeches. Now it has a new job: "70% of AI transformations fail — but we know how to make you one of the 30%."

Scare you first, then sell you the cure. The formula hasn't changed in thirty years. This article does one simple thing: it goes and checks who actually measured that number.

Here's the answer up front: nobody ever did. What follows is its résumé.

The life of a number

1993: it is born, with "unscientific" written on the birth certificate. The two authors of Reengineering the Corporation estimated that among companies attempting process reengineering, 50% to 70% "do not achieve the dramatic results they intended." They said so themselves: this was an "unscientific estimate," it covered process reengineering only, and "failure" meant "fell short of the ambitious goals they had set" — not "made a mess of things."

1995: the author takes it back in person. Lead author Hammer wrote, specifically, that this descriptive observation had been widely twisted into a law — reengineering has no inherent success rate or failure rate. Almost nobody has ever cited the retraction. The number lived on; the disclaimer died.

2000: it mutates. Two Harvard professors wrote in Harvard Business Review: "The brutal fact is that about 70% of all change initiatives fail." No footnote, no study, no source. "Process reengineering" had quietly become "all change," and an "unscientific estimate" had been promoted to "brutal fact."

2008-2009: it acquires a fake pedigree. A McKinsey article wrote: "In 1995, John Kotter published research that revealed only 30 percent of change programs are successful." We checked Kotter's 1995 article from start to finish: the number is not in it — there is no overall failure rate of any kind. All he offered was an impression: most change efforts land somewhere between great success and utter failure, "with a distinct tilt toward the lower end of the scale." The "published research that revealed" was invented out of thin air. (Kotter himself first said 70% in 2008, in his own book — prefaced with "I estimate.")

2011: scholars perform the autopsy. The British scholar Hughes traced every published 70% claim back to its source. The conclusion: no valid and reliable empirical evidence supports the claim. That is a formal, peer-reviewed verdict.

After 2011: it becomes a zombie. McKinsey kept writing "70% of change programs fail" in 2015, with no source; in 2021 it went ahead and cited Kotter's 2008 book as backing — and the "research pedigree" behind Kotter's number was the one McKinsey itself had invented years earlier. The citation chain loops back into a circle. Then, starting in 2024, "70% of AI transformations fail" hit the market in bulk.

The life of '70%': circulation above, debunking below ● top: the number cited / strengthened ● bottom: retraction & autopsy (rarely cited) 1993 2009 2026 1993: born 'unscientific', reengineering only 2000: upgraded to 'brutal fact', unsourced 2008-09: Kotter 'I estimate' + McKinsey invents source 2024: new host, AI transformation 1995: authors retract, 'no inherent failure rate' 2011: Hughes autopsy, 'no valid evidence'
Schematic: every mutation got louder (top track) while both corrections went uncited (bottom) — section 1 unpacks each dot

So what is the real failure rate?

The honest answer comes in three layers.

Layer one: the most famous "measurement" doesn't survive inspection. The software industry's "most projects fail" lore comes from Standish's CHAOS reports (1994: only 16% of projects "succeeded"). But its definition of success was on time, on budget, delivering all planned features — hit all three or you lose. Two Dutch researchers redid the math on 1,211 real projects and found the method actually measures how conservatively your company writes its budgets: for the same company, flip the estimating habits and the "success rate" swings from 6% to 94%. Confronted directly, the Standish chairman's answer was that all the data in its reports "should be considered Standish opinion and the reader bears all risk in the use of this opinion."

Layer two: the people who measured carefully found "a few disasters," not "most fail." The Oxford scholar Flyvbjerg's database holds more than ten thousand projects. Average overruns are actually manageable (IT projects average 73% over budget, measured from the point of approval); the truly scary part is that one project in six is a "black swan" — averaging 200% over budget. It's like a health checkup: most people's numbers are just a bit elevated, and the danger sits with a few — it is not "70% of everyone is critical."

Layer three: academia's formal answer is "this question has no single answer." A systematic review gathered every published failure-rate estimate: they run from 7% to 90%. Why so scattered? Because "failure" is a word you can define however you like. Two live examples:

Remember the trick: set the passing bar at "perfect," and the failure rate can be as high as you want. The organizations selling transformation services happen to be the same ones holding the power to set the bar.

The definition machine: one dataset, two report cards BCG 2020: digital transformations (825 executives, self-graded) 30% met targets 44% created value, missed targets 26% limited value → packaged as: '70% fall short' Bain 2024: business transformations (400+ executives, self-graded) 12% ~75% got at least halfway, short of full ambition ~13% → packaged as: '88% fail' (same data: ~87% achieved at least half)
Schematic: set the pass bar at 'perfect' and the failure rate is whatever you need; segments are each report's own breakdown — section 2.4 unpacks both

What the last round of transformations left in the file

Before AI, the software industry had just been through two industry-wide transformations: Agile and DevOps. Their files are the most useful thing we have today.

Agile's file: a questionnaire filled out for seventeen years. The industry runs an annual State of Agile survey (vendor-run and self-selected — remember both). The most valuable thing it ever measured: "culture and leadership" has been named the biggest obstacle for well over a decade running — a movement that sold itself on culture change, whose own survey says year after year that the culture didn't change. In recent years the survey itself has been ebbing: satisfaction fell from 71% to 59%, the latest edition drew only 349 respondents, and 24% of companies are cutting their agile investment.

The framework business's evidence vacuum. SAFe, the best-selling scaling framework, advertises that "Seventy percent of Fortune 100 companies... have certified SAFe professionals" — note: "have employees who passed the exam," not "are using it." In nine years, not one independent controlled study has shown it works; the largest independent comparison (about 15,000 people) concluded that which framework you pick makes almost no practical difference to team effectiveness — what matters is the team's experience with agile. A 2019 US Air Force memo put it flatly: rigid frameworks like SAFe are strongly discouraged.

The homework the whole industry copied — the authors say it was made up. The "Spotify model," imitated by countless companies, was later disowned by Spotify's own agile coach: when they wrote it, they weren't even working that way themselves — it was half vision, half approximation.

The exits have started too. In January 2023, Capital One cut about 1,100 agile roles (agile coaches, delivery leads). The official line: the organization had matured, and agile should fold into engineering practice. You can read that as graduation or as layoffs; what's certain is that the full-time agile layer is disappearing.

To be fair: agile practices themselves may well be good. A Norwegian study of 196 projects found that projects using agile methods really did turn out better. The failure file belongs to "agile transformation," the organizational ritual — not to iterative delivery, the practice. Paying for the box and throwing away the pearl inside: that was the last round's biggest lesson.

DevOps's file: measured the most seriously — and it measured a ceiling. The DORA reports are a step up from opinion surveys — they measure outcome metrics like deployment frequency and recovery from failure. But it had an accident in 2022: the "elite" cluster, present every prior year, vanished entirely, and the official explanation was an unverified pandemic guess. Its own FAQ admits that each year's clusters are recomputed from that year's different respondents — year-to-year comparisons don't hold. The loudest success story collapsed too: GE announced in 2015 that it would build a $15 billion software business, poured in billions of dollars, and when it announced at the end of 2018 that the digital unit would become an independently run but still wholly GE-owned company, annual revenue was $1.2 billion — 8% of the target.

And one more thing needs saying out loud: nearly every transformation success story you've heard comes from winners volunteering to get on stage. The companies that failed don't host talks.

The "frozen middle" is a scapegoat

Every transformation story has an official villain: middle managers — "the top wants change, the bottom wants change, the middle is frozen solid."

First, check the pedigree: the line is usually attributed to a General Motors CEO in the 1980s, but no primary source from the time can be found — only consulting blogs citing one another. Same species as the "70%."

Then check the evidence, which mostly points the other way: research finds middle managers' involvement in strategy is positively correlated with performance; the common reason change goes sideways is that top leadership pulls back and middle managers are left guessing in an information vacuum — not deliberate resistance; and middle managers do block change — when the change hurts their interests. In other words, the middle is a thermometer, not permafrost: the three dials — incentives, trust, participation — all sit upstairs. The largest practitioner survey points the same way: the number-one factor in successful change is "active and visible executive sponsorship," mentioned about three times as often as the runner-up; "engagement with middle managers" ranks seventh.

One honest aside: even the golden rule that "involve employees and they won't resist" rests on an original experiment with a few dozen participants — and it failed to replicate in another country. It's a rule of thumb, not a law.

On the employee side, there is one solid number: in 2022 the average employee went through 10 company-level changes; in 2016 it was 2. Change fatigue is real — the last round's ritual abuse has already spent this round's patience in advance.

AI transformation: how much of the script is replaying?

Hold the file above against the 2024-2026 record of AI adoption: most of the machinery is replaying as-is, and three things are genuinely different.

Replaying:

Genuinely different:

Last wave's script vs this wave's reality Mechanisms replaying Conditions that broke Herd: 64% of CEOs invest before seeing valueMandates: Duolingo, mandate→retreat in 12 monthsRitual: only 21% redesigned workflowsCert boom: AI certs near 30%, ~20x pre-ChatGPTInverted: 78% BYOAI, 57% hide their useCapability doubles ~every 7 monthsSeat-based SaaS layer (half a break)
Schematic: the organizational machinery replays (left) while three structural conditions have no precedent in the last archive (right) — section 7 unpacks each

How to Check This Article's Claims

From hardest evidence to softest:

  1. "70% of transformations fail" has no empirical source — born in 1993 as an "unscientific estimate," retracted by its author in 1995, rewritten in 2000 as a sourceless "fact," and fitted with a fake research pedigree in 2008-09. Every link can be checked against the original texts.
  2. Kotter's 1995 article, the most commonly named source, does not contain the number at all. He first said 70% in 2008, calling it "I estimate."
  3. Academia performed the formal autopsy in 2011 (verdict: no valid and reliable evidence), and the number circulates to this day, AI version included.
  4. The real failure rate has no single answer: published estimates run from 7% to 90%; the same data can be packaged as "88% failed" or as "87% got at least halfway" — the number is decided by whoever sets the passing bar.
  5. Measured software projects show "a few disasters," not "most fail": average overruns are manageable, and one project in six is a black swan averaging 200% over budget.
  6. The agile survey named "culture/leadership" the biggest obstacle for seventeen years straight; scaling frameworks went nine years with zero independent controlled studies; the largest independent comparison found framework choice makes no practical difference.
  7. Agile practices themselves correlate with better outcomes — what died was the transformation ritual, not the practices.
  8. The "frozen middle" has no traceable source, and the evidence points the other way: the top success factor is executive sponsorship; measured AI usage runs leaders 44% > managers 30% > frontline 23% — the frozen-middle story gets even the direction wrong.
  9. AI transformation is replaying the last round's organizational machinery (fear-driven investment, metric-based mandates, ritual adoption, a certification industry), while genuinely differing in two places: employees hiding their use (78% bring their own tools, 57% conceal it) and tool capability doubling every 7 months.

What to watch: whether Meta's AI performance metric survives 2026 (on the Duolingo script, it retreats); whether the "21% redesigned workflows" number rises in next year's survey; and whether the next candidate zombie number — "95% of AI pilots fail" — walks the same road as the "70%."

The Things That Matter Most

  1. When you hear "research shows X% fail," ask three questions: Who measured it? How was "failure" defined? What is the speaker selling? Thirty years of the "70%"'s résumé show that these three questions almost always dismantle the scare opener.
  2. A failure rate is not a constant of nature; it is a function of the definition. Set the passing bar at "perfect" and the failure rate can be whatever you want; the danger that actually shows up in measurement is the fat tail — a few disasters, not majority failure.
  3. Buy the practices, not the rituals. Last round, practices like iterative delivery worked, while "transformation frameworks" produced not one piece of independent controlled evidence in nine years. This round, spend the budget on redesigning workflows (only 21% of organizations have) — not on renaming things and issuing certificates.
  4. Don't force people to use AI; first make it safe to admit they already do. Your employees are most likely already using it (78% bring their own tools) — swap "grade the usage" for "show them safe, compliant ways to use it." Duolingo has already stepped on this rake for you.
  5. When a transformation fails, the rot starts upstairs. The evidence gives the top factor to executive sponsorship and incentive design, not middle-manager attitude. Before cursing the "frozen middle," check whether your own sponsorship has gone offline.
  6. Keep score with your own delivery data. The best legacy of the last round is outcome measurement (deployment frequency, failure rate, recovery time). Whether your AI transformation is going well, your dashboard will tell you — stop letting a made-up percentage do the navigating.

For the full argument, sources, and counter-evidence behind every conclusion, read the deep dive.