The more a claim is cited, the more credible it looks. But here is the counterintuitive fact: the number of citations and the actual evidence behind a claim can be completely decoupled. Authority is often something the citation network grows on its own — sustained not by how much data supports it, but by how many people have repeated it. You think you're trusting the evidence; you're really trusting the density of the retelling.
Three channels can amplify razor-thin evidence into what looks like rock-solid consensus. Bias: cite the supporting studies, ignore the contradicting ones, so the literature looks one-sided. Amplification: reviews cite reviews, layer upon layer, and no one goes back to read the original data — second-hand becomes first-hand. Invention: cite a paper that contains no such data, or even contradicts the claim, as "support," turning what was a hypothesis into settled fact. Stack the three, and a mere conjecture can acquire near-axiomatic status in the literature.
A 2009 study in the British Medical Journal performed a "citation-network autopsy" on a widely accepted claim in biomedicine: 242 papers and 675 citations woven into one web. Trace the citations back to their origin, and only a tiny handful of papers actually contained supporting experimental data; a large share of the "supporting" citations either held no data at all or pointed the opposite way. The entire consensus was held up almost entirely by citations propping up other citations — not by the evidence underneath. Numeric authority can be hollow.
Software dependencies are isomorphic: when you trust a third-party package, you're really trusting what it imports, and what that imports — any hollow link in the transitive trust chain is invisible to you. In machine learning, training a model on model-generated data lets errors reinforce themselves through "self-citation," eventually growing into a confident hallucination. Even a stretch of legacy code no one dares touch often draws its "authority" not from having been verified, but from having been depended on so long that no one goes back to check whether it's right.
The most dangerous sentence in a technical decision is "the industry best practice is to do it this way." Next time you hear "everyone does it like this" or "community consensus," dig one layer down: is that consensus backed by a few first-hand benchmarks and one real load test, or by a ring of blog posts citing each other, none of whose authors ever ran the numbers? Keep "how many people say it" and "how much evidence there is" on separate ledgers.
Take one technical tenet your team now treats as default. If you were forced to defend it using only first-hand evidence — not "others say so too" — how many cards would you still hold?
A number turns from "someone's rough guess" into "an established fact" not because anyone verified it, but because every retelling quietly strips away its qualifiers — the uncertainty, the definition, the source. Qualifiers are a number's immune system; strip them off, and it becomes an invulnerable "fact."
It's a four-step decay chain, and every step makes the number harder and falser. First someone offers a rough estimate, hedged with caveats. In retelling, the "roughly" and "unscientific estimate" get deleted — factualization. A naked number needs authority, so it's assigned a source that never actually said it. Finally, even when someone performs an "autopsy" and disproves it, the number has long since broken free of the original text and circulates on its own — the autopsy report can never catch up.
"70% of change efforts fail" — one of the most-quoted mantras in management. Trace it to its nearest source, and you find a self-limiting line in a 1993 business-reengineering classic: "Our unscientific estimate is that as many as 50 to 70 percent of reengineering efforts fail to achieve the dramatic results they intended." Note two things: the authors explicitly call it unscientific, and they say "fail to achieve dramatic results" — not the same as "fail." Over three decades, "unscientific" was cut, "dramatic results" was swapped for "fail," and fabricated attributions to authoritative studies were bolted on. In 2011 a scholar published a formal "autopsy" in a journal, concluding the number rests on no reliable evidence whatsoever — yet it still opens countless slide decks today.
This is lossy compression. Information theory says every retelling must discard something, and the first thing discarded is the metadata — error bars, sample size, the boundaries of the definition — because it's "hard to say." A news headline stripping qualifiers is the same move: the body says "may, under specific conditions," the headline keeps only "does." Distributed systems face it too: after a datum passes through several hops, its lineage (where it came from, on what definition) is the first casualty, and downstream you get a bare value with no context.
Before you cite any "X% of teams / projects / users…" in a report or proposal, run a "qualifier checkup": did the original say "roughly" or "under certain conditions"? Are you quietly reading "fell short of expectations" as "failed"? Trace the line you want to quote back to its source, and you'll often find it far more modest than you remembered.
The number you last used to "settle the matter" in a decision — if it were required to travel with its error margin, its definition, and its first-hand source attached, would it still dare stand in front of the conclusion?
A number thoroughly disproven often keeps right on living — that's a "zombie statistic." The reason is plain: it serves a narrative need, not a need for truth. As long as the story still wants it as a hook, disproving evidence can't kill it. A number survives not by being true, but by being useful.
Debunking is inherently a weak signal: it arrives late, spreads narrowly, and is dreadful to tell ("there's actually no reliable data" grips far less than a punchy percentage). The original number, meanwhile, is round, dramatic, and useful — all natural transmission advantages. Worse still is the psychological "continued influence effect": once a person accepts a claim, it keeps shaping their later judgments even after they're corrected to their face. Deleting a piece of information is far harder than implanting it.
"We only use 10% of our brains" — neuroscience demolished it long ago (imaging shows every brain region gets used over the course of a day; there is no vast, permanently dormant reserve), yet it remains one of the most stubborn popular claims, because it's perfect for self-help, sci-fi, and advertising: 90% still to unlock, how tempting. "You must drink 8 glasses of water a day" is the same — trace it to the earliest recommendation in the 1940s and the original had a second half: "most of this quantity is already contained in the foods we eat." The first half became law; the inconvenient second half was thrown into the dustbin of history. Trivial to disprove; impossible to kill.
The epidemiological lens fits well: a rumor behaves like a pathogen and a correction like a vaccine — the trouble is the vaccination rate is too low and comes too late to stop an outbreak already spreading. Marketing exploits the same law in reverse: a "useful but untrue" line routinely out-transmits a "true but tedious" one. In any narrative-driven system, "useful" carries more weight than "true."
When you want to correct a mistaken belief in a child or a team, "debunking" alone will almost certainly fail — you've left a hole where a story was, without filling it. What actually works is offering an equally tellable, equally useful replacement narrative to take its place, not merely announcing the old one is wrong. The mind dislikes a vacuum; give it no new story, and it retrieves the old one.
The last time you corrected someone's mistaken claim, did you hand them a "that's wrong," or a new story — more useful, more memorable — good enough to replace it?
Fighting zombie statistics isn't about "more skepticism" — blanket suspicion is both exhausting and unsustainable. What works is a fixed routine: any number that makes your fingers itch to forward it or write it into a report first passes three gates — trace to the source, check the definition, ask who's selling. Turn screening from an attitude into a process.
Trace to the source: follow the citations back to the original paper and see what it actually measured, on how large a sample, under what conditions. Check the definition: how exactly is "failure," "adoption," or "growth" defined, and what's the denominator — change the operational definition of the same phenomenon and the answer can shift by an order of magnitude. Ask who's selling: who's spreading this number, and what do they sell with it (cui bono — who benefits). If it fails any one gate, demote it to "unverified" and keep it out of your conclusion.
Numbers that can't pass the three gates tend to share a face: a round integer (exactly 70%), no denominator, no time window, and a vague "studies show" for a source rather than a named one. Conversely, numbers that hold up usually drag a string of "tedious" qualifiers behind them and read anything but catchy. So a practical heuristic: the clean, tidy number that perfectly supports the conclusion you wanted and is oh-so-shareable is precisely the one to stop and check first — it's too useful, useful to the point of suspicious.
This habit has mature counterparts in engineering: distributed systems speak of "data lineage," where every value can be traced to how it was computed; software supply chains use a "bill of materials" (SBOM) that demands you account for where each component came from; science pushes "preregistration plus open raw data," institutionalizing exactly "trace to the source, check the definition." Information literacy is just moving those engineering disciplines inside your own head.
This dovetails with the calibrator Day 65 built for "I've got it" — now install one for your information intake too: any number destined for a decision, a report, or even a message to your child passes the three gates before use. It also previews the next stop: Day 66 was about a number's origins (where it came from); "Ask About the Ruler First" will be about a number's scale (how it was actually measured) — the same question, measured with a different ruler, can yield two wildly different answers.
The key number you last wrote into a report or decision — if you forced it right now through "trace to the source, check the definition, ask who's selling," how many of the three gates would it survive?