DAY 61

Health & Longevity: Evidence Literacy
How to Read Health Research Without Being Fooled

2026-07-17 · BigCat's Vitality Protocol
This issue is a meta-skill: no new protocol, but a toolkit for reading every protocol. Grounded in evidence-based-medicine methodology and classic case studies.
SUB · Study Design / Causal Inference
The Evidence Pyramid: Not All "Studies Show" Are Equal
Design, Sample Size, and a Control Group
Bottom Line
When you see "studies show," ask three things first: what design, how many people, was there a control. Observational studies establish only correlation; only an RCT approaches causation—the two differ by roughly an order of magnitude in evidentiary weight.
Evidence Grade
Expert consensus: the OCEBM (Oxford Centre for Evidence-Based Medicine) Levels of Evidence (2011) and the GRADE system are the standard teaching frameworks worldwide.
Science & Mechanism
Systematic Review / Meta-analysis
Randomized Controlled Trial (RCT)
Prospective Cohort
Case-Control
Cross-sectional / Correlational
Animal / Mechanistic
Expert Opinion / Case Report
▲ Higher = better control of confounding, closer to causation
Correlation is not causation, and there are three common escape routes: confounding (a third variable drives both ends—"red-wine drinkers live longer," but the real culprit is often socioeconomic status), reverse causation (not X causes Y but Y causes X—"low cholesterol tracks with cancer," when in fact early cancer lowers cholesterol), and selection bias. An RCT uses randomization to average out known and unknown confounders at once—that is the sole reason it outranks observational work. Bradford Hill's nine criteria (1965)—strength, dose-response, temporality, consistency, biological plausibility, and more—help move cautiously from correlation toward causation.
Actionable Protocol
DesignWhat it answersTypical trap
Meta / Systematic reviewSynthesizes many studiesGarbage in, garbage out; heterogeneity
RCTCausationSmall n, narrow population, short follow-up
Prospective cohortStrong correlation + temporalityResidual confounding
Case-controlRare disease, fastRecall bias
Cross-sectionalGenerates hypothesesNo temporality, no causal claim
Animal / mechanisticPathways, feasibilityDose & species extrapolation
Note for Women
Many "classic findings" come from predominantly male samples. Female inclusion in animal research has long lagged—Beery & Zucker (2011) found male:female ratios reaching 5:1 across several fields. Extrapolating results from male rodents or male subjects to women carries real risk. Reading a study, first check the sex composition of subjects and whether sex-stratified analysis was done.
Common Myths
Myth 1: "An RCT is always the most trustworthy." Small, short, surrogate-endpoint RCTs mislead too; a well-designed large cohort often beats a poor RCT.
Myth 2: "If the mechanism makes sense, it works." Countless mechanistically plausible therapies have failed in RCTs (the CAST trial in the next card is the classic).
Key References
• OCEBM Levels of Evidence Working Group. 2011.
• Hill AB. The environment and disease: association or causation? Proc R Soc Med. 1965;58:295-300.
This Week + Reflection
THIS WEEK
Pick one health claim you've recently believed (e.g. "X is anti-inflammatory / neuroprotective") and trace it back to its original source: is it a human RCT, or observational, animal, or in-vitro work?

Reflection: if a recommendation rests only on mechanism and animal data, would you change your behavior for it? Where should the bar sit?
SUB · Risk Communication / Statistical Literacy
Relative vs Absolute Risk: The Headline's Favorite Trick
Relative vs Absolute Risk & NNT
Bottom Line
"Cuts risk by 50%" carries almost no information unless you know the baseline—going from 2% to 1% is also "down 50%." Always ask for the absolute number and the NNT (number needed to treat).
Evidence Grade
Methodological consensus: Cook & Sackett (1995, BMJ) established NNT; Gigerenzer's body of risk-literacy work shows relative-risk framing systematically misleads both doctors and patients.
Science & Mechanism
Relative risk reduction (RRR) inflates perception; absolute risk reduction (ARR) is what you actually get. NNT = 1/ARR, meaning "how many people must be treated to prevent one event." Example: a statin in primary prevention cuts 5-year heart-attack risk from 2.0% to 1.5%—RRR = 25% (the headline's darling), ARR = 0.5%, NNT = 200. That is, treat 200 people for 5 years to prevent one heart attack; the other 199 get no benefit yet bear cost and side effects. Weigh the downside with NNH (number needed to harm) the same way.
Actionable Protocol
Three ways to state the same dataPhrasingHonesty
RRR 25%"Risk cut by a quarter"Exaggerated
ARR 0.5%"Half an event fewer per 100"Honest
NNT 200"Treat 200 to prevent 1"Most useful
Pocket formula: ARR = control event rate − intervention event rate; NNT = 1 / ARR. Any report giving only relative risk and no absolute number—treat with suspicion.
Note for Women
Women bear the brunt of poor risk communication: breast-cancer screening and HRT benefits/risks are routinely presented as relative risk, breeding fear. WHI (2002) reported HRT raised "breast-cancer risk by ~26%" (relative); the absolute increment is roughly 8 extra cases per 10,000 users per year—same fact, but the absolute framing makes individual decisions far more rational.
Common Myths
Myth 1: "Higher 5-year survival = living longer." Earlier diagnosis (lead-time bias) can lift survival rates without extending life.
Myth 2: Applying RRR to yourself as if it were ARR, overstating your personal benefit.
Key References
• Cook RJ, Sackett DL. The number needed to treat. BMJ. 1995;310:452-454.
• Rossouw JE, et al. (Women's Health Initiative). JAMA. 2002;288:321-333.
This Week + Reflection
THIS WEEK
Next time you see a "cuts risk by X%" health story, go find the absolute risk and NNT in the original—often the "stunning effect" shrinks to a single decimal place.

Reflection: would your willingness to take a drug with NNT = 200 really equal that of one with NNT = 8?
SUB · Statistical Traps / Reproducibility
p<0.05 Doesn't Mean "True": The Triple Trap of Significance
p-Values, Power & Publication Bias
Bottom Line
p<0.05 says only "if the intervention did nothing, the chance of getting this data is under 5%"—it says nothing about effect size, and even less about whether the conclusion is true. Small n + multiple comparisons + publishing only positives = a pile of false positives.
Evidence Grade
Methodological consensus: the American Statistical Association's 2016 statement on p-values (Wasserstein & Lazar); Ioannidis (2005, PLoS Med) argued that most published findings are false.
Science & Mechanism
Three independent traps:
① Statistical ≠ clinical significance—with a large enough sample, a trivial difference (e.g. 0.3 kg of weight loss) can hit p<0.05.
② Underpowered studies (too small n)—real effects go undetected (false negatives); and positives that squeak through tend to overstate the effect (the "winner's curse").
③ p-hacking & publication bias—trying many endpoints/subgroups until one lands at p<0.05; negative results get filed away unpublished, tilting the whole literature positive. This is exactly why so many "breakthroughs" fail to replicate. Ioannidis (2005) made the formal case that most published conclusions are false.
Actionable Protocol
A checklist for reading a study:
• How big is n? Is the effect size (absolute difference, Cohen's d) reported?
• Was the primary endpoint pre-registered, or cherry-picked afterward from a pile of results?
• How wide is the confidence interval? Does it cross the "no effect" line?
• If it's a meta-analysis, is the funnel plot symmetric (a check for publication bias, Egger 1997)?
• Treat a single "stunning" study with doubt; wait for replication.
Note for Women
Too few female subjects directly lowers the power to detect effects in the female subgroup—so even when a study is positive overall, "it works for women too" may just mean no difference was detected, not that there truly is none. Check whether a "sex × intervention" interaction test was run, not merely a line in the footnote.
Common Myths
Myth 1: "p=0.049 and p=0.051 are worlds apart." 0.05 is an arbitrary threshold; the two carry almost identical evidentiary weight.
Myth 2: "Bigger n is always better, never mind effect size." Large samples make even trivial differences "significant."
Key References
• Wasserstein RL, Lazar NA. The ASA statement on p-values. Am Stat. 2016;70:129-133.
• Ioannidis JPA. Why most published research findings are false. PLoS Med. 2005;2:e124.
• Egger M, et al. BMJ. 1997;315:629-634.
This Week + Reflection
THIS WEEK
Find a study you care about and look only at its sample size and confidence interval first (skip the abstract's conclusion for now); judge for yourself whether the conclusion is solid, then compare.

Reflection: if an effect needs n = 10,000 to barely reach significance, what does it mean for you personally?
SUB · Spotting Pseudoscience / Critical Thinking
Five Fingerprints of Pseudoscience
Five Fingerprints of Pseudoscience
Bottom Line
Surrogate endpoints posing as hard ones, animal/test-tube results extrapolated straight to humans, the health halo, cherry-picked cases, and refusal to be falsifiable—the more it hits, the more suspect.
Evidence Grade
Expert consensus + classic cases (the CAST trial below is the textbook example of a hard-endpoint / surrogate-endpoint conflict).
Science & Mechanism
Unpacking the five fingerprints:
① Surrogate posing as hard endpoint—an intermediate marker standing in for the outcome you truly care about. Counterexample CAST (Echt 1991, NEJM): antiarrhythmic drugs successfully suppressed premature ventricular beats (surrogate) yet doubled mortality (hard endpoint); a plausible mechanism can still be lethal.
② Over-extrapolation—works in a test tube/mouse ≠ works in humans, and lab doses are often tens of times human levels.
③ Health halo—"natural," "detox," "immune-boosting" and other vague terms with no operational definition.
④ Cherry-picking—citing only the studies that flatter the claim.
⑤ Unfalsifiable—"it didn't work because you didn't stick with it / too many toxins," rationalizing any outcome.
Actionable Protocol
Run a "mine-sweep" checklist before you buy:
• Is it about a hard endpoint (death, heart attack, fracture) or some surrogate marker?
• Is the evidence from a human RCT, or mice, test tubes, personal experience?
• Are there specific doses and falsifiable predictions?
• Is the seller also the researcher? Is the conflict of interest disclosed?
• Does it pile up undefined jargon like "quantum," "detox," "boost your energy"?
Note for Women
Women's health is a pseudoscience hotspot: "hormone balancing," "womb detox," "period cleanses" and the like mostly lack RCTs and use vague symptoms (fatigue, bloating) as proof of efficacy. For pregnancy and perimenopause products especially, insist on hard endpoints and regulatory status.
Common Myths
Myth 1: "Traditional/natural = safe and effective." Natural isn't safe (see Day 45), and long-standing isn't the same as effective.
Myth 2: "It worked for me, so that's evidence." Placebo, spontaneous recovery, and regression to the mean can all manufacture the illusion that "I tried it and it really helped."
Key References
• Echt DS, et al. (CAST). Mortality and morbidity in patients receiving encainide, flecainide, or placebo. N Engl J Med. 1991;324:781-788.
• Ioannidis JPA. PLoS Med. 2005;2:e124.
This Week + Reflection
THIS WEEK
Take a supplement or wellness plan you're using or thinking of buying and tick it against the five fingerprints—how many does it hit?

Reflection: if a plan "can call itself effective no matter the outcome," what does that fact alone tell you?
Deeper Reflections
① If RCTs are strongest, why does nutrition science have almost no lifelong dietary RCTs?
Neither ethics nor adherence allow it: making a group strictly eat a prescribed diet for decades is neither realistic nor humane. So nutrition leans on cohorts + mechanism, and its evidence grade is inherently a notch lower—which is exactly why nutrition conclusions keep flip-flopping. Demanding cardiovascular-drug-grade evidence from nutrition is unrealistic; the mature stance is to work with the best currently available evidence rather than wait for a perfect RCT that will never arrive.
② Is an intervention with tiny absolute risk (huge NNT) always not worth doing?
Not necessarily. If the intervention is cheap, harmless, and population-wide (iodized salt, vaccines), even a large individual NNT can add up to enormous absolute benefit at the population level—Geoffrey Rose's "prevention paradox." The optimal answer differs between the individual view (is this worth it for me) and the public-health view (is it worth it across millions); both are right, they just ask different questions.
③ Can pre-registration and open data cure publication bias?
They ease it greatly but can't fully cure it. Pre-registration locks in the primary endpoint and shrinks the room for p-hacking; mandatory registries like ClinicalTrials.gov make "vanished negative trials" harder to hide. But negative results can still go unsubmitted or uncited. Journals accepting Registered Reports (reviewing methods first, then publishing regardless of result) is the more thorough fix. The system is moving the right way, but the reader's skepticism is still non-optional.
④ Facing contradictory studies, how should an individual actually act?
Three heuristics. First, weigh evidence level and consistency—several high-quality studies pointing the same way beat one dazzling result. Second, distinguish decision types: "stop a harmful habit" can demand low evidence, while "adopt a new intervention" demands high. Third, sort by cost and risk: try the cheap and harmless first, wait for replication on the costly and risky. You needn't wait for certainty to act—but the action should be proportional to the size of the uncertainty.