This is the plain-language edition. The full argument, the caliber record for every figure, all sources and all verification verdicts are in the deep dive. This is an evidence review, not medical advice. Nothing here should be taken as a reason to use any drug.
Arguments about longevity drugs and longevity diets almost never happen at the level of "is this study real." Most of the studies are real. The problem is which tier those real studies get moved to.
Evidence for any longevity intervention can only sit in one of three tiers:
The four star interventions sit in wildly different places, and their fame has almost nothing to do with where they sit. We checked the four most-quoted signature numbers one at a time, and the result is that all four fail if quoted as they usually are. Each is the product of the same class of small move — not fabrication, but taking a number from the caliber where it is true to one where it is not.
There are four such moves:
One at a time.
This is the only one of the four with no real argument at the animal level. The US National Institute on Aging runs a dedicated mouse lifespan testing program (the ITP): three labs run the same protocol in parallel, on mice bred so that every animal is genetically different, and it commits in writing to publishing results whether they are positive or negative — a rare piece of honest design in this field. Rapamycin is the largest and most dose-stable hit in that program.
But the widely circulated "extends lifespan by 14%" cannot be quoted as is. The original says "on the basis of age at 90% mortality" — a measure of longest lifespan, not the median most people picture — and dosing began at 600 days of age, already late life. Another pair of numbers from the same paper (9% and 13%), routinely quoted as medians, are introduced in the original as mean lifespan.
The dose differences are startling: the lowest dose did essentially nothing in male mice (+3%, not significant), while the highest reached +23%. So "rapamycin extends mouse lifespan by X%" means nothing without a dose and a start age.
The human tier is another world. The most comprehensive systematic review screened over eighteen thousand articles and included just 19 human studies — and those 19 cover rapamycin's close relatives too, not rapamycin alone. And the largest human trial on this line is a failure: 1,024 participants, primary endpoint missed, with the result pointing the wrong way.
That widely repeated "boosted immune response in the elderly by 20%" is not actually about rapamycin — it used a related drug (everolimus); the criterion for calling it a success was an 80% Bayesian posterior probability rather than conventional statistical significance; and treatment-related side effects ran at roughly twice the placebo rate.
The trial in healthy people (PEARL) was run by a telehealth company that sells the drug. Its registered main goal was reducing visceral fat, and it was not met. The heavily cited positive result rests on eight women — and the paper itself, a percentage inside the same sentence, and the company's own webinar give three different numbers for how many.
In 2026 someone finally tested head-on the premise the entire off-label community relies on: does taking it once a week give you the benefit while avoiding the harm? A 40-person, 13-week trial found no support — the main measure pointed toward placebo, and a blood-sugar marker (HbA1c) rose slightly. The trial is small and self-described as exploratory, so it settles nothing on its own. But it is the only head-on test so far, and the "once a week is safe" story has never had any human evidence behind it — while in mice, an independent study partly contradicted it.
The only anti-aging rapamycin trial with "how long they live" as its main goal is in dogs. It targets 580 of them, gives each a year of dosing plus two years of monitoring, and its design paper publishes no reporting date — so it is still years away.
This is the cleanest of the four cases, because it went the whole distance.
In 2014, a British study produced a conclusion that startled everyone: diabetics taking metformin outlive people without diabetes. It was cited for twelve years.
Its actual caliber: a "median survival time ratio" of about 15% extrapolated by a statistical model, over a median follow-up of only about two and a half years. The paper concedes it did not adjust for weight, blood sugar, blood pressure, cholesterol or kidney function, because those data were largely missing for the control group. And the crucial detail: in the raw data the difference was not significant (p=0.054), with the authors writing that the two survival curves "show that overall there was little discernible difference." The study was funded by two pharmaceutical companies, one of whose employees was a co-author.
In 2022, a Danish team ran the same comparison on national registry data and the result reversed. What makes this reversal persuasive is what they did next: they put their own mortality rates beside the British ones, and three of the four cells match almost exactly. Two countries with different health systems and different populations agree in three cells. That rules out "the two countries are just different" and compresses the disagreement into the one remaining cell.
In 2023, a Welsh team with nearly 130,000 patients did not just contradict the number — they explained it. Re-running the analysis across time windows of different lengths, they found metformin's "survival advantage" does exist in the first three years and reverses after five. The 2014 conclusion can be reproduced in independent data — as an artifact of following people for too short a time.
There is also randomized evidence. Over three thousand adults with elevated blood sugar but not yet diabetes were randomized and followed for twenty-one years. Metformin did reduce diabetes onset — it works on its own indication. But the hazard ratio for death from any cause was 0.99: nothing moved in twenty-one years. To be fair, the trial was never designed to measure mortality, so this is "no difference detected," not "proven ineffective."
As for the big trial meant to prove metformin slows aging (TAME), awaited for eleven years: as of July 2026 it has no clinical trial registration number and has never enrolled a single participant. Its designer has publicly predicted its imminent start every year since 2015. Meanwhile websites already report the "results" of this never-started trial as fact.
Evidence running the other way rarely appears next to the good news: two randomized trials found metformin blunts the benefits of exercise. In one (MASTERS), the researchers had predicted it would enhance the training response and found the opposite. Even the most enthusiastic advocate says people under 50 without diabetes should not take it.
And the study reported worldwide as "metformin delays aging in monkeys" dosed six male monkeys and states explicitly that it did not observe survival. The advocate described it publicly as showing aging delayed by eight years; the largest number in the paper is 6.86 years, on a different measure — that "eight years" appears to have been carried over from a separate human study of his own.
The first three interventions have at least one real animal lifespan signal. The NAD+ supplement line has a problem further upstream: its major premise does not hold in humans.
The whole industry rests on one sentence: NAD+ declines with age, so put it back. The human evidence for that sentence is a 49-person skin biopsy study and a 17-person brain scan study (split into groups of seven, four and six). A later, larger brain study — funded by a food company, so with incentives pointing toward confirming a decline — failed to reach significance.
Then in May 2026, a 32-author European team measured seven population cohorts and found that whole-blood NAD+ in humans does not decline with age. That study has an obvious weakness of its own: the group built specifically to compare ages was only 20 versus 20 people, which is underpowered. But it is not alone: a 1,518-person Chinese study using a completely different method on a different population likewise found no consistent decline.
So this needs two separate sentences. "Human blood NAD+ does not decline appreciably with age" now has two independent supports. "Human NAD+ does not decline with age" outruns the evidence — the tissue-level evidence points the other way, and the 2026 paper itself lists measuring only blood as its first limitation. Worth noting: the two most-cited studies supporting a decline share an author and a university — the same independence problem this article charges elsewhere.
Supplements do raise blood NAD+; that isn't disputed. The next step is the problem: to do anything, it has to rise in the tissue doing the work. The only trial to biopsy muscle directly in older people found muscle NAD+ did not rise significantly, and grip strength did not change (p=0.96). Two other independent trials, using different methods, double the dose and four times the duration, got the same answer.
Pooled analyses of functional endpoints are essentially null: no effect on muscle mass or strength, no significant metabolic benefit. A review that assessed all 25 human studies then published concluded the effects are few, clinically marginal, and systematically overstated in the literature. (To be fair: two other pooled analyses point the other way, so this is not unanimous.)
The real identity of two signature numbers: "NR lowers blood pressure by 9 mmHg" is a post-hoc subgroup inside a 24-person trial, about which the authors explicitly wrote that no statistical inferences can be made. "NR improves walking in peripheral artery disease" was judged positive against a pre-loosened one-sided standard, and the effect was 17.6 metres — above the paper's own "small" threshold for clinical importance (about 8 m) but below its "large" one (20 m).
The two most important facts: in that three-lab mouse lifespan program, NR clearly failed (males p=0.252, females p=0.612). And NMN's most famous clinical result (improved insulin sensitivity) had only 13 versus 12 participants and was challenged in Science on the grounds that randomization failed — baseline liver fat differed more than twofold between arms, and liver fat is exactly what this drug targets. More decisively: the same lab ran a bigger, longer, higher-dose repetition and it did not replicate. The results are already posted to the trial registry.
One more thing worth knowing: a 2026 head-to-head trial proposed that gut bacteria convert these expensive precursors into nicotinic acid — a very cheap vitamin — which then raises NAD+ throughout the body. If that mechanism holds, the expensive one is doing the cheap one's job. And a head-to-head trial of NR or NMN against plain nicotinic acid does not exist on any clinical endpoint.
This is the only one of the four tested in multi-year randomized trials in humans, which makes its evidence the sturdiest — and the best illustration of this field's ceiling.
The famous two-year caloric restriction trial (CALERIE) has two facts you need. First, it prescribed a 25% calorie cut and participants actually achieved 11.9% — and not evenly: about 19.5% in the first six months, only 9.1% for the following eighteen. So any claim about "what happens when humans cut calories 25%" is really describing about 12%. That is not the researchers' failure; it is the ceiling of human adherence. Second, both of its registered main goals were null at two years (one was met at one year).
The widely cited "caloric restriction slowed aging" result comes from a post-hoc analysis of stored samples. It tested 11 "biological age" measures, and one moved.
We were going to write that up as cherry-picking one out of eleven. The review process struck that down on hard grounds: the ten measures that did not move quantify a level in years, and a one-year intervention mathematically cannot shift them; the one that moved measures the rate of aging, and is the only measure that could have moved on this timescale. Its significance also clears a strict multiple-comparison correction. So the effect is real.
The problem is what it gets converted into. The paper's own discussion converts it to "a reduction in mortality risk of as much as 10-15%, similar to the effect of smoking cessation." A Norwegian study then measured the exact link that conversion needs — the same people measured eleven years apart, testing whether a change in aging rate predicts death — and found nothing. Across different populations, the mortality risk per unit of this measure ranges from 1.23 to 1.99, a spread too wide to support any precise conversion. And to this day, no independent team has measured this marker in an independent randomized caloric-restriction trial.
The cost side deserves the same precision. Two years of restriction did significantly reduce bone density at the spine, hip and femoral neck by about 2% — but not at the wrist or whole body; the researchers calculated that for a 50-year-old woman this raises ten-year fracture risk by less than 0.5%; and they state it is the change expected from losing that much weight by any means, not a toxicity specific to restriction. The companion safety analysis concluded it was "safe and well tolerated." A real cost, and a small one.
The famous monkey mystery has been solved, and the answer is instructive. The Wisconsin study said restriction extended monkey lifespan; the National Institute on Aging study said it did not. A 2017 joint analysis found the key difference was what the control monkeys ate: sucrose was 45% of total carbohydrate in Wisconsin's chow versus under 7% at NIA, and NIA's "controls" were not free-feeding but given measured portions. So how much of Wisconsin's benefit was "eating less" and how much was "not continuously eating a high-sugar refined diet" cannot be separated in that design. Also: Wisconsin's 2009 announcement used "age-related deaths" as its endpoint, while all-cause mortality in the same paper was not significant.
Time-restricted eating (16:8, for example) is essentially about calories, not timing. A trial that provided meals and held calories fixed found no benefit at all; a 12-month trial with both arms calorie-restricted detected no difference. But timing is not entirely irrelevant: one trial with both arms calorie-restricted found that putting the eating window in the morning lost 2.3 kg more, and alternate-day fasting has a small edge over continuous restriction.
Finally, a case of hype in both directions. In 2024, "an 8-hour eating window raises cardiovascular death risk by 91%" swept the world. That 91% came from the association's press release, not the scientific abstract; it was an unreviewed conference abstract; and each person's "eating window" label came from exactly two days of diet recall questionnaires, used to characterize an average of eight years of follow-up.
It was published in full in 2025, and the direction is the opposite of what most people would guess: after peer review the figure was +135%, larger than the press release's 91%. The lesson is not "the scary headline was later debunked" but that publishing ahead of peer review is itself the problem, whichever way the number later moves.
Hype runs the other way too: the fasting-mimicking diet's "2.5 years younger" is the treated group compared with itself before and after, even though the trial had a control group; and the circulating "11 years younger" is not a measurement at all but a computer simulation assuming the effect persists for twenty years.
All of the above keeps hitting the same wall: tier three — people actually healthier and longer-lived — is nearly empty. The reason is not mysterious.
"Aging" is not a registrable indication in any country. A review that screened 3,780 papers found no country with a regulatory framework for drugs targeting aging. And — this deserves naming — no primary FDA document has ever accepted "aging" as an indication, even though multiple peer-reviewed reviews assert in print that "the FDA has approved TAME." One of the reviews that says so is the very paper we used to establish that no framework exists.
So substitutes are used, and the most popular substitute is the epigenetic "biological age clock." Its problems come in two kinds that are completely different in nature:
Measuring the same DNA twice gives different answers — in the worst case, one of six clocks differed by 9 years. (Though typical error across the clocks is 0.9 to 2.4 years, so quoting the worst case as typical is exactly the move this article criticizes.) Notably, the paper documenting the problem was also selling the fix: its improved method is licensed to a consumer biological-age testing company, which also supplied the data.
Sample a different tissue and the results diverge much more. Run a blood-trained clock on an oral sample and the same person at the same moment can be assigned a "biological age" decades apart. But that sentence has to travel with its other half: blood versus blood is essentially identical (all seven clocks non-significant). So this is not "biological age can't be measured" but a highly consistent, calibratable systematic bias — take the blood ruler to the mouth and it reliably reads far off. Two independent cohorts show the same structure, at a magnitude of 9 to 20 years.
What that does and does not license: it does not license "different vendors will report ages decades apart" — the vendors that sample oral tissue specifically do not use blood-trained clocks. It does license: biological ages from different tissues are not on a common scale and cannot be compared; and no study has ever split one person's sample across several commercial vendors and published how far apart the answers land.
One more number says more about this field's real shape than anything above. We queried the US clinical trials registry: a search for "aging" (which the registry expands to synonyms automatically) returns 3,076 interventional studies, while only 1,902 actually list Aging in the condition field; of the 3,076, 71.5% enroll no more than 100 people. But the composition is the more informative part: within that same result set, "skin aging" returns 429 studies, "facial aging" 368 and "wrinkle" 233 — far more than metformin at 23, rapamycin at 30, NAD at 35 and taurine at 5. By count, registered "anti-aging clinical research" is mostly cosmetic medicine, not geriatric medicine.
Put it all together and a pattern appears:
The explanation is not mysterious: what can be retailed at scale must be a supplement; a supplement legally cannot be a drug; and what actually works at the animal lifespan tier is usually a drug. So market size is set not by how large the effect is but by regulatory category. Rapamycin cannot be packaged into a capsule and sold to you not because its evidence is weak, but precisely because it is a drug.
That also explains why all four signature numbers are products of caliber-switching. When no qualified endpoint exists, and market access depends on regulatory category rather than effect size, the function of "evidence" degrades from deciding what to sell into decorating what is already being sold. These moves are not individual researchers' moral failures; they are the predictable output of that incentive structure.
Each item below states what would prove it wrong. Ordered from sturdiest to most in need of watching.
1. Taking NR or NMN raises NAD+ in blood but not in resting muscle. Three independent datasets agree, one of which doubled the dose and quadrupled the duration with no change. How to test: any new biopsy trial must report the muscle NAD+ figure itself, not only its downstream products.
2. Metformin failed to extend lifespan in that three-lab mouse program. Males +7% but not significant, no effect in females — and the program has run metformin only once. How to test: the program re-running it at a higher dose or earlier start age and publishing the result.
3. The 2014 "diabetics live longer" finding is an artifact of too-short follow-up. Two countries' independent data reverse it, and the Welsh team reconstructed the artifact: present in the first three years, reversed after five. How to test: any study claiming to restore that conclusion must compare against people without diabetes, not against another glucose-lowering drug.
4. Time-restricted eating's weight benefit comes from eating less, not from timing. Hold calories fixed and the benefit disappears. How to test: any claim of an independent timing effect must reproduce under provided-meal, fixed-calorie conditions. Note that counter-evidence already exists (a morning eating window lost 2.3 kg more), so this claim covers time-restricted eating only, not all forms of fasting.
5. Human whole-blood NAD+ does not decline appreciably with age. Two independent teams, two countries, two methods agree. How to test: any "NAD+ declines with age" claim must specify the tissue — blood now has two independent nulls, and the tissue evidence points the other way.
6. The 8-to-12% small wins in that mouse program sit at the limit of what it can detect. The best evidence is what the program says about itself: it wrote in print that these small results may be chance. The same set of experiments contains a near-identical pair of gains (+8.3% counted as a miss, +8.7% as a hit) landing on opposite sides of the statistical threshold. How to test: whether any sub-10% result survives an independent re-run at the original dose and start age.
7. That eleven-years-awaited metformin aging trial has never started. No registration number, not one participant. How to test: a registration number appearing with actual enrollment above zero.
8. "Take rapamycin once a week" was not supported the first time it was tested head-on. The main measure pointed toward placebo, alongside a small rise in a blood-sugar marker — but the trial is small and exploratory. How to test: an adequately powered replication that also measures the drug's actual biological action; plus the dog trial whenever it reports.
9. Caloric restriction really did slow the "rate of aging" measure, but converting that into mortality risk does not hold. The effect survives strict correction; a Norwegian study measuring directly whether a change in that rate predicts death found nothing. How to test: any independent team measuring this marker in an independent randomized restriction trial — it has never been done.
10. Biological ages from different tissues cannot be compared, but this is a calibratable systematic bias, not measurement noise. Blood versus blood is essentially identical; oral versus blood differs by 9 to 20 years in two independent cohorts. How to test: split one person's sample across several commercial vendors and publish the spread — a study that does not yet exist.
11. "Aging" is not a registrable indication anywhere, and no "biological age clock" has been validated as a qualified substitute measure in the regulatory sense. How to test: a study showing that how much an intervention slows a clock predicts how much outcomes improve; or an aging-related entry appearing on the FDA's qualified biomarker list.
What to watch: the dog rapamycin trial is this field's first rigorous drug trial with "how long they live" as the main goal, and the only data in the next few years that could genuinely change the conclusions. Beyond that, watch the annual reports from that three-lab mouse program — urolithin A and taurine are being tested in it right now, results are a few years out, and they are the two loudest longevity claims of recent years.
When you see any longevity number, ask which tier it's on. Animal lifespan, a human marker, humans actually healthier — the distance between these three tiers is larger than any argument inside one of them. Most misleading claims are not invented numbers; they are tier-one numbers used as tier-three numbers.
Read "the main goal was missed but a subgroup improved" as a negative result. Trials register what they will measure precisely to prevent picking afterwards. PEARL missed its main goal and its positive result rests on eight people; NMN's signature result had 13 versus 12 participants, and the same lab could not replicate it.
"The treated group compared with itself" is not a randomized comparison. The fasting-mimicking diet's "2.5 years younger" and NMN's "25% improvement" are both before-and-after within one group — even though those trials had control groups available. Having a control group and not using it is a signal.
The prescribed dose is not the dose taken. That two-year restriction trial prescribed 25% and delivered 11.9% — and there is still no human data on what 25% restriction does, because nobody can sustain it.
Take a serious look at that mouse program's ledger. It is the only scoreboard in this field that serves nobody's position, and it commits to publishing whatever it finds. Metformin failed. NR failed. Fisetin failed. The ones that passed — rapamycin, acarbose, 17α-estradiol — are mostly prescription drugs, and their effects often appear in only one sex.
None of this is advice about what to take, including no advice to take nothing. This article grades evidence strength; it does not make personal decisions. Anyone seriously considering any of these interventions needs a physician, not an article — especially for off-label use of a prescription drug, where the risks are specific and individual. The one thing this article can confirm is that no human data currently support the idea that once a week is both effective and safe.