DAY 58

Health & Longevity: Digital Health Tools
Wearables, CGM, Apps & AI — Read Them Critically

2026-07-14 · BigCat's Vitality Protocol
Evidence this issue: device accuracy from validation studies + cohorts; digital therapeutics & screening from RCTs; AI diagnosis from imaging RCTs + symptom-checker evaluations
SUB · Wearables / Data Literacy
Wearable Data — Read Trends, Not Absolutes
Your watch is a trend tracker, not a medical device
Bottom Line
A band or watch is a trend tracker, not a medical device. Resting heart rate is accurate; sleep staging, HRV and SpO2 are only relative references. Compare yourself to your own baseline, not to someone else's absolute numbers — and don't let a "recovery score" run your life.
Evidence Grade
Validation + cohort: optical PPG heart rate is accurate at rest, but degrades during high-intensity/interval exercise and on darker skin (Bent 2020, npj Digit Med). Against polysomnography, consumer devices call "asleep vs awake" reasonably well but stage (light/deep/REM) with only moderate accuracy. The Apple Heart Study (2019, NEJM) showed a watch can flag AFib, but with high false positives in low-risk people — it's a prompt to get checked, not a diagnosis.
Science + Mechanism
Watches use optical PPG (green light sensing blood-flow pulsation) for heart rate, an accelerometer for steps and sleep movement, and beat-to-beat (RR) intervals for HRV. All are indirect estimates: skin tone, tattoos, fit and motion artifact all interfere. So-called "recovery/readiness scores" are each vendor's undisclosed black-box algorithm blending HRV, resting HR and sleep into one number — useful for your own trend, but with no unified clinical validation and no diagnostic authority.
Protocol
Track the two most reliable signals: morning resting heart rate and HRV, under fixed conditions (lying still on waking, same time), and watch the 7–14 day trend
Reading the trend: sustained higher resting HR + sustained lower HRV usually means fatigue, poor sleep, incoming illness or alcohol — a body signal, not a device fault
AFib / arrhythmia alerts: don't panic, don't ignore — take the record to a proper ECG for confirmation
Don't over-attribute: minor day-to-day wobble in sleep stages or SpO2 is worth a glance, not daily anxiety
For Women + Common Myths
The menstrual cycle shifts readings: in the luteal phase (premenstrual), resting HR rises, HRV drops and body temperature climbs — this is normal physiology, not "poor recovery / overtraining." Don't needlessly cut training or blame yourself. Some devices use the temperature shift to help estimate the fertile window, but precision is limited.
Myth 1: "Low sleep score = bad sleep" — obsessing over the score fuels anxiety and insomnia (orthosomnia, the "perfect-sleep compulsion"); the more you chase it, the worse you sleep.
Myth 2: "SpO2 of 95% is dangerous" — consumer pulse oximetry has wide error; a few points of drift with no symptoms is usually measurement noise, not hypoxia.
Try This Week + Reflection
THIS WEEK
Log your morning resting heart rate under fixed conditions for 7 straight days; look only at the trend line, not any single day's absolute value. Cross-reference this week's sleep, alcohol and stress to find your own pattern.
Reflection: when a black-box recovery score conflicts with how your body actually feels, which do you trust — and why?
SUB · Continuous Glucose Monitoring / Metabolism
CGM — Revolution for Diabetes, Experiment for the Healthy
A post-meal glucose rise is normal physiology, not disease
Bottom Line
Continuous glucose monitoring (CGM) is revolutionary in diabetes management; used on healthy people, the evidence that "personalized glucose response" improves weight or longevity is still thin. In a healthy person, a post-meal glucose rise is normal physiology, not a disease.
Evidence Grade
RCT + cohort: CGM lowers HbA1c and reduces hypoglycemia in type 1 and insulin-treated type 2 diabetes (DIAMOND 2017, JAMA). The personalized-nutrition PREDICT study (Berry 2020, Nat Med) confirmed that the same food produces widely different glucose responses across people, but "healthy people using CGM to adjust diet improves hard outcomes" currently lacks evidence.
Science + Mechanism
CGM measures interstitial glucose, which lags blood by ~5–10 minutes, with larger error during rapid changes or hypoglycemia. Its value is turning "a few finger-stick readings a day" into a continuous curve — revealing post-meal peaks, nocturnal lows and the dawn phenomenon. A healthy person eating a bowl of rice, rising to 7–8 mmol/L and settling back, is insulin working normally. Treating every "peak" as an enemy is applying a diabetes frame to a healthy body.
Protocol
If diabetic: use as prescribed; the core metric is Time in Range (TIR) >70% and fewer lows, not fixating on single points
If healthy and curious: a 2-week "know yourself" experiment is fine, but don't pathologize a normal post-meal peak
Before buying a CGM, do the well-proven things: pair starches with fiber/protein, take a 10–15 minute walk after meals (markedly blunts the peak), keep regular sleep — no device needed
Read the shape, not the point: peak height + speed of return matters more than any one reading
For Women + Common Myths
Insulin sensitivity fluctuates across the cycle: the luteal phase is relatively more insulin-resistant, so the same meal may produce a higher curve — this is normal. CGM use in gestational diabetes (GDM) is accumulating evidence but needs obstetric guidance; don't self-interpret.
Myth 1: "The flatter the glucose curve, the healthier" — healthy people fluctuate; chasing flatness slides toward restrictive disordered eating.
Myth 2: "The CGM number is the whole metabolic truth" — glucose is one signal; ApoB, visceral fat and insulin resistance are harder metabolic markers.
Try This Week + Reflection
THIS WEEK
No CGM required: take a 10–15 minute walk after each main meal this week — a repeatedly proven way to blunt the peak — and make "move after eating" your default.
Reflection: why does a single quantifiable metric (glucose) so easily make us ignore the more important whole that isn't easy to measure?
SUB · Health Apps / Digital Therapeutics
Health Apps — Most Are Unproven, a Few Are Real Medicine
Knowing how to pick is what makes them useful
Bottom Line
The hundreds of thousands of health apps in the stores mostly have no clinical evidence; but a small class of "digital therapeutics" (e.g. digital CBT for insomnia) is RCT-backed and can even be first-line. Knowing how to pick is what makes them useful.
Evidence Grade
RCT: digital CBT for insomnia (e.g. Sleepio) significantly improves insomnia across multiple RCTs (Espie 2012/2019) and is guideline-recommended as first-line. The flip side: reviews of the vast app pool find very few have peer-reviewed evidence (Larsen 2019, npj Digit Med). Meditation apps (Headspace/Calm) have modest RCT evidence for stress.
Science + Mechanism
The logic of a digital therapeutic (DTx) is to standardize and scale an already-validated evidence-based protocol (like CBT) and deliver it to everyone's phone — it delivers a therapy, not just a "reminder feature." That is the essential difference from an ordinary tracking app: the former has a defined mechanism of action and clinical endpoints; the latter is often a pretty-interfaced step counter plus marketing copy.
Protocol
Demand evidence and credentials: prefer products with a published RCT and regulatory clearance (FDA/NICE, etc.), not just downloads and ratings
For insomnia, choose digital CBT-i first: strong evidence, no drug side effects, a reasonable long-term first-line try
Meditation / mindfulness: pick one with research backing; consistency matters more than which brand
Check privacy: health data is sensitive — see whether it gets packaged and sold; with a free product, you are often the product
For Women + Common Myths
The "safe window" prediction of period/fertility apps is not reliable contraception; even an FDA-cleared one (Natural Cycles) has a non-trivial real-world failure rate. Treating a cycle app as your only contraceptive is high-risk.
Myth 1: "Hospital-grade / AI-powered" marketing ≠ evidence — buzzwords are not clinical validation.
Myth 2: "Five-star reviews mean it works" — ratings reflect experience, not efficacy; whether a controlled trial exists is the dividing line.
Try This Week + Reflection
THIS WEEK
Inventory the health apps on your phone, pick the one you rely on most, and search "does it have a published randomized controlled trial?" If not, don't hand it your health decisions.
Reflection: digital therapeutics can scale evidence-based psychotherapy to reach people — what does that mean for access to care, and where are its limits?
SUB · AI Diagnosis / Human-AI Collaboration
AI Diagnosis — Sharp at Narrow Tasks, Not Your GP
Use it to assist, not to replace the doctor — keep a human in the loop
Bottom Line
On specific narrow tasks (e.g. image recognition), AI already matches or exceeds experts; but consumer-facing symptom checkers and chatbot self-diagnosis are unreliable. Use it to assist, not to replace the doctor — keep a human in the loop.
Evidence Grade
Imaging RCT/validation: deep learning reaches expert level on narrow tasks like diabetic retinopathy (Gulshan 2016, JAMA) and mammography screening (McKinney 2020, Nature). The flip side: symptom checkers have low accuracy — the correct diagnosis is in the top three only about half the time (Semigran 2015, BMJ). Large language models (as of 2025) are fluent but confidently fabricate (hallucinate) and carry no medical accountability.
Science + Mechanism
Medical AI's strength is pattern recognition over massive labeled data: one retinal image, one pathology slide — a clear task boundary and a definite answer, and AI is solid. But it is narrow AI — put it in an open scenario of rare disease, interacting systems, and history-taking with common sense, and it stumbles. A subtler risk is automation bias: once a human sees "the AI says so," they stop questioning and treat assistance as a verdict.
Protocol
Trust "narrow," not "everything": trust cleared, workflow-embedded specialized AI (retinal/skin screening); be skeptical of general "cure-all" self-diagnosis
Treat AI as an assistant, not a doctor: use it to decode lab-report terms and organize a list of questions, not to self-diagnose or stop medication
Verify every AI health conclusion: check against authoritative guidelines or a doctor — especially for drugs, doses, and "do I need to see someone?"
Keep human judgment: AI gives direction; the final diagnosis and decision belong to an accountable clinician
For Women + Common Myths
Training-data bias can hurt you: many models under-represent women and darker-skinned people, so they perform worse in those groups (e.g. dermatology AI is weaker on darker skin). When using an AI health tool, notice whether it was validated on "people like you."
Myth 1: "AI is more accurate than doctors" — true only on a few narrow tasks; it fails on edge cases, rare disease, and situations needing human communication.
Myth 2: "Whatever the AI says is right" — large models confidently invent evidence that doesn't exist; treat their output as a draft, not a verdict.
Try This Week + Reflection
THIS WEEK
Before your next appointment, use AI to organize your symptoms and questions into a list to bring in — let AI assist and the doctor diagnose, and feel how the "human-AI collaboration" division of labor should work.
Reflection: as an AI practitioner, how would you design a medical-AI interface that keeps humans skeptical rather than captured by automation bias?
Deeper Questions
① Digital health tools let us "measure" more — but are we actually healthier?
Measuring ≠ changing behavior ≠ improving outcomes. Wearables and CGM explode the data, but most RCTs show that simply giving people data produces limited long-term behavior change. What works is wiring data into an intervention that guides action (like digital CBT). A tool's value isn't in "seeing," but in whether it drove the right next step.
② Why does a "quantifiable single metric" so easily hijack our attention?
Because numbers confer a sense of certainty and control. A glucose curve or sleep score is concrete, comparable, shareable — so we optimize the visible and neglect the invisible-but-more-important (social connection, meaning, visceral fat). It's "looking for keys under the streetlight" — searching where the light is, not where the keys fell. Beware letting the easily measured hijack the truly important.
③ If medical AI is superhuman on narrow tasks, why is clinical deployment slow?
High accuracy is just the entry ticket. Deployment still stalls on: performance drop from data distribution shift across hospitals, liability (who's responsible when it errs), integration with existing workflow, regulatory approval, and the new misdiagnosis mode from automation bias. The technical ceiling stopped being the bottleneck long ago — the social-institutional-trust fit is.
④ "Data as the product" — is the privacy cost of health data routinely underestimated?
Free apps often price your health data as the payment. Once heart rate, menstruation, location and sleep are packaged, they can feed insurance pricing and targeted ads, and a leak is irreversible. When choosing a tool, "how does it make money" matters as much as "is it accurate" — especially for women's physiological data.
⑤ For someone pursuing the "AI-augmented individual," what's the right stance on digital health?
Treat AI and devices as a lever that amplifies judgment, not a crutch that replaces it: let narrow AI handle the pattern recognition it's good at, and reserve scarce human attention for integration, skepticism and decision. Be neither a slave to the data (orthosomnia) nor a blind believer in the tech (automation bias) — keep a "human-in-the-loop, accountable" collaborative structure.