Meta-Knowledge: One Dataset, Two Universes

July 30, 2026 · Cross-Disciplinary Core Concepts
Day 72
Measurement Methodology Effect Sizes Statistical Inference Reading Evidence

The Construct Ladder

Exposure and outcome each have rungs
Stand one rung off and the conclusion flips
Core Insight

In the fight over adolescents and screens, one side calculated that digital technology use explains 0.4% of the variance in well-being, and casually filed it alongside eating potatoes. The other side calculated correlations of 0.15 to 0.22 between social media and depressive symptoms in girls, and called it a public health crisis. Both sides used the same data. The difference isn't in the arithmetic; it's in which rung of two ladders each side is standing on. Exposure can narrow all the way from "digital technology use" down to one specific behavior on one specific platform. Outcome can harden all the way from "self-rated well-being" up to self-harm hospitalization and suicide. A construct with a fixed name is in fact a container that stretches.

Mechanism

Each ladder has its own logic. On the exposure side the logic is dilution: bundle several behaviors into one composite index and, if only one of them relates to the outcome, the correlation gets diluted in proportion by the irrelevant components. Watching television, doing homework on a tablet, and scrolling short videos at midnight all sit inside the single rung called "screen time," and the signal gets averaged away. Narrow one rung and the effect surfaces — but sample size and generalizability shrink with it. On the outcome side the logic is trade-off: softer outcomes (self-report scales) have high base rates and ample statistical power, but carry reporting bias and whatever mood the respondent was in that day. Harder outcomes (hospitalization, death) are almost immune to self-report contamination, but events are so rare that the confidence interval can accommodate both "double the risk" and "no effect at all." So a position need only pick one rung on each ladder to arrive legitimately at the magnitude it wanted.

▸ Two ladders, four positions
EXPOSURE (broad → narrow) OUTCOME (soft → hard) Digital technology use Screen time Social media One behavior, one app Self-rated well-being Reported symptoms Clinical diagnosis Self-harm admission Death by suicide
Broad × soft → 0.4% of variance Narrow × symptom scale → r ≈ 0.15
Same data. Move one rung on each ladder and the magnitude shifts by an order of magnitude
Counterintuitive Example

The analysis that produced the potato comparison was not sloppy work. It ran every one of the thousands of analytic specifications a researcher could have chosen and looked at the whole distribution of effects rather than reporting a favorite — which is precisely the rigorous defense against cherry-picking. The effect came out tiny exactly because exposure was pinned to the broadest rung. Within the same data, narrow exposure to social media and the population to girls, and the correlation climbs to around 0.15. Both findings replicate. What actually conflicts isn't the data — it's two different claims published under one headline.

Cross-Disciplinary Transfer

Composite endpoints in cardiovascular trials have the same structure: bundle myocardial infarction, stroke, and repeat revascularization into one measure and the positive result is often driven by the softest component. In nutrition, "red meat" merges processed and unprocessed into a single rung; separate them and the sign can flip. Distributed systems do it with "latency" — mean latency, p99, and the p99 of one specific endpoint are three rungs on three ladders, and swapping rungs during capacity planning shifts the number by orders of magnitude.

For BigCat

When evaluating AI coding tools, "engineering productivity" is the stretchable container par excellence. Go down a rung: are you measuring feature lead time, code review turnaround, or the time to write a single function? Same on the outcome side — self-reported satisfaction, merge cycle time, production defect rate, maintainability a year out: hardness increases as measurability decreases. Fix which two rungs you are standing on before you look at the numbers. Reverse that order and you are choosing rungs to fit the numbers.

Question

Take your team's most-quoted metric and move both its exposure and its outcome one rung harder. How much effect survives? If your instinct says "about the same," which rung is that instinct itself standing on?

Nominal vs. Analytic Sample

The big N in the abstract isn't the N behind the claim
How sample size decays down the analysis chain
Core Insight

The "355,358 adolescents" opening an abstract almost never equals the number of people supporting the sentence that follows. Sample size decays down the analysis chain: total recruited → those asked about this exposure → those also measured on this outcome → those with no missing covariates → those inside the subgroup you care about. By the last rung, the number that made the headline is often carried by fewer than eight thousand people. The big N is printed in the abstract, the small n is buried in a supplementary table, and the reader's confidence is anchored on the former.

Mechanism

The decay has three sources. First, variable availability: different cohorts ask different questions, so after datasets are pooled, any single analysis can only use the subset that was asked that particular item. Second, listwise deletion: each additional covariate drops another batch of incomplete cases, and five or six of them routinely cost a third of the sample — while missingness is rarely random, since people who skip the income question differ from those who don't, damaging power and unbiasedness at once. Third, subgroup splitting: stratify by sex, age band, or usage intensity and each cut halves the denominator. Stack all three and you get the triple jump. The danger runs both ways: with enormous n almost any faint association reaches significance, so significance stops carrying information; once you cut to a subgroup, n shrinks and the interval widens enough to accommodate anything. Both situations can coexist in one paper, and the abstract reports only the large number.

Counterintuitive Example

Clinical trials had the same disease until journals mandated a flow diagram — screened, randomized, lost to follow-up, analyzed — as a required figure, spelling out how many people fell out at each step. That requirement entered the standards precisely because, without it, readers systematically overestimated the thickness of the evidence. Observational research and in-house corporate analysis mostly still have no such gate. In your own A/B test report, the number exposed, the number that actually triggered the experiment logic, and the number entering the statistical test are three different figures — and usually only the first is reported.

Cross-Disciplinary Transfer

In data engineering it shows up as row collapse after a join: several tables of a million rows each yield tens of thousands once joined, while the dashboard still says "based on millions of records." In epidemiology it is attrition bias — the people who come back for follow-up are healthier. In machine learning it is the effective size of an eval set: a benchmark advertises ten thousand items, but the capability subclass you actually care about has a few dozen, and leaderboard flips usually happen inside those few dozen.

For BigCat

Add a flow diagram to your own analyses, even if it is only four lines: total volume entering the system, records with complete fields, records passing quality filters, records behind the conclusion. Writing those four lines costs almost nothing, and it often convinces you before it convinces anyone else — more than once you will find the conclusion resting on a sample you would not dare cite.

Question

In your last presentation, what is the ratio between the largest number you showed and the number actually supporting your core conclusion? If they differ by an order of magnitude, did you mention it?

Incommensurable Effect Sizes

One association, five legal phrasings, five emotions
Choosing the unit is choosing the position
Core Insight

The same association can be legitimately written five ways: explains 0.4% of the variance; correlation of 0.06; odds ratio of 1.3; 30% higher risk in the heavy-use group; a difference of 0.15 standard deviations. The entire emotional range from "negligible" to "major public health problem" is available without altering a single data point. This isn't fraud, it's unit conversion — and since most readers have no habit of converting units in their heads, the choice of unit quietly becomes the choice of position.

Mechanism

A few conversions are worth memorizing. A correlation must be squared to become a variance proportion: r = 0.2 explains only 4%, which is why "explains just a few percent" is born looking trivial. An odds ratio approximates a risk ratio when the outcome is rare and badly overstates it when the outcome is common — the same data can yield an odds ratio of 3.0 for a risk ratio of 1.5. Relative risk detached from an absolute base rate cannot be weighed at all: a 30% increase on a base rate of one in a thousand is an absolute increment of three in ten thousand. Standardized differences depend on within-group standard deviation, so the more homogeneous the population, the smaller the denominator and the larger the number — the same effect is natively more impressive in a pre-screened sample. These quantities cannot be compared across types, yet printed on one page they look perfectly comparable.

Counterintuitive Example

Recast r = 0.15 into a different display — what proportion of the high group versus the low group lands on the positive side of the outcome — and it corresponds to roughly 57% versus 43%. The same "explains 2% of variance" number suddenly looks substantial. The reverse case is more famous: aspirin's effect on preventing heart attack, expressed as a correlation, is about 0.03, squarely in the range textbooks call negligible — and it was enough to rewrite clinical guidelines, because the outcome was death and the intervention was nearly free. Effect size never determines importance on its own; it has to be multiplied by the severity of the outcome and divided by the cost of the intervention.

Cross-Disciplinary Transfer

Under severe class imbalance, accuracy, F1, and AUC tell three different stories, and which one you report largely determines the reader's impression. Availability of 99.9% sounds close to perfect; converted to 43 minutes of downtime a month, less so. Any field where one reality admits several legitimate measures grows its own craft of choosing the unit to choose the position.

For BigCat

When defining internal metrics, impose a hard rule: every effect is reported in at least two units, one relative and one absolute. "Conversion up 12%" must sit next to "three more per thousand requests." The rule's greatest value isn't defending against other people's spin — it's preventing your own unconscious switching between units. Relative for wins, absolute for costs is the most common and least noticeable drift there is.

Question

Last time you said "improved by X percent," what was the absolute increment? If you had led with the absolute number, would the project still have been approved?

Difference in Significance ≠ Significant Difference

The most common broken step in subgroup drill-down
Two verdicts against zero are not a verdict about each other
Core Insight

A paper reports "a significant association in girls, none in boys," and the reader immediately hears "there is a sex difference." That step does not hold statistically. Comparing each group against zero yields two independent conclusions; to claim the two groups differ from each other, you must test the interaction term directly. Skipping that step reads as seamless — precisely because the conclusion sounds like exactly what the data is obviously saying.

Mechanism

Significance is effect size divided by standard error, and standard error is set by subsample size. Split a sample in half and both subgroups lose power, so p = 0.04 in one and p = 0.09 in the other can correspond to nearly identical effect sizes — the difference comes from their respective noise, not from reality. Testing "the difference between groups" requires the standard error of the difference, which is roughly 1.4 times that of a single group; in practice, detecting a between-group difference of the same magnitude with the same confidence needs about four times the sample. Eyeballing confidence intervals is equally unreliable: two 95% intervals can overlap visibly while the difference is still significant, and the seemingly safe "they don't overlap, so they differ" fails when the two groups have unequal variances. The only reliable move is to treat the difference itself as the parameter being estimated.

▸ Illustrative: one significant, one not, difference not significant
effect = 0 Girls r=0.17, p=0.003 → significant Boys r=0.08, p=0.09 → not significant Difference 0.09, interval spans 0
Girls Boys Difference
Illustrative figures. The two groups reach opposite verdicts, yet the claim "the groups differ" spans zero comfortably
Counterintuitive Example

This is not a rare slip. A systematic audit of more than five hundred neuroscience papers found that over a hundred and fifty of them made an argument that required an interaction test, and roughly half of those jumped straight from "significant in one group, not in the other" to "the groups differ." It happened in the most rigorously peer-reviewed journals, which tells you the problem is not carelessness but that this inference natively feels correct. Its mirror image: slice the data into enough subgroups and one cell will always be significant, and that cell goes into the abstract.

Cross-Disciplinary Transfer

Clinical trials wrote rules for this: a subgroup effect discovered after the fact may only be a hypothesis, and must be pre-specified and pass an interaction test before it enters the conclusions. A/B test drill-down works the same way — "significant in one country" is almost always a product of the number of slices. Machine learning ablations offend routinely: module A beats baseline significantly and module B does not, which does not mean A beats B; that claim requires comparing the two directly.

For BigCat

Your team has almost certainly discussed whether AI tools help senior engineers differently from juniors. To hold that as a conclusion you must put an interaction term in one model and estimate in advance whether the sample supports it — most teams' experiments are only large enough to answer "does it work overall," nowhere near large enough for "who does it work better for." The honest statement is "the data cannot yet distinguish the two groups," not "we saw no effect in the senior group."

Question

Which "works for group A, not group B" does your team currently believe? Has anyone tested the interaction? If not, is it evidence — or a hypothesis that has been used as a conclusion for a long time?

Four questions before you read a number

Before any rate, correlation, or percentage enters your decision, locate four things:

  1. Which layer of exposure is this?
  2. Which layer of outcome is this?
  3. Which statistic is this?
  4. Which denominator is this — nominal sample or analytic sample?

Only once all four are answerable does it become worth arguing about whether the number is big. Day 70 asked how strict the ruler is; this issue asks where in the coordinate system you are standing. Together they make one complete reading.