Day 72 · 2026.07.31

When the Evidence Doesn't Support the Decision You Already Made

Topic: Decision Under Inconvenient Evidence·4 principles
"Irrefutability is not a virtue of a theory (as people often think) but a vice." — Karl Popper, 1963

Through 2025–2026, half the Western world legislated phone bans in schools. Over the same window, four rigorous evaluations came back. Norway, national registry data (Abrahamsson, Journal of Human Resources, 2026): no effect on the full sample, but a significant effect for girls. England, 30 schools compared (Goodyear et al., The Lancet Regional Health – Europe, 2025): restrictive policies showed no association with better mental wellbeing, and did not reduce students' total daily phone use. United States, national lockable-pouch study (Allcott, Dee, Gentzkow et al., NBER 2026, 43,000+ secondary schools): average test-score effects near zero; wellbeing fell in year one and turned positive in later years. Florida, student-level panel (Figlio & Özek, NBER 2025): suspensions rose in year one, concentrated among Black students, and test scores improved significantly in year two.

This is not "phone bans don't work," and it is not "they do." It is a decision already made and announced publicly, colliding with a messy body of evidence — which happens in a large company every quarter. Four moves: separate the two kinds of reason; track the dose, not the label; stay falsifiable after you've announced; read the source when someone cites a study at you.

PRINCIPLE 01

Worth doing ≠ proven to work Two Kinds of Reason, Two Kinds of Credit

two kinds of reasonborrowed credibilityfalsifiability
There are only two kinds of reason to back a policy: a values reason ("we believe classrooms shouldn't have phones, and we'll pay the cost") and an evidence reason ("research shows it improves mental health"). Both are legitimate, but dressing a values reason up as an evidence reason means that when the evidence turns, the policy takes your credibility down with it — when it could have stood perfectly well on the values reason alone.
"A theory which is not refutable by any conceivable event is non-scientific. Irrefutability is not a virtue of a theory (as people often think) but a vice." — Karl Popper, Conjectures and Refutations, ch. 1
Situation: at an all-hands, your VP announces "AI coding assistants are mandatory across the org — target is 30% efficiency in six months," and asks you to carry it to your team.
✗ Both of these hurt you

"Research shows AI-assisted coding delivers 30%, so we're adopting it." — you have hung a company strategy bet on a causal claim you never verified; when the number doesn't land, you are the one who falls.
"There's no evidence for that number, so I'll reserve judgment." — pure resistance with no alternative. It reads as not backing the company's direction, and the line gets delivered anyway, by someone else.

✓ Try this: pull the two reasons apart on the spot

"Two separate things. Why we're doing it: this is a company strategy bet — a team that isn't using AI in three years won't be competitive. That judgment doesn't need a study to authorize it. What we actually know: there is no reliable evidence for 30%. So we adopt on the strategic reason, and we treat 30% as a hypothesis to be tested, not a known fact. We'll know in six months, because I've already written down what counts as working. Until then, anyone claiming it works or doesn't is not citing evidence."

  • For the change I'm pushing, is the real reason a values reason or an evidence reason? Which one am I giving in public?
  • If a strong study said tomorrow that it doesn't work, would my policy still stand? If it would, why did I reach for evidence to back it in the first place?
  • That number I keep quoting — have I read the source, or did I lift it off someone's slide?
  • Borrowing credibility from evidence: an evidence reason sounds like a law of nature made the decision for you, and it hides the decision-maker — so people reach for it instinctively. The price is that you also give up your room to change your mind.
  • The reverse error: "there's no evidence for it" gets used to kill every new attempt. Absence of evidence isn't evidence of absence — what you want there is a values reason plus an evaluable pilot.

Do: take one change you're pushing and write two lines. "Even if this were proven ineffective, we'd still do it, because ___." "If ___, we should stop." Can't write the first line? It's being propped up by borrowed evidence. Can't write the second? It isn't falsifiable.

Reflect: when did you last change a publicly stated position because of evidence?

PRINCIPLE 02

You implement a dose, not a label A Zero Result on a Label Tells You Nothing

label ≠ interventiondose ladderaverages hide distributions
"Phone ban" isn't one thing; it's a spectrum of intensity, from "silent mode in class" to "handed in at the door, unreachable all day." In the Norwegian data the benefit tracks the intensity. England is the sharper finding: restrictive school policies did not reduce students' total daily phone use — the label changed, the dose didn't. When you evaluate a label instead of a dose, a zero result tells you nothing.
"Until new behaviors are rooted in social norms and shared values, they are subject to degradation as soon as the pressure for change is removed." — John P. Kotter, Leading Change, ch. 1, on the error of "not anchoring changes firmly in the corporate culture"
Five doses under one name
Level 1written in the handbook
Level 2asked for verbally
Level 3a gate in the process
Level 4impossible to bypass
Level 5old path removed
Norway's gains sit in the two right-hand levels. You think you're evaluating the measure; you're evaluating its weakest implementation.
Situation: your boss says, "We've had the RFC process for six months and engineering efficiency hasn't moved. Should we drop it?"
✗ Both directions are wrong

"Let's give it another quarter." — no new information, just delay.
"Let's drop it, clearly it doesn't work." — you're rejecting an intervention using a dose that was never actually administered.

✓ Try this: report the dose before you argue about the effect

"First, let's be clear what we'd be dropping. Of 41 designs this half, 9 went through an RFC with a second signer. The other 32 were documents written after the code had already merged. We didn't run an RFC process; we ran an RFC folder.
To judge whether it works, it has to reach that dose first: for the next two months, no cross-team interface change merges without an RFC — that gate only, roughly 20% of designs. The zero we get in two months is a real zero; and if it's still zero, I'll be the first to propose dropping it."

  • When I say "we've already rolled out X," can I quote a dose number — coverage, compliance rate, times it actually blocked something?
  • Does this measure have a way around it? What share of cases take it?
  • Before I call it ineffective: has it run at full strength for one complete cycle?
  • The most common cause of a null effect isn't that the intervention failed — it's that the intervention never happened. A policy gets diluted three times over: translated into process, waived under deadline, forgotten at the next manager handover.
  • Averages hide distributions: Norway's full-sample effect was zero — significant for girls, zero for boys. The organizational analogue is everywhere. A policy applied "equally to everyone" (return-to-office mandates, a wider on-call rotation, meetings pushed into the evening) can average out to zero while its costs land on the people carrying more caregiving load, who are disproportionately women. Break out a subgroup whenever you report an average.

Do: measure the real dose of your most important "already rolled out" measure: coverage, bypass rate, number of decisions it actually changed. Numbers, not impressions.

Reflect: how many of your team's measures are in the state of "name still there, dose zero"? They occupy the slot marked "we're already doing this," which is exactly what blocks actually doing it.

PRINCIPLE 03

Staying falsifiable after you've announced Write the Failure Criteria While You Can Still Be Neutral

cognitive dissonancecriteria up fronteffects have a time shape
Once a decision is public, the cost of changing your mind isn't just admitting you judged wrong — it's admitting you had the whole team work for six months for nothing. That isn't a willpower problem, it's a structural one: a public position automatically recruits you to defend it. There is one countermeasure: at the moment you announce, write down what counts as failure and hand it to someone else to hold — because that version of you is the last one that can still be neutral.
"A man with a conviction is a hard man to change. Tell him you disagree and he turns away. Show him facts or figures and he questions your sources. Appeal to logic and he fails to see your point." — Leon Festinger et al., When Prophecy Fails (1956), opening of ch. 1

Criteria also need a time axis. In the pouch study wellbeing fell in year one and turned positive later; in Florida, test scores only improved in year two. Real organizational change almost always gets worse before it gets better — or just stays worse. Look at it without separate time windows and you'll pass final judgment at the bottom of the trough.

Situation: six months ago you pushed through a mandatory AI-assisted review gate on cross-team interface changes. Q2 data: defect escape rate flat, average review time up 20%. You present next week.
✗ Both directions are wrong

"It's early, the tooling is still improving — I'd give it two more quarters." — each clause is defensible; together they're a cheque that never comes due.
"I was wrong, let's kill it." — if you're at the bottom of the transition trough, you're cutting something that would have worked, at its lowest point — and next time nobody will dare push anything.

✓ Try this: read out the criteria you wrote six months ago

"Going by the criteria we set then: at six months we look at adoption and review time (transition metrics); at twelve months, defect escape rate (the outcome metric). The six-month bar was 'coverage >80% and review-time increase <30%' — we're at 87% and +20%, so that gate passes. The outcome criterion isn't due yet, so I'm not drawing a conclusion on it today. One cost we didn't anticipate: cross-timezone PRs wait about 4 hours longer. If defect escape hasn't dropped 15% at twelve months, then per what we wrote, we remove the gate and keep advisory mode only."

✓ If someone says "you just won't admit it isn't working"

"Possibly. Which is why the criteria aren't mine as of today — they were set six months ago, they're on page 3 of the proposal, and Dana signed off on them. The only freedom I have today is to read them out."

  • When I announced this change, did I also write down "how much / by when / what we do if it misses"?
  • Do the criteria separate transition metrics from outcome metrics, with different due dates?
  • Has a second person signed the criteria, or do they exist only in my head — which is to say, not at all?
  • Rewriting the criteria the moment they come due: when the data looks bad, the tempting move isn't stubbornness — it's "redefining the metric." That's more dangerous than stubbornness, because it looks like being data-driven. One defense only: any change of definition must leave a trace, stating what the number was before the change.

Do: dig up your exact words from six months ago (email, slide, meeting notes) and compare them line by line with what you say now, to see whether the goal quietly moved. If it did, say so unprompted at the next review.

Reflect: in the last two years, did you have a project that "quietly faded" rather than being declared a failure? Who should have called it?

PRINCIPLE 04

When someone cites a study at you Both Sides Are Quoting Real Papers

read the sourceboth sides quote real papersfactual vs values disagreement
On phone bans, both sides are quoting real papers: supporters cite Norway's girls and Florida's second year, opponents cite England's null result and the pouch study's average. Nobody is fabricating; each side reports only the layer that suits them. Which is why "studies show" carries almost no information in a meeting — the information is in the questions below.
"Thinking like a scientist... means being actively open-minded. It requires searching for reasons why we might be wrong—not for reasons why we must be right—and revising our views based on what we learn." — Adam Grant, Think Again, ch. 1
Situation: in a cross-team review, a peer kills your proposal with an industry report — "the analysts say 70% of migrations like this fail."
✗ Both of these lose

"That report's methodology is flawed." — you haven't read it. It's a bluff, and it turns the argument personal.
Say nothing and go find evidence later. — the decision in that room gets made today.

✓ Try this: break it into three checkable factual questions

"I'd like to use that number, but three things first. How does the 70% define failure — does over-budget count? What migrations are in the sample — does it include same-stack moves like ours? What's the time window? I don't have those answers, so I won't use it to support my proposal, and I'd suggest we not use it to kill it either — I'll read the source this week and bring three answers next week."
If they call that a delay: "I'm narrowing the disagreement. We don't disagree about whether to migrate; we disagree about whether that 70% applies to us. That's a factual disagreement solvable in a week, not a values disagreement."

  • Level: is the figure from the full sample or a subgroup? "Significant in a subgroup, zero overall" is the most common way to mislead with true statements.
  • Magnitude and dose: how large is the effect ("significant" is a statistics word, not a business word)? Is its intervention intensity the same thing you're proposing to do?
  • Time window and cost side: is this year one or year three? Are costs reported at all — the Florida study reported a short-run rise in suspensions concentrated among Black students, which rarely appears on the slide that cites it.
  • Has anyone in this room read the source, including me? And is this a factual disagreement or a values one?
  • Two cheap postures: citing without checking ("studies show"), and dismissing without checking ("academic research doesn't apply to industry"). Neither carries information.
  • The women's angle: asking to check a source in a meeting reads as "rigorous" from a man and more easily as "not a team player / slowing us down" from a woman. The lower-friction phrasing is to frame the move as progress rather than a brake: not "I disagree with that number," but "I'll go verify it and bring answers next week, so we don't have to argue about it again." The trade-off worth naming: this hands you extra work — a real cost, and not one you should be carrying alone. It's the cheapest available path in the current environment, not a fair arrangement.

Do: find a decision your team is making on the basis of a number "everyone knows," and chase it to its original source. If you can't find one, put that fact in the notes of the next review.

Reflect: if a strong study tomorrow overturned your firmest technical conviction, what would you doubt first — the study, or yourself?

Going deeper

Is "wait for the evidence" realistic in a fast-moving industry?
Mostly not — it never arrives, and waiting has its own cost. The real first option is to build it as an evaluable pilot: narrow scope, full dose, criteria fixed in writing, a due date. "Wait" is only right when the action is irreversible and very expensive — layoffs, external commitments, deleting data. The dividing line is reversibility, not a preference for speed.
My boss already announced it and I'm not the decision-maker. What can I do?
You can't change the decision, but you can almost always change the record: add one line to the notes — "this decision rests on a strategic judgment, not on a verified effect." It challenges nobody, yet it fixes the terms. Then claim the job of writing the criteria, because whoever writes the criteria defines what counts as success. A year later, if it worked you're the reliable operator; if it didn't, you're the only person in the room with clean data.
Zero overall, significant in a subgroup — which do I decide on?
First ask where the zero came from: genuinely no effect, or opposing effects cancelling out (some people gain, some lose)? In the second case, deciding on the average means systematically sacrificing the people who lose — and you won't see them in the report. The practice: for any org-wide policy, name two or three subgroups you have reason to worry about in advance, and report them separately afterwards — hunting for a significant subgroup after the fact is just another kind of fabrication.
Doesn't all this make me look indecisive?
It does, if you only do the first half. "I'm not sure" with nothing after it reads as having no spine. The complete posture is "I'm not sure + I'm doing it + here are my criteria" — the uncertainty is about the effect, the commitment to act doesn't waver. Over time, the least credible person isn't the one who changes their mind often, it's the one who has never changed it for any evidence at all.