DEEP READ · OFF-LIST

Superforecasting

The Art and Science of Prediction · Philip Tetlock & Dan Gardner · 2015

中文 →

In One Sentence

Tetlock spent twenty years doing something almost nobody had done—writing down thousands of judgments about the future, numbering them, and then scoring every one against what actually happened—and found two things. First, the most famous and most confident experts on television perform, over the long run, at roughly the level of chance. Second, scattered through the general public is a small group of ordinary people who beat professional intelligence analysts, consistently. What they have is not talent, access or a towering IQ, but a set of habits you can learn: break the big question into checkable small ones, start from base rates, revise in small steps, put an actual number on the answer, then go back and settle the account item by item. The real claim of the book is not that the future can be known, but that judgment is a craft that can be scored—and therefore improved.

Where It Sits

Tetlock is a professor of psychology and political science at the University of Pennsylvania; Gardner is a journalist and his co-author. The book is the sequel to Expert Political Judgment (2005), which tracked more than twenty thousand forecasts from 284 experts beginning in 1984 and produced the much-abused finding that the average expert barely beat chance. This book asks the obvious follow-up: if nobody is any good, is anybody any good? From 2011 the US intelligence community's research arm, IARPA, ran a four-year forecasting tournament. The Good Judgment Project—run by Tetlock with the psychologist Barbara Mellers, his wife and collaborator—fielded thousands of volunteers and won it, with amateurs beating rival university teams. The book is the story of that tournament and of what was excavated from it.

The Central Claims

The Core Ideas, One by One

Hedgehogs and foxes: one big idea vs. many small pieces

The image comes from a fragment of the Greek poet Archilochus—"the fox knows many things, the hedgehog one big thing"—and was used by the historian of ideas Isaiah Berlin to sort thinkers. Tetlock turned it into a measurable variable. The hedgehog expert has one grand framework that explains everything—markets always clear, empires always decline, this or that ideology must collapse—and files every new event into it. The fox improvises: bits taken from several conflicting sources, no demand for theoretical elegance, judgments held loosely.

HedgehogFox
Way of seeingOne theory that covers everythingBits from many conflicting sources
On contrary evidenceFilters it out as noise; digs in harderTakes it as signal; nudges the estimate
How they talkCrisp, decisive, great story—bookable on TV"Probably," "unless"—unsatisfying
Long-run accuracyPoor; the most famous may trail chanceClearly better

The same traits that make someone famous are the ones that make them wrong.

The data are unkind: foxes are clearly more accurate over time, and hedgehogs—especially the most celebrated ones—may do worse than chance. The mechanism is no mystery: the tighter and more powerful a framework, the more easily contrary evidence gets classified as noise, so errors cannot correct themselves and instead get pushed further out.

The sharpest line in the book concerns the inverse relationship between fame and accuracy. Media want people who will commit, who deliver the clean line—which is exactly the hedgehog's specialty. The qualities that get someone booked and the qualities that make them right are close to opposites. That alone should change how you consume commentary: when an analysis makes a tangled situation sound crisp and settled, the first thought to have is not "how clear-headed" but "that framework is filtering the evidence for them."

Calibration and resolution: two separate scores

To judge forecasts you first need something to judge them with. The book uses the Brier score—the average squared distance between the probability you gave and what happened (1 if it occurred, 0 if not). Lower is better: 0 is perfect, and 0.5 is what you get by saying "fifty-fifty" to everything. It matters because it punishes two distinct failings, which are worth separating:

The tension between them is the whole difficulty of the craft: caution keeps you calibrated, nerve makes you useful, and you need both. Superforecasters are willing to say 5% or 95%, and then work to earn it through constant small adjustments.

Buried here is a rule with wide everyday value: only a number can be wrong, and only what can be wrong can improve. The book's classic cautionary case is the Joint Chiefs' assessment to Kennedy before the 1961 Bay of Pigs invasion, which used phrasing like "a fair chance"—and afterwards nobody agreed on what probability that had meant. Vague words are two-way insurance: if it works out you were right, and if it fails you never committed. They protect the speaker at the cost of the organisation learning anything at all.

Fermi-ising: break the unanswerable into the checkable

The physicist Enrico Fermi was famous for asking students to estimate how many piano tuners work in Chicago—not because the number matters, but to train them to decompose an impossible quantity into a few quantities each of which can be roughly estimated (city population → households → share with pianos → tuning frequency → tunings one person can do in a year), then multiply back. The point is that errors partly cancel: several rough estimates pulling in different directions usually beat one confident guess.

Superforecasters do the same to current events. The book's well-known example is whether polonium would be found in Yasser Arafat's remains. Rather than argue about whether the conspiracy theory was true, they broke it into checkable pieces: what is polonium's half-life, how long had the body been buried, would traces still be detectable, which lab was testing, what have comparable tests found before, what incentives does the announcing party have. Each sub-question has a fact you can look up or a rate you can estimate, and only assembled do they yield a number.

Why it matters: most bad judgments go wrong at the moment the question is posed, because the question is too big. "Will this company make it?" is unanswerable; "can they sign three enterprise customers in nine months?" is not. The act of decomposing is itself already an improvement, because it forces the vague parts of your intuition out into small claims that can be refuted one by one.

Outside view first: base rates before specifics

This is the most portable move in the book, inherited from Kahneman and Tversky. The inside view stares at the case in front of you: how good this team is, how clever the plan, how special the circumstances. The outside view first files the case into a class and asks: "Of cases like this one, what fraction historically ended this way?" That fraction is the base rate.

Why outside first? Because the inside view is natively optimistic and easily drowned in detail. A team sitting down to estimate a project's duration talks entirely about its own plan—and almost never looks up how much the department's last twenty comparable projects overran. That dull number is usually more accurate than the whole afternoon's discussion.

The superforecaster's routine is: anchor on the base rate, then move it up or down with the specifics of this case, and move it sparingly. The order cannot be reversed—dive into the details first and you will simply select whichever data confirm the impression you already formed. The habit cures the single most expensive assumption people make, which is that this time is different. Almost always, your case is the same as their case.

Perpetual beta: small updates, and two opposite ways to die

Superforecasters do not call it once. They treat a forecast as a position that needs continuous maintenance: news arrives, the number moves from 63% to 58%, then to 61%. Behind this is the spirit of Bayesian updating—no need to run the formula, just hold on to the idea: your new view should be the old view plus the shift the new evidence warrants, and stronger evidence warrants a bigger shift.

What is worth memorising are the two opposite failure modes, equally fatal:

Frequent small revisions work precisely because they preserve the weight of all the earlier evidence: moving a little each time is a way of conceding that the latest item is one piece of a large puzzle. The tournament data bear this out—the best forecasters updated more often and in smaller increments. The habit transfers everywhere: hold an opinion as a number you can nudge, not a flag you plant.

Growth mindset and the premortem: turning errors into information

Tetlock reports that the most consistent trait of superforecasters is not intelligence but "perpetual beta"—treating their own judgment as a product that is never finished and always in need of debugging. Underneath sits Carol Dweck's growth mindset: treat ability as trainable and failure becomes data; treat it as fixed and failure becomes a verdict, to be avoided or dressed up.

Two concrete practices are worth stealing outright. The first is line-by-line review: not the vague "was I right last quarter?" (memory quietly flatters), but pulling up the numbers you actually wrote and checking each. This is also the only antidote to hindsight bias—after the fact your brain will sincerely believe it knew all along, and only the contemporaneous record can catch it out.

The second is the premortem: before starting, assume that a year from now the project has failed completely, and ask everyone to write down why. It beats asking "what are the risks?" because it changes the question from "might this go wrong" (which invites defence) to "it went wrong—how?" Once the ending is stipulated as fact, imagination switches from defending to explaining, and the quiet misgivings come out.

Teams, groupthink, and the extremizing algorithm

Putting superforecasters into collaborating teams raised scores again—provided the team does not collapse into an echo chamber. The guardrail the book offers is constructive confrontation: members are explicitly asked to attack each other's reasoning (not each other), and to ask precise questions—not "what do you think?" but "when you say 'likely,' what number is that? What would you have to see to change your mind?"

It is worth naming what groupthink actually is: not that everyone agrees, but that nobody is willing to voice dissent for the sake of keeping the peace, so the group converges on a conclusion no one has really tested. (Irving Janis's textbook case was, again, the Bay of Pigs.) Teams outperform individuals only when the inputs are genuinely independent; the moment people start reading the room, extra members amplify the error instead.

One technical detail is worth knowing because it is counterintuitive. When you pool many forecasts, a simple average comes out too timid—each person holds only part of the information, and averaging also averages everyone's hedging. The Good Judgment Project weighted the better forecasters and then applied extremizing: pushing the pooled probability further toward 0 or 1 (say 0.7 up to 0.8). The logic: if a group each holding a different piece of the puzzle all say "probably," then once the pieces are assembled the group should be more confident than any individual. (The caveat matters: this works only when the participants' information is genuinely diverse. If everyone read the same article, extremizing just magnifies the mistake.)

The black swan argument: where the method stops working

Nassim Taleb's objection deserves an honest hearing, because it identifies the method's real boundary. His position runs roughly: history is driven by a handful of unpredictable extreme events, and for those, probability estimates are not merely useless but actively dangerous, since they manufacture false security. The right response is not forecasting better but restructuring your exposure so that being wrong does not destroy you.

Tetlock's reply has two parts. First, tournament questions mostly resolve within a year, and such short- and medium-horizon questions make up the overwhelming majority of real decisions—and they demonstrably are predictable. Refusing to improve them on the grounds that long-run black swans are unforecastable is changing the subject. Second, at the deeper level the two men are not opposed: build robustness where robustness belongs, buy precision where precision belongs. Treating the two as alternatives is a false dichotomy. But Taleb's core warning stands: a well-calibrated probability is not the same thing as being prepared for the tail.

The Argument in Outline

The book runs in a straight line, in five steps. Step one, establish that there is a problem: score decades of expert forecasts and the average is dismal, while fame tends to run against accuracy—because the market for commentary rewards confidence and narrative, not correctness.

Step two, fix the instrument before discussing the skill: without a scoring rule there can be no improvement. So forecasts are forced into numbers, evaluated by Brier score, and split into calibration and resolution. This is the real foundation; every later conclusion stands on it.

Step three, run the experiment and find the people: four years, thousands of volunteers, millions of judgments. Roughly the top two per cent stay ahead, and the lead persists across years, so it is not luck.

Step four, reverse-engineer what they do: Fermi-ise the question, take the outside view first, update often and in small steps, hunt actively for disconfirming evidence, commit to numbers, and settle the account line by line afterwards—held together by the perpetual-beta mindset. Step five, show it is teachable: brief training produces measurable gains, teamwork adds another layer, and weighting plus extremizing at the aggregation stage adds one more.

So what does the book actually establish? Not that the future is knowable, but that at horizons of about a year there are real and stable differences in the quality of judgment, that these differences come mostly from learnable practices, and that they can therefore be trained. It moves forecasting out of divination and into craft—and the first discipline of the craft is not cleverness, it is bookkeeping.

Misreadings & Serious Objections

Ten Sentences

1. An unscored forecast is not a forecast; it is a performance. Experts go decades without improving because nobody writes down what they said and checks it.

2. Fame runs roughly opposite to accuracy: media reward the people who commit and tell a clean story, and those are the least accurate people in the sample.

3. Hedgehogs have one theory that explains everything; foxes assemble scraps from conflicting sources. Foxes are more accurate, because a big framework files inconvenient evidence as noise, so the error can never correct itself.

4. A good forecast carries two separate scores: calibration (of the things you call 70%, do 70% happen?) and resolution (will you leave 50%?). Someone who says "about even" to everything is perfectly calibrated and perfectly useless.

5. "A fair chance" is cover for avoiding the scoreboard. Vague wording means you were right if it happens and never committed if it doesn't—it protects the speaker while the organisation learns nothing. Only a number can be wrong, and only what can be wrong can improve.

6. Fermi-ise: split an unanswerable question into a few estimable ones and multiply back, so errors partly cancel. Most bad judgments went wrong at the moment the question was posed too broadly.

7. Outside view first: ask what fraction of comparable cases ended this way (the base rate), then adjust sparingly for the specifics. Reverse the order and you will only collect evidence that flatters your first impression. Almost always, your case is the same as their case.

8. Treat a forecast as a position under continuous maintenance and revise in small steps. Two opposite deaths are equally fatal: refusing to move because it looks like an admission, and leaping from 20% to 70% on one headline—mistaking noise for signal.

9. Review by pulling up the numbers you wrote at the time, because hindsight bias will have you sincerely believing you knew all along; and before starting, run a premortem—stipulate that a year from now it failed, and have everyone write down why. The quiet misgivings surface immediately.

10. The boundary worth remembering: this craft governs clean, roughly one-year questions, not the black swans that rewrite history—and a well-calibrated probability is not the same as being prepared for the tail. A forecast is half a decision; the other half is what you can afford to lose.