Is failure a defect to be eliminated at any cost, or the mechanism by which a system perceives reality, calibrates itself, and even grows stronger? These four books answer the same counterintuitive question from four different levels.
2026 · Book Recommendations · Issue 43
Our instinct is to treat failure as an endpoint to be eliminated. Each of these books inverts it from a different level. Taleb argues that a whole class of systems doesn't merely survive volatility and stress but grows from it—"antifragility" is a third property beyond resilience. Harford argues that a complex world cannot be planned correctly in advance; success can only emerge through the evolutionary trial-and-error of variation—survivability—selection. Petroski argues that engineering's true limits are revealed only by failure: a successful structure proves only that it "didn't collapse," never how far it stood from collapse. Syed argues that learning from error is not automatic—it is a cultural choice: some systems turn every failure into data, others bury it. Four mechanisms, one shared fuel: failure.
| Book | Author | Year | The one thing it makes clear |
|---|---|---|---|
| Antifragile | Nassim Nicholas Taleb | 2012 | Resilience only means "takes a hit and stays the same"; antifragility means "takes a hit and gets stronger"—a whole class of systems craves volatility rather than merely enduring it |
| Adapt | Tim Harford | 2011 | No one can plan the right answer to a complex problem in advance; the only reliable path is the trial-and-error of "dare, survive, select" |
| To Engineer Is Human | Henry Petroski | 1985 | Success only proves the structure didn't collapse, never the safety margin; a run of successes accumulates hidden risk until the next failure resets what we know |
| Black Box Thinking | Matthew Syed | 2015 | Learning from error isn't automatic—aviation treats accidents as data, medicine blames the individual, so the same mistakes recur |
Taleb begins from a gap in language: the opposite of "fragile" is not "sturdy." Sturdy and robust merely mean unchanged after a shock; but there is a third class of things that get better after a shock—and it had no name. He coined one: antifragile. Muscle grows under load, bone densifies under stress, the immune system strengthens after exposure to germs, an economy improves by weeding out the weak in small crises. These systems need a dose of stress, volatility, and randomness; wrap them in sterile comfort and they atrophy.
The key is nonlinearity. Fragility means a concave response to shocks: one big shock does far more damage than the same total delivered in many small doses. A porcelain cup shatters when dropped a meter, yet survives a thousand one-millimeter taps—its exposure to volatility is concave. Antifragility is the reverse—a convex response: the cumulative gain from many small, controllable setbacks exceeds the gain from dead calm. This is why Taleb says he cares about "exposure," not "prediction": you can't predict where the next shock comes from, but you can choose to stand on the convex side of the curve.
This yields an operational strategy: the barbell. Don't spread resources across the "moderate risk" middle—the worst spot, with neither a safety cushion nor big upside. Instead go to both extremes: most of your resources (say 85%) extremely conservative, near-zero risk, and a small remainder in high-risk bets with huge upside but capped downside (worst case, you lose only that small slice). A single failure can't touch the core, while any windfall can change everything—you've plugged yourself into optionality: limited loss, open-ended gain. It also explains his fondness for the "Lindy effect": the longer something (a book, a technology, an idea) has lived, the longer its remaining life expectancy—surviving the shock of time is itself proof of antifragility.
Taleb's sharpest cut targets the cult of "intervention" and "stability." Artificially suppressing volatility—bailing out markets, stamping out every small fire, shielding a child from every setback—looks protective but merely converts risk from "frequent small failures" into "rare catastrophe." Suppress small fires and the forest accumulates fuel until one blaze can't be put out; this is his "iatrogenics"—harm caused by the treatment itself. Antifragile systems require the failure of their lower units to keep the whole healthy: restaurants must go bankrupt one by one for the restaurant industry to be excellent overall.
The prose is self-assured and combative; the mockery of academics, economists, and "fragilistas" sometimes overpowers the argument. The core concept is extremely strong, but the book is long and digressive, hammering the same insight repeatedly. How the "barbell" cashes out in concrete asset allocation is left largely unquantified.
The barbell maps directly onto the time allocation of an "AI super-individual." Most people spread their learning time across the middle—chasing a pile of "maybe-useful" mid-tier tools, dabbling, neither going deep nor hitting a jackpot. The antifragile arrangement: put 85% of your energy on rock-solid ground (your existing depth in distributed systems/AI, your body, your family), and stake only 10–15% on experiments with capped downside and open upside—take one of your deep interests (consciousness, complexity science) × frontier AI and make a few "at most wastes a weekend, but if it lands it's something others can't build" bets. To try next week: list three things on your plate and ask which is "convex" (failure costs only time, success is unbounded), then give your next open slot only to that—rather than learning yet another middling framework.
Harford's central diagnosis is the "God complex": however complex a problem, people carry an unshakeable certainty that their judgment is right. Yet reality is otherwise—the knowledge behind even a humble object (he cites a pencil) is dispersed across thousands of people worldwide; no single person could make one alone. When no one can hold the whole picture, "planning the right answer" is blocked at the root. The only reliable alternative is to mimic biological evolution: generate lots of variation, let the environment weed out the failures, amplify the survivors.
He compresses this into the three principles of the Soviet engineer Palchinsky: first, seek out new ideas and try new things, expecting some to fail; second, when trying something new, do it on a scale where failure is survivable; third, seek out feedback and learn from your mistakes as you go. The three interlock, none dispensable: dare-but-don't-survive (missing the second) bets it all in one shot; play-it-safe (missing the first) never progresses; try, fail, but ignore feedback and deny it (missing the third) is paying tuition for nothing.
The most counter-human of the three is the second—"survivability." Trial and error presupposes that failure isn't fatal. An entrepreneur who stakes everything, a financial system with no redundancy, an institution "too big to fail"—these have precisely forfeited the right to experiment, because they can't afford even one loss. Harford stresses repeatedly: the real skill isn't avoiding failure but designing structures that make failure cheap and recoverable—try small, try in parallel, isolate each experiment so one error can't infect the whole system.
And the third—learning from feedback—is hard because it requires being able to tell that you actually failed. In the real world, failure is routinely hidden by narrative: reword it, pick a flattering metric, blame luck, and the error "disappears." Through cases from the Iraq War to the financial crisis to aid projects, Harford shows that the larger and more self-assured an institution, the fewer its mechanisms for admitting failure—so the same error can recur, at enormous cost, for years.
The cases are wide-ranging but scattered—leaping from the Iraq War to financial derivatives to corporate innovation, leaving the reader to extract the through-line. Where "trial and error" hits its limits in moral or irreversible settings (one-shot major medical decisions, the critical windows of parenting) is treated thinly—not every failure is "recoverable."
Palchinsky's second principle is most useful when leading a team into AI adoption. Don't make the one-shot big bet of "roll out one unified AI process company-wide"—if the direction is wrong, the cost is so high that no one dares admit failure, and the error drags on buried under narrative. Instead run parallel small experiments: have 3–4 groups each use different tools/approaches on their own most painful use case, each capped at "even failure wastes only two weeks," with an agreed, quantifiable feedback metric (the third principle). To try next week: split some current "big, all-encompassing AI plan" into three mutually isolated, two-week-verdict experiments—and write down in advance "what result counts as failure"—because without a clear failure criterion, the third principle can never be executed.
Petroski is a structural engineer, and he states from inside the profession a truth outsiders find counterintuitive: engineering advances mainly through failure, not success. A bridge or an aircraft that has run safely for ten years proves only that it did not collapse under those conditions—it tells you nothing about how much safety margin remains or how close it stood to the limit. Success is silent; only failure speaks, pointing precisely at "here, this is where it fails." So every major structural disaster advances engineering knowledge more than countless successful builds.
Hence an unsettling paradox of success: a run of successes breeds the next failure. Designers see predecessors as "too conservative, too much material," and generation by generation trim redundancy, widen spans, push toward the limit—each step "proven" safe by prior success, until one generation crosses the line no one had marked, disaster strikes, and knowledge is reset. Petroski notes that major bridge collapses have historically recurred roughly every thirty years—about one generational turnover of engineers: those who lived through the last failure retire, the lesson fades, and redundancy gets eaten away again.
He also dismantles the misreading of the "factor of safety." Engineers build in several times the needed redundancy; Petroski says it should really be called the "factor of ignorance"—it measures not how safe we are, but how little we understand about material fatigue, wind-induced vibration, resonance, and unexpected loads. Redundancy is insurance bought against "what we don't know we don't know"; once success tempts us to trim it, what we're really buying is disaster. The Comet airliners breaking apart from metal fatigue at the corners of square windows, the Tacoma Narrows Bridge twisting apart in resonance with the wind—all were failure modes designers did not anticipate and could not have, forced into the textbooks only by the failures themselves.
Written in 1985, its cases center on the mechanical structures of civil engineering and aviation, not the failure modes of newer systems like software and networks. The prose is essayistic; the theoretical frame is looser than later books of its kind. It says little about "how to institutionalize the memory of failure"—it is better at diagnosis than at prescription.
"Factor of safety = factor of ignorance" is a ready-made diagnostic blade for anyone building distributed systems. A system runs two years without incident, and the team begins trimming redundancy—cutting backups, raising utilization, removing "seemingly superfluous" degradation plans—each step "proven" safe by past stability, until one traffic spike crosses the line no one marked. To try next week: pick one piece of redundancy that "hasn't caused trouble, so you're about to optimize it away," and write down which class of failure-you-don't-yet-understand it was originally insurance against. Can't articulate it? Then you truly understand it and may trim it. Can articulate it? Then what you're about to cut is the factor of ignorance. And second: turn major-incident post-mortems into an institutional memory that doesn't vanish with staff turnover, to fight Petroski's "thirty-year cycle"—don't let the lesson reset to zero the moment those who lived it retire.
Syed builds the whole book on a stark contrast of two industries: aviation vs medicine. Aircraft carry two black boxes; every accident is investigated by an independent body with no blame attached, and the conclusions are fed back into the entire industry by force—so aviation grows ever safer. Medicine is the reverse: errors are often concealed, blamed on individuals, waved past under the name "complication," lacking a mechanism of no-fault investigation. The result: the same preventable errors take lives in hospitals over and over. The difference lies not in the intelligence or goodwill of practitioners, but in how the system as a whole treats failure.
He abstracts the two attitudes into "closed loop" and "open loop." A closed loop is where error information is denied or reinterpreted away, so the loop seals shut, the system learns nothing, and it spins in place; an open loop is where failure is recorded honestly, treated as signal, and fed back—so the loop opens and the system keeps evolving. What decides your fate is not whether you err, but where the information flows after the error.
Why is the closed loop so common? Syed's answer is the psychology of "cognitive dissonance": when the facts clash with one's self-image ("I'm competent / I'm right"), people don't admit error cleanly—they unconsciously rewrite the narrative: find excuses, cherry-pick data, blame luck, disparage the evidence, so the self-image survives intact. The more expert, the more authoritative, the more one's reputation is staked on "never being wrong," the stronger this defense. So the very people who should most learn from failure are the ones who most resist admitting it. It isn't bad people covering up—it's ordinary people protecting the self.
The way out Syed calls "marginal gains": break a big goal into countless measurable small links, let each link expose its tiny failures, and optimize them one by one. Britain's cycling team turned a one-percent improvement in every detail into an overwhelming advantage exactly this way. It works precisely because it shrinks "failure" and strips its stigma—each micro-failure is neither fatal nor shameful, so people are willing to report it honestly, and the loop opens.
The argument is clear but somewhat thin—"learn from failure" is proven by example after example, and the second half feels repetitive. The aviation-vs-medicine dichotomy is simplified: medicine's recent patient-safety movement is in fact moving toward the aviation model. It says too little about "which failures shouldn't be risked at all"—it is better at encouraging the open loop than at drawing its boundaries.
The "cognitive dissonance → closed loop" chain speaks most directly to parenting. The moment a child (or an adult) errs, if what follows is blame or shame, cognitive dissonance drives her to learn to conceal, make excuses, avoid risk—turning home into a closed loop. To make it an open loop, the mechanism isn't "more encouragement" but decoupling error from self-worth: in the debrief, ask only "what does this data tell us to adjust next time," never "whose fault was it." To try next week, a family-version "black box": agree that once a week, each person shares one of their own failures that week and one thing learned from it—parents go first and model admitting error—so the child sees that owning a failure is not shameful but a respected act in this home. The same applies to team retrospectives: swap the "blame meeting" for a "no-fault investigation," and information will actually flow back into the system.
The key is "nonlinearity" and "optionality." Ask two things: (1) Would one big shock harm you far more than the same total delivered as many small setbacks? Yes = you're on the concave side, fragile. (2) Is your downside capped and your upside open? Yes = you're on the convex side, antifragile. Most people mistake themselves for "solid" when they've merely hidden the risk in a "rare but fatal" tail—that isn't resilience, it's fragility disguised as it.
Each has its own way of killing you. Missing the first (don't dare try) = long stagnation, rotting safely; missing the second (scale too large) = one fatal bet, no second chance; missing the third (ignore feedback) = paying tuition repeatedly and learning nothing. Most smart people are stuck on the third—not that they don't try, but that after failing, cognitive dissonance quietly rewrites the narrative, so "failure" vanishes in their own eyes. Honest check: when was the last judgment you explicitly admitted "this one was wrong"?
Watch two signals together. Petroski signal: is some piece of redundancy being optimized away "because nothing's gone wrong," with no one able to say what it originally guarded against?—that's likely trimming the factor of ignorance. Syed signal: after the last problem, was the debrief a "no-fault investigation, only what to fix," or "find whose fault it was"? If the latter, your system is a closed loop—error information is flowing toward blame, not improvement. Both signals red means you're personally brewing a "thirty-year-cycle" scale of failure.