Issue 43 · Themed Reading List

Failure & Resilience

Is failure a defect to be eliminated at any cost, or the mechanism by which a system perceives reality, calibrates itself, and even grows stronger? These four books answer the same counterintuitive question from four different levels.

2026 · Book Recommendations · Issue 43

Introduction

Our instinct is to treat failure as an endpoint to be eliminated. Each of these books inverts it from a different level. Taleb argues that a whole class of systems doesn't merely survive volatility and stress but grows from it—"antifragility" is a third property beyond resilience. Harford argues that a complex world cannot be planned correctly in advance; success can only emerge through the evolutionary trial-and-error of variation—survivability—selection. Petroski argues that engineering's true limits are revealed only by failure: a successful structure proves only that it "didn't collapse," never how far it stood from collapse. Syed argues that learning from error is not automatic—it is a cultural choice: some systems turn every failure into data, others bury it. Four mechanisms, one shared fuel: failure.

The Four Books at a Glance

BookAuthorYearThe one thing it makes clear
AntifragileNassim Nicholas Taleb2012Resilience only means "takes a hit and stays the same"; antifragility means "takes a hit and gets stronger"—a whole class of systems craves volatility rather than merely enduring it
AdaptTim Harford2011No one can plan the right answer to a complex problem in advance; the only reliable path is the trial-and-error of "dare, survive, select"
To Engineer Is HumanHenry Petroski1985Success only proves the structure didn't collapse, never the safety margin; a run of successes accumulates hidden risk until the next failure resets what we know
Black Box ThinkingMatthew Syed2015Learning from error isn't automatic—aviation treats accidents as data, medicine blames the individual, so the same mistakes recur

The Four Books in Detail

Antifragile
Antifragile: Things That Gain from Disorder · Nassim Nicholas Taleb · 2012
Random House · ~519 pages
Resilience just means taking a beating and staying up. The truly scarce property is "getting stronger from the beating"—and it has a recognizable mathematical shape.
Core Insight

Taleb begins from a gap in language: the opposite of "fragile" is not "sturdy." Sturdy and robust merely mean unchanged after a shock; but there is a third class of things that get better after a shock—and it had no name. He coined one: antifragile. Muscle grows under load, bone densifies under stress, the immune system strengthens after exposure to germs, an economy improves by weeding out the weak in small crises. These systems need a dose of stress, volatility, and randomness; wrap them in sterile comfort and they atrophy.

The key is nonlinearity. Fragility means a concave response to shocks: one big shock does far more damage than the same total delivered in many small doses. A porcelain cup shatters when dropped a meter, yet survives a thousand one-millimeter taps—its exposure to volatility is concave. Antifragility is the reverse—a convex response: the cumulative gain from many small, controllable setbacks exceeds the gain from dead calm. This is why Taleb says he cares about "exposure," not "prediction": you can't predict where the next shock comes from, but you can choose to stand on the convex side of the curve.

Three responses to the same shock
payoff shock size → antifragile robust fragile concave: loss accelerates convex: gain accelerates

This yields an operational strategy: the barbell. Don't spread resources across the "moderate risk" middle—the worst spot, with neither a safety cushion nor big upside. Instead go to both extremes: most of your resources (say 85%) extremely conservative, near-zero risk, and a small remainder in high-risk bets with huge upside but capped downside (worst case, you lose only that small slice). A single failure can't touch the core, while any windfall can change everything—you've plugged yourself into optionality: limited loss, open-ended gain. It also explains his fondness for the "Lindy effect": the longer something (a book, a technology, an idea) has lived, the longer its remaining life expectancy—surviving the shock of time is itself proof of antifragility.

Taleb's sharpest cut targets the cult of "intervention" and "stability." Artificially suppressing volatility—bailing out markets, stamping out every small fire, shielding a child from every setback—looks protective but merely converts risk from "frequent small failures" into "rare catastrophe." Suppress small fires and the forest accumulates fuel until one blaze can't be put out; this is his "iatrogenics"—harm caused by the treatment itself. Antifragile systems require the failure of their lower units to keep the whole healthy: restaurants must go bankrupt one by one for the restaurant industry to be excellent overall.

Key Quotes
"Wind extinguishes a candle and energizes fire. Likewise with randomness, uncertainty, chaos: you want to use them, not hide from them. You want to be the fire and wish for the wind."
— Antifragile, Prologue
"Antifragility is beyond resilience or robustness. The resilient resists shocks and stays the same; the antifragile gets better."
— Antifragile, Prologue
Limitations

The prose is self-assured and combative; the mockery of academics, economists, and "fragilistas" sometimes overpowers the argument. The core concept is extremely strong, but the book is long and digressive, hammering the same insight repeatedly. How the "barbell" cashes out in concrete asset allocation is left largely unquantified.

Application for BigCat

The barbell maps directly onto the time allocation of an "AI super-individual." Most people spread their learning time across the middle—chasing a pile of "maybe-useful" mid-tier tools, dabbling, neither going deep nor hitting a jackpot. The antifragile arrangement: put 85% of your energy on rock-solid ground (your existing depth in distributed systems/AI, your body, your family), and stake only 10–15% on experiments with capped downside and open upside—take one of your deep interests (consciousness, complexity science) × frontier AI and make a few "at most wastes a weekend, but if it lands it's something others can't build" bets. To try next week: list three things on your plate and ask which is "convex" (failure costs only time, success is unbounded), then give your next open slot only to that—rather than learning yet another middling framework.

Adapt
Adapt: Why Success Always Starts with Failure · Tim Harford · 2011
Little, Brown · ~309 pages
In a complex world, no expert can plan the right answer in advance—the systems that improve all run on a discipline that dares to fail.
Core Insight

Harford's central diagnosis is the "God complex": however complex a problem, people carry an unshakeable certainty that their judgment is right. Yet reality is otherwise—the knowledge behind even a humble object (he cites a pencil) is dispersed across thousands of people worldwide; no single person could make one alone. When no one can hold the whole picture, "planning the right answer" is blocked at the root. The only reliable alternative is to mimic biological evolution: generate lots of variation, let the environment weed out the failures, amplify the survivors.

He compresses this into the three principles of the Soviet engineer Palchinsky: first, seek out new ideas and try new things, expecting some to fail; second, when trying something new, do it on a scale where failure is survivable; third, seek out feedback and learn from your mistakes as you go. The three interlock, none dispensable: dare-but-don't-survive (missing the second) bets it all in one shot; play-it-safe (missing the first) never progresses; try, fail, but ignore feedback and deny it (missing the third) is paying tuition for nothing.

The most counter-human of the three is the second—"survivability." Trial and error presupposes that failure isn't fatal. An entrepreneur who stakes everything, a financial system with no redundancy, an institution "too big to fail"—these have precisely forfeited the right to experiment, because they can't afford even one loss. Harford stresses repeatedly: the real skill isn't avoiding failure but designing structures that make failure cheap and recoverable—try small, try in parallel, isolate each experiment so one error can't infect the whole system.

And the third—learning from feedback—is hard because it requires being able to tell that you actually failed. In the real world, failure is routinely hidden by narrative: reword it, pick a flattering metric, blame luck, and the error "disappears." Through cases from the Iraq War to the financial crisis to aid projects, Harford shows that the larger and more self-assured an institution, the fewer its mechanisms for admitting failure—so the same error can recur, at enormous cost, for years.

Key Quotes
"First, try new things, expecting that some will fail; second, do it on a scale where failure is survivable; third, seek out feedback and learn from your mistakes as you go along."
— Adapt, Chapter 1 (the Palchinsky principles)
"Success comes through rapidly fixing our mistakes rather than getting things right first time."
— Adapt, Introduction
Limitations

The cases are wide-ranging but scattered—leaping from the Iraq War to financial derivatives to corporate innovation, leaving the reader to extract the through-line. Where "trial and error" hits its limits in moral or irreversible settings (one-shot major medical decisions, the critical windows of parenting) is treated thinly—not every failure is "recoverable."

Application for BigCat

Palchinsky's second principle is most useful when leading a team into AI adoption. Don't make the one-shot big bet of "roll out one unified AI process company-wide"—if the direction is wrong, the cost is so high that no one dares admit failure, and the error drags on buried under narrative. Instead run parallel small experiments: have 3–4 groups each use different tools/approaches on their own most painful use case, each capped at "even failure wastes only two weeks," with an agreed, quantifiable feedback metric (the third principle). To try next week: split some current "big, all-encompassing AI plan" into three mutually isolated, two-week-verdict experiments—and write down in advance "what result counts as failure"—because without a clear failure criterion, the third principle can never be executed.

To Engineer Is Human
To Engineer Is Human: The Role of Failure in Successful Design · Henry Petroski · 1985
St. Martin's Press / Vintage · ~272 pages
A successful structure tells you one thing only: it didn't collapse. The real boundary of knowledge is always drawn by failure.
Core Insight

Petroski is a structural engineer, and he states from inside the profession a truth outsiders find counterintuitive: engineering advances mainly through failure, not success. A bridge or an aircraft that has run safely for ten years proves only that it did not collapse under those conditions—it tells you nothing about how much safety margin remains or how close it stood to the limit. Success is silent; only failure speaks, pointing precisely at "here, this is where it fails." So every major structural disaster advances engineering knowledge more than countless successful builds.

Hence an unsettling paradox of success: a run of successes breeds the next failure. Designers see predecessors as "too conservative, too much material," and generation by generation trim redundancy, widen spans, push toward the limit—each step "proven" safe by prior success, until one generation crosses the line no one had marked, disaster strikes, and knowledge is reset. Petroski notes that major bridge collapses have historically recurred roughly every thirty years—about one generational turnover of engineers: those who lived through the last failure retire, the lesson fades, and redundancy gets eaten away again.

How success eats the safety margin until failure resets it
safety margin time / successive successes → danger line (failure occurs) trimming redundancy… failure → reset

He also dismantles the misreading of the "factor of safety." Engineers build in several times the needed redundancy; Petroski says it should really be called the "factor of ignorance"—it measures not how safe we are, but how little we understand about material fatigue, wind-induced vibration, resonance, and unexpected loads. Redundancy is insurance bought against "what we don't know we don't know"; once success tempts us to trim it, what we're really buying is disaster. The Comet airliners breaking apart from metal fatigue at the corners of square windows, the Tacoma Narrows Bridge twisting apart in resonance with the wind—all were failure modes designers did not anticipate and could not have, forced into the textbooks only by the failures themselves.

Key Quotes
"The colossal disasters that do occur are ultimately failures of design, but the lessons learned from those disasters can do more to advance engineering knowledge than all the successful machines and structures in the world."
— To Engineer Is Human, Chapter 1
"No one wants to learn by mistakes, but we cannot learn enough from successes to go beyond the state of the art."
— To Engineer Is Human
Limitations

Written in 1985, its cases center on the mechanical structures of civil engineering and aviation, not the failure modes of newer systems like software and networks. The prose is essayistic; the theoretical frame is looser than later books of its kind. It says little about "how to institutionalize the memory of failure"—it is better at diagnosis than at prescription.

Application for BigCat

"Factor of safety = factor of ignorance" is a ready-made diagnostic blade for anyone building distributed systems. A system runs two years without incident, and the team begins trimming redundancy—cutting backups, raising utilization, removing "seemingly superfluous" degradation plans—each step "proven" safe by past stability, until one traffic spike crosses the line no one marked. To try next week: pick one piece of redundancy that "hasn't caused trouble, so you're about to optimize it away," and write down which class of failure-you-don't-yet-understand it was originally insurance against. Can't articulate it? Then you truly understand it and may trim it. Can articulate it? Then what you're about to cut is the factor of ignorance. And second: turn major-incident post-mortems into an institutional memory that doesn't vanish with staff turnover, to fight Petroski's "thirty-year cycle"—don't let the lesson reset to zero the moment those who lived it retire.

Black Box Thinking
Black Box Thinking: Marginal Gains and the Secrets of High Performance · Matthew Syed · 2015
John Murray · ~336 pages
Learning from failure is not automatic. It is a culture you can choose—or refuse—and the contrast between two industries proves it.
Core Insight

Syed builds the whole book on a stark contrast of two industries: aviation vs medicine. Aircraft carry two black boxes; every accident is investigated by an independent body with no blame attached, and the conclusions are fed back into the entire industry by force—so aviation grows ever safer. Medicine is the reverse: errors are often concealed, blamed on individuals, waved past under the name "complication," lacking a mechanism of no-fault investigation. The result: the same preventable errors take lives in hospitals over and over. The difference lies not in the intelligence or goodwill of practitioners, but in how the system as a whole treats failure.

He abstracts the two attitudes into "closed loop" and "open loop." A closed loop is where error information is denied or reinterpreted away, so the loop seals shut, the system learns nothing, and it spins in place; an open loop is where failure is recorded honestly, treated as signal, and fed back—so the loop opens and the system keeps evolving. What decides your fate is not whether you err, but where the information flows after the error.

Why is the closed loop so common? Syed's answer is the psychology of "cognitive dissonance": when the facts clash with one's self-image ("I'm competent / I'm right"), people don't admit error cleanly—they unconsciously rewrite the narrative: find excuses, cherry-pick data, blame luck, disparage the evidence, so the self-image survives intact. The more expert, the more authoritative, the more one's reputation is staked on "never being wrong," the stronger this defense. So the very people who should most learn from failure are the ones who most resist admitting it. It isn't bad people covering up—it's ordinary people protecting the self.

The way out Syed calls "marginal gains": break a big goal into countless measurable small links, let each link expose its tiny failures, and optimize them one by one. Britain's cycling team turned a one-percent improvement in every detail into an overwhelming advantage exactly this way. It works precisely because it shrinks "failure" and strips its stigma—each micro-failure is neither fatal nor shameful, so people are willing to report it honestly, and the loop opens.

Key Quotes
"For organizations beyond aviation, it is not about creating a literal black box; rather, it is about the willingness and tenacity to investigate the lessons that often exist when we fail, but which we rarely exploit."
— Black Box Thinking, Introduction (defining "black box thinking")
Limitations

The argument is clear but somewhat thin—"learn from failure" is proven by example after example, and the second half feels repetitive. The aviation-vs-medicine dichotomy is simplified: medicine's recent patient-safety movement is in fact moving toward the aviation model. It says too little about "which failures shouldn't be risked at all"—it is better at encouraging the open loop than at drawing its boundaries.

Application for BigCat

The "cognitive dissonance → closed loop" chain speaks most directly to parenting. The moment a child (or an adult) errs, if what follows is blame or shame, cognitive dissonance drives her to learn to conceal, make excuses, avoid risk—turning home into a closed loop. To make it an open loop, the mechanism isn't "more encouragement" but decoupling error from self-worth: in the debrief, ask only "what does this data tell us to adjust next time," never "whose fault was it." To try next week, a family-version "black box": agree that once a week, each person shares one of their own failures that week and one thing learned from it—parents go first and model admitting error—so the child sees that owning a failure is not shameful but a respected act in this home. The same applies to team retrospectives: swap the "blame meeting" for a "no-fault investigation," and information will actually flow back into the system.

Questions to Ask Yourself After Reading

  1. In Taleb's terms, look at what you're currently doing: is it fragile, robust, or antifragile? If an unexpected shock hit, would you be crushed, stay put, or actually benefit?
    A framework

    The key is "nonlinearity" and "optionality." Ask two things: (1) Would one big shock harm you far more than the same total delivered as many small setbacks? Yes = you're on the concave side, fragile. (2) Is your downside capped and your upside open? Yes = you're on the convex side, antifragile. Most people mistake themselves for "solid" when they've merely hidden the risk in a "rare but fatal" tail—that isn't resilience, it's fragility disguised as it.

  2. Check your most recent "attempt" against Harford's three principles: Do you dare try new things? Would you survive if it failed? Did you actually seek feedback and admit the failure? Which of the three are you most missing?
    A framework

    Each has its own way of killing you. Missing the first (don't dare try) = long stagnation, rotting safely; missing the second (scale too large) = one fatal bet, no second chance; missing the third (ignore feedback) = paying tuition repeatedly and learning nothing. Most smart people are stuck on the third—not that they don't try, but that after failing, cognitive dissonance quietly rewrites the narrative, so "failure" vanishes in their own eyes. Honest check: when was the last judgment you explicitly admitted "this one was wrong"?

  3. Look at your system (team, family, codebase) through Petroski + Syed together: Is it quietly trimming redundancy amid success, creeping toward a line no one marked? When failure occurs, does the information flow back into the system, or get buried under narrative?
    A framework

    Watch two signals together. Petroski signal: is some piece of redundancy being optimized away "because nothing's gone wrong," with no one able to say what it originally guarded against?—that's likely trimming the factor of ignorance. Syed signal: after the last problem, was the debrief a "no-fault investigation, only what to fix," or "find whose fault it was"? If the latter, your system is a closed loop—error information is flowing toward blame, not improvement. Both signals red means you're personally brewing a "thirty-year-cycle" scale of failure.