Day 42 · Failures & Disasters of Technology

A disaster does not happen on the day it happens. It was written long before.

8 August 2026 (Saturday) · BigCat's Time Machine
A bridge, a ship, a rocket, a reactor. The four accidents share no technical details at all, yet the inquiries keep arriving at the same finding: inside an organisation, bad news travels slower than the fault does.
EVENT · 01

The only man empowered to stop the work never visited the siteThe Quebec Bridge Collapse · 29 August 1907

29 Aug 1907Quebec · St. Lawrence RiverRoyal Commission 1908 · Petroski

Theodore Cooper was the most celebrated bridge engineer in North America, and his contract gave him final authority over every drawing of the Quebec Bridge. He was nearly seventy, in poor health, and visited the site exactly once during the entire build. To cut costs the main span was stretched from 1,600 to 1,800 feet (549 m, the longest cantilever span in the world at the time) — and the dead-load estimate was never recalculated. Design and construction both sat with a single firm, Phoenix Bridge.

From June 1907 the resident engineer Norman McLure reported repeatedly that compression chords on the south arm were visibly bent, and the bend was growing; Phoenix insisted it was old deformation from shipping. The argument dragged on for two months. On 27 August Cooper wired from New York to add no further load — the telegram never became a stop-work order. At half past five on 29 August, nineteen thousand tons of steel fell into the St. Lawrence in fifteen seconds. Seventy-five men died, thirty-three of them Mohawk ironworkers from Kahnawake.

The 1908 Royal Commission found the dead-load error by Cooper and the designer Peter Szlapka to be the direct cause. Henry Petroski (To Engineer Is Human, 1985) reads a deeper layer: the technical knowledge was adequate; what was missing was institutional structure — design, construction and review compressed into one firm, with the veto handed to a man who knew the site only through letters. The constraint is hard: even if the telegram had arrived that morning, nobody on site was authorised to halt the work. Canada afterwards made independent external review mandatory on major works.

A critical design signed off by one remote "architect", with the review living in an email thread rather than in a process; a floor where doubts exist but nobody holds the right to stop the line.

Concentrating the veto in the most experienced person while giving no one on site the right to stop makes system reliability depend on whether one man happened to get his mail that day.
In your team, can the most junior person bring everything to a halt? Who absorbs the cost when they do?
EVENT · 02

There were enough lifeboats — by the rules of 1894Titanic and the Board of Trade Rules · 14–15 April 1912

14 Apr 1912North AtlanticWreck Commissioner 1912 · SOLAS 1914

In 1894 the British Board of Trade required ships above ten thousand tons to carry sixteen lifeboats. The rule went unchanged for eighteen years while ships kept growing. Titanic measured 46,328 gross tons, carried twenty boats with capacity for 1,178, and had 2,224 people aboard — she was not merely compliant, she exceeded the legal minimum. The reason is that lifeboats were not modelled for full evacuation at all, but for ferrying people to a rescue ship coming alongside.

On 14 April the ship received at least six ice warnings. At 23:40 the berg was sighted; the officer ordered hard-a-port and engines astern. The impact fell along the starboard bow, opening roughly 90 metres of intermittent damage and flooding six watertight compartments — the design margin was four. The bulkheads rose only to D and E decks and were not capped, so water spilled over each divider in turn. She sank at 02:20; more than 1,500 died.

The most popular counterfactual is ramming the berg head-on: structural analysis suggests a bow-first collision might have destroyed only the first two compartments. Physically plausible, decisionally impossible — the officer had about 37 seconds. The variable that could really have been changed sat eighteen years earlier: both the British and American inquiries pointed to lifeboat requirements scaled by tonnage rather than by passengers carried. The debate forks here: one camp blames Captain Edward Smith for not slowing in an ice field, another notes that every North Atlantic liner then did the same. The first SOLAS convention of 1914 followed, requiring lifeboat capacity for everyone aboard.

Compliant ≠ safe: regulatory metrics lock in the previous generation's failure modes. Autonomous driving judged by disengagements per thousand miles, and finance measuring risk by historical volatility, can both fail well above the passing line.

The most dangerous accidents often occur inside fully compliant systems: the rules encode the previous generation's imagination of how things fail.
The safety metric you rely on — which year's system was it designed for? Does the failure mode it assumes still hold?
EVENT · 03

"Take off your engineering hat and put on your management hat"Challenger, STS-51-L · 28 January 1986

28 Jan 1986Cape CanaveralRogers Commission · Vaughan · Feynman

The segmented joints of the shuttle's solid rocket boosters were sealed against hot gas by two O-rings. Roger Boisjoly, an engineer at contractor Morton Thiokol, had written a memo in July 1985 warning that the joint could fail catastrophically, and that cold stiffens the rubber and slows its rebound. Recovered boosters from earlier flights repeatedly showed O-ring erosion by hot gas — each time judged to be "within the acceptable range of experience."

On the night of 27 January a cold front hit Florida; launch temperature was forecast at about 2℃, against a previous coldest launch of 12℃. On the teleconference the engineers recommended delay. NASA's Larry Mulloy shot back: "My God, Thiokol, when do you want me to launch — next April?" After a five-minute caucus, executive Jerald Mason told engineering vice-president Robert Lund: take off your engineering hat and put on your management hat. The company then signed off under management's name. Liftoff came at 11:38 the next morning; the vehicle broke apart 73 seconds later, killing all seven. At the hearings Richard Feynman pressed a piece of O-ring into ice water and it failed to spring back. In Appendix F he recorded another number: management put the odds of failure at one in a hundred thousand, working engineers at one in a hundred.

Diane Vaughan's The Challenger Launch Decision (1996) demolished the popular story of managers knowingly gambling. Her central concept is the normalisation of deviance: nobody suppressed data; NASA repeatedly treated "eroded but did not explode" as evidence the system was still inside its margin, and the standard drifted downward across a decade. Edward Tufte blamed presentation instead: the charts that night were ordered by flight number rather than temperature. Vaughan's reply is that the revealing curve can only be drawn once you know the answer. The hard constraint remains: delaying one launch would not have fixed the joint design.

Alert fatigue in SRE practice: each near-miss logged as "nothing happened" ratchets the tolerance up a notch, and the consumption of margin fires no alarm of its own.

Organisations are rarely destroyed by one wrong decision. They are destroyed by a run of rationalisations that begin "it was fine last time."
Is there an anomaly on your watch that keeps recurring and keeps being ruled acceptable? When exactly did the threshold quietly loosen?
EVENT · 04

The flaw had been known for eleven years — and classifiedChernobyl Reactor No. 4 · 26 April 1986

26 Apr 1986Pripyat · UkraineINSAG-1 (1986) · INSAG-7 (1992) · Plokhy

The Soviet RBMK-1000 graphite-moderated boiling-water reactor had two lethal properties: a positive void coefficient — when coolant boils into steam, reactivity rises rather than falls, a positive feedback loop — and control rods tipped with graphite, so that the first stage of insertion increases reactivity. In 1975 Unit 1 at the Leningrad plant had already suffered partial fuel melting from the first of these. The accident report was classified, and operators at other stations knew nothing of it.

In April 1986 Unit 4 was to complete a test postponed for years before shutdown. On the 25th the Kyiv grid called for power, delaying the test ten hours and handing it to a night shift untrained for it. During the power reduction xenon-135 built up and output fell unexpectedly to 30 MW thermal, far below the planned 700. Procedure required abandoning the test; deputy chief engineer Anatoly Dyatlov ordered it continued, and operators withdrew a great many control rods to drag output back to 200 MW — the core was now in a state the procedures expressly forbade. The test began at 01:23:04; coolant boiled and the positive void coefficient amplified itself. Thirty-six seconds later the AZ-5 emergency scram was pressed, the graphite tips entered first, and reactivity rose another step. Two explosions lifted a thousand-ton lid. Thirty-one died at the scene or of acute radiation sickness.

The IAEA's INSAG-1 report of 1986 blamed the operators; the 1992 INSAG-7 revised that sharply: design flaws were the primary cause, and the operators had entered, unknowingly, a state the reactor should never have permitted. Serhii Plokhy (Chernobyl, 2018) pushes the root cause one level higher — secrecy severed organisational learning: had the 1975 lesson been published, procedures and rod design had eleven years in which to change. A hard constraint also applies: the USSR chose the RBMK because it could be refuelled online and required no heavy pressure-vessel forging industry — its safety deficit was a deliberate trade under industrial constraints, not an oversight.

Civil aviation is the inverse case: accident reports published worldwide and anonymous reporting make one lesson binding across the whole industry. Where safety incidents stay in-house, every firm must fall into every hole itself.

A system's true safety level equals the distance it lets bad news travel.
How far did your organisation's last incident review actually reach? Only the team, or also the people who have never hit that particular hole?

The gap between the warning and the disaster

None of the four lacked a warning. What was missing was a channel that turned warnings into action.
Quebec Bridge
bent chord reported → collapse: about 10 weeks
Titanic
lifeboat rule finalised → sinking: 18 years
Challenger
Boisjoly memo → break-up: 6 months
Chernobyl
Leningrad accident → explosion: 11 years

Four accidents, four severed loops

The usual account is about individuals: who miscalculated, who was arrogant. What actually operates is the loop.
Case / year
The usual account
The actual mechanism
Quebec Bridge · 1907
the engineer got the dead load wrong
design and build in one firm, review remote and single-point: the check loop existed only on paper
Titanic · 1912
arrogance and too few lifeboats
rules encoding an obsolete failure model, so compliance itself became the blind spot
Challenger · 1986
management overruled the engineers
normalisation of deviance: "nothing happened" taken as evidence of safety
Chernobyl · 1986
the operators broke the rules
secrecy cut cross-organisational learning; a known flaw went eleven years unfixed

Deeper reflection

Question 1: Of the four loops, which is hardest to repair in your own organisation?
The right to stop work is the easiest to write into policy and the easiest to lose in practice — whoever exercises it bears the whole cost of delay, while the losses avoided are never visible. Harder still is cross-organisational learning: it requires telling your competitors about your own embarrassments. Aviation manages it because a single crash damages confidence in the entire industry. Where an accident only harms one firm's reputation, information sharing will not emerge on its own; it has to be compelled.
Question 2: Vaughan's "normalisation of deviance" or Tufte's "the chart was drawn wrong" — which is the better guide to action?
Tufte's prescription is concrete — organise evidence by physical variable rather than by sequence number — but it carries the risk of hindsight. Vaughan's diagnosis generalises better: record every "anomaly without consequence" as margin consumed rather than as evidence of safety, and set a cumulative trigger per anomaly class, after which the design must be re-argued rather than the individual case re-assessed. The accounting method matters more than any single judgement.
Question 3: What do these four failures say about complex systems in general?
Charles Perrow (Normal Accidents, 1984) argued that when a system combines interactive complexity with tight coupling, accidents are a normal product rather than an exception; the high-reliability organisation theory of Weick and Sutcliffe answers that carrier decks and air traffic control stay accident-free under the same conditions through preoccupation with weak signals and deference to expertise at the front line. The four cases favour the latter — each failure points to a loop that could have existed. This matches the distributed-systems intuition: what decides the outcome is the isolation boundary and the rollback path, not how reliable any single component is. One further inference deserves caution: safety rules are almost entirely retrospective, so a field that has not yet had its disaster has not yet had its rules written.