中文EN
← Deep Research
Deep Research · Deep dive

The Ironies of Automation: The Better the AI, the Faster the Human Veto Decays? (Deep Dive)

This is the deep-dive edition · read the plain-language edition →
TL;DR
'The better the automation, the faster the human veto decays' was an analytical prophecy in a 1983 process-control paper — half of which was actually engineering prescriptions. The forty-year report card: skill decay (dose-dependent, cognitive-first) and the vigilance ceiling are confirmed; automation bias is multi-source with RCT causality — error-laced LLM advice cost physicians with 20 hours of AI training 14 points of diagnostic accuracy, and even very experienced radiologists fell from 82.3% to 45.5%. The 2025 endoscopy study delivered the first real-world signal of 'worse after AI is removed' (ADR 28.4%→22.4%, observational, unconfirmed), pairing with the in-use gain (RCT meta RR 1.24) as the two ledgers of deployment evaluation. Forty years grew no falsification school, only a mitigation literature: experienced failure, outcome accountability, decision-time nudges, and verifiability design have measured effects; exhortation and generic training don't. Vigilance physiology, unrehearsable unknown failures, and base-rate arithmetic are the three unengineerable floors. Eleven testable claims close the essay.
96 votes · 32/32 survivedin-use RR 1.24 vs post-AI −6.0ppengineerable + 3 hard floors11 testable claims

Every empirical citation in this essay went through graded verification: 32 load-bearing claim groups each received 3 independent fidelity votes (verbatim checks against primary sources, caliber recomputation; 96 votes, 0 groups overturned, 40+ caliber corrections), and 5 single-source load-bearing empirics additionally faced a contradiction-search seat (hunting for independent-team, independent-data measurements) and a methods-audit seat (hostile refereeing with veto power; 9 verdicts). The audit changed how this essay is written in several places: the endoscopy study's adjusted odds ratio is withheld because a published correction touched it; "cognitive forcing has a threefold cost" shrank to "the only significant cost is perceived complexity"; "explanations reduce overreliance" acquired a boundary condition imposed by an independent measurement. Labels: 【multi-source】= independent teams, independent data, same direction; 【single-source, verified】= one measurement, fidelity-checked and past methods audit; 【direction contested】= independent measurements disagree; 【verified】= verbatim-checked quotes and mechanistic facts from primary/official documents; 【vendor caliber】【industry self-report】【unverified, source】as labeled. Fidelity ≠ truth — for single-source claims fidelity is the verification ceiling, which is exactly why the grades exist. Source index at the end.

0. A 2025 medical study, and a 1983 prophecy

In August 2025, The Lancet Gastroenterology & Hepatology published a Polish study: months after four endoscopy centers introduced AI polyp detection, nineteen physicians averaging 28 years in practice saw their adenoma detection rate in non-AI colonoscopies drop from 28.4% to 22.4%. Every percentage point of that rate has a measurable correspondence to patients' colorectal cancer risk. It was the first time anyone had tied "humans got worse after using AI" to a clinical quality metric — though it is an observational study, and we will give it a full methods physical shortly. 【single-source, verified; see Chapter 5】

Forty-two years earlier, UCL cognitive psychologist Lisanne Bainbridge wrote this script in a five-page paper, "Ironies of Automation." She wasn't talking about AI but about process control in chemical plants; the logic is identical: the more advanced the automation, the more the human is left with only the tasks the designer couldn't automate; skills decay without use; yet the system counts on precisely this rusted skill set to take over at the worst possible moment. Her words: "There is no way in which the human operator can check in real-time that the computer is following its rules correctly" — "The human monitor has been given an impossible task." 【verified】

For this site the question is not academic. When Code Becomes Cheap, AI Code Review, A Foundation Inspection of Scalable Oversight, and The Machine-Judge Atlas all converged on the same seat: the verdict stays with the human. If that seat corrodes under automation, the whole line of argument stands on crumbling ground. So this issue attacks the thinnest part of our own position: after forty years of human-factors evidence, is "the human holds the veto" a death sentence, or an engineerable constraint?

The verdict up front: it is a constraint, it is engineerable, but there are three floors that cannot be engineered away — and every lever with a measured effect comes with a measured cost. The evidence, oldest first.

1. How to read the numbers: three anti-scam rules

This literature has three systematic traps; every conclusion below should be read with them in mind:

2. Rereading the original: what those five 1983 pages said, and didn't

Bainbridge 1983 (Automatica 19(6):775-779) deserves a verbatim rereading, because it has been cited for forty years (roughly 2,800 Google Scholar citations by end-2023) and most citers remember only the word "irony." 【verified】

What she said. The core irony: "the more advanced a control system is, so the more crucial may be the contribution of the human operator" — because the designer who tries to eliminate the operator "still leaves the operator to do the tasks which the designer cannot think how to automate." The skill layer: "physical skills deteriorate when they are not used" — a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one, while the takeover moment demands someone "more rather than less skilled, and less rather than more loaded, than average." The vigilance layer: citing Mackworth's 1950s vigilance studies, humans cannot maintain effective visual attention on a low-event source for more than about half an hour — so posting a human to watch a machine that almost never fails is a structural design error. 【verified】

What she didn't say. She did not say "therefore automation is a dead end." Roughly half the paper (§2-3) is engineering prescriptions: failures on a seconds timescale must have reliable automatic backup, and if that is impossible and failure costs are unacceptable, the plant should not be built; automatic systems should "fail obviously"; emergency responses should be overlearned on a high-fidelity simulator; each shift should retain a period of real hands-on control. In other words, the original is itself a "constraints plus mitigations" list; the pure-death-sentence image is a later cropping. 【verified】

One more thing posterity forgets: this is an analytical paper in a process-control context with essentially no original data. In 1983 it was prophecy; it needed the next forty years of evidence to live or die.

3. Forty years in the lab: two robust phenomena, and the gaps they admit themselves

Human factors caught Bainbridge's prophecy with two concepts: automation complacency (not watching when you should) and automation bias (believing whatever the machine says). Parasuraman & Manzey's 2010 integrative review is the ledger. 【verified】

The positive side. The founding experiment (Parasuraman, Molloy & Singh 1993; students; the MATB multitask platform): with constantly reliable automation, mean detection of automation failures was just 33%; with variable reliability, 82%; in single-task pure monitoring, about 97%. Three numbers, one sentence: misses are not laziness but multitask attention allocation — and automation that never fails is precisely the most dangerous kind, because it teaches attention to leave. 【verified】 On the automation-bias side, the review's conclusion is hard: the phenomenon "(a) can be found in different settings, (b) occurs in both naive and expert (e.g., pilots) participants, ... (d) cannot be prevented by training or explicit instructions to verify the recommendations of an aid" (complacency showed up in pilots averaging ~440 flight hours and in veteran air traffic controllers; up to 60 minutes of extra training didn't reduce it). 【verified】 The cruelest expert data point is medical: experienced readers using mammography CAD — when the machine failed to mark a cancer, human detection of "unmarked" cancers fell from 46% unaided to 21% aided. The machine's silence was read as evidence of absence. 【verified; Alberdi 2004/2008 via P&M 2010】

Skill decay has a quantitative floor too: the Arthur et al. 1998 meta-analysis (53 articles; retention analysis on 178 data points, N=8,719) — skill loss deepens with disuse, reaching about δ=-1.27 beyond one year (note: only 3 data points in that bin; the more widely circulated -1.4 from the abstract does not match the paper's own table, so this essay cites the table); and cognitive tasks decay faster than physical ones (δ=-1.15 vs -0.75). Bad news for the AI era: cognitive tasks are exactly what AI automates. 【single-source, verified (meta-analysis layer)】

The negative side — listed by the same review. Complacency still has no consensus definition. Moray & Inagaki's critique has gone unanswered for decades: not watching a machine that almost never fails may be rational attention allocation, not dereliction — no study has shown humans sample less than the normative optimum; Lee & See 2004 found no credible evidence that high trust causes poor monitoring. Plus the lab-failure-rate problem from Chapter 1. 【verified】 Keep this page of the ledger in mind — it is the evidence ceiling of the deskilling narrative, and Chapter 5's medical data will run into it again.

The 1993 founding experiment: detecting automation failures Constantly reliable automation · multitasking33%Occasionally failing automation · multitasking82%Single-task pure monitoring (control)~97%
Schematic: Parasuraman, Molloy & Singh 1993 (students, MATB multitask platform) — automation that never fails teaches attention to leave; lab failure rates far exceed real systems, so extrapolate with care (Chapter 3)

4. Aviation: the most complete field archive, and a two-sided verdict

Aviation is the only industry that has run the full loop — automation dependency written into accident investigations, regulatory documents, controlled measurements, and intervention cycles. Its archive is the best rehearsal material for AI workflows.

The accident side. AF447 (2009, 228 dead): the autopilot dropped out over iced pitot tubes and handed a perfectly flyable aircraft back to three pilots; the stall warning sounded continuously for 54 seconds; neither flying pilot ever mentioned it or formally identified the stall. The BEA final report's verdict deserves verbatim quotation: when initial control and diagnosis both fail, the human-as-backup safety model enters "common failure mode" — losing control makes the situation unreadable, and unreadability deepens the loss of control. The report put the absence of high-altitude manual flying training into the causal chain: "Current training practices do not fill the gap left by the non-existence of manual flying at high altitude." 【verified】 Four years later, Asiana 214: the flying pilot had 9,684 hours, had been an instructor, had passed every check — and the NTSB still found he lacked critical manual flying skills. The NTSB also dismantled a self-reported number along the way: the airline claimed 77.7% of its 2012 777 landings were manual; the NTSB noted the data never said at what altitude the autopilot came off — most crews likely took over below 1,000 feet. "We still hand-fly plenty" (industry self-report) and "manual skills are degrading" (official finding) can both be true. 【verified】

The statistics side. The FAA's 2013 automation working group reviewed 26 accidents plus 20 major incidents: over 60% of the accident reports identified a manual handling error as a factor; in over 50%, investigation boards found pilots "out of the control loop and peripheral to the actual operation of the aircraft and therefore not prepared to assume control when necessary" — Bainbridge's sentence, almost verbatim, in a regulatory document. Roughly a quarter showed overconfidence in the automation. 【verified】 The same year, FAA SAFO 13002 wrote both sides into one paragraph: autoflight systems "have improved safety and workload management," but "continuous use of autoflight systems could lead to degradation of the pilot's ability to quickly recover the aircraft from an undesired state." 【verified】

The controlled-measurement side. Haslbeck & Hörmann 2016: 126 airline pilots, stratified-random sampled by fleet × rank, hand-flying a raw-data ILS in full-flight simulators — the long-haul fleet (fewer hand-flying opportunities) was significantly worse than short-haul, fleet main effect ηp²=.45 (localizer .39 / glideslope .38); and recent flight practice predicted fine-motor performance better than total experience. Skill decay is dose-dependent: what rusts is "haven't practiced lately," not "haven't flown much." 【single-source, verified】 A finer blade in Casner et al. 2014 (16 pilots, small sample): stick-and-rudder and scanning skills were "mostly intact"; what degraded was the cognitive layer — navigation reasoning, anomaly recognition. What needs practice is the thinking, not the hands. 【single-source, verified (abstract level)】

The two-sided verdict. Now the other pan of the scale: aviation got dramatically safer in the automation era — Boeing's statistical summary has the fatal accident rate down 60% and the total accident rate down 35% over the past two decades while departures grew more than 20% 【industry self-report (the industry-standard dataset)】; by Airbus's count, the fourth-generation fly-by-wire fleet with envelope protection had zero fatal loss-of-control (LOC-I) accidents in the last decade, and its LOC-I hull-loss rate is 89% lower than generation 3 【vendor caliber — Airbus builds the FBW, interest declared】. But note: AF447 itself was a generation-4 aircraft — it crashed after the protections degraded to alternate law. The two sentences must be said together: automation saves lives in aggregate while creating a new failure mode concentrated at the handback moment. Not a contradiction; the same coin.

The intervention side (a lesson in implementation discount). After 2013 the FAA added 14 CFR 121.423 "Extended Envelope Training" (full stall and upset recovery, compliance by 2019) — note it mandates extreme-state recovery, not routine manual-flying upkeep, which is exactly the gap the auditors flagged: the DOT Inspector General found in 2016 that after issuing the hand-fly-more SAFO, the FAA never verified whether carriers actually increased manual opportunities; only 5 of 19 reviewed simulator training programs mentioned monitoring skills; and airspace modernization keeps shrinking line opportunities to hand-fly. 【verified】 Since then, IATA reported zero LOC-I accidents worldwide in both 2020 and 2025 【industry self-report; causality untested — multiple safety investments ran in parallel】. Aviation's lesson cuts both ways: evidence that interventions work exists; verification that interventions happen is missing.

5. Medicine: the second archive, and this essay's money chart

The medical archive opens with a counterintuitive lesson: the most famous "humans ignore the machine" number is not human dereliction at all. Clinicians override drug-safety alerts 49-96% of the time (about 25% for serious overdose alerts) — but in reviewed cases, expert panels agreed with the clinician's decision to override a valid alert 95.6% of the time (a single-study figure), and adverse events were observed after only 2.3%-6% of overridden alerts — with the review declining to attribute even those to the override itself. The main driver of high override rates is terrible alert specificity — an engineerable alarm-design problem, not a corroding veto seat. 【verified, van der Sijs 2006】 This rhymes with Moray/Inagaki's "rational sampling" critique in Chapter 3: before convicting the human of unreliability, check whether the system is spamming. The Argus mechanism from The Machine-Judge Atlas — low error rate × very low base rate = false alarms drowning true ones — ran for twenty years in hospital alerting before AI judges rediscovered it.

But the automation-bias bill is real. The last generation of diagnostic aid — mammography CAD — at scale: Fenton 2007 (43 facilities, 222,135 women, 429,345 screens): with CAD, specificity fell 90.2%→87.2%, PPV fell 4.1%→3.2%, biopsy rates rose 19.7%, sensitivity didn't significantly improve, and overall accuracy (AUC) was lower with CAD (0.871 vs 0.919). Lehman 2015 (625,625 screens, 271 radiologists): CAD improved nothing; hardest to explain away is the within-reader comparison — the same radiologists had 83.3% sensitivity with CAD and 89.6% without. A generation paid tuition on "give the human a machine prompter": the assist didn't materialize, the behavior changed anyway. 【verified】 The mechanism experiment twists the knife: Dratsch 2023, 27 radiologists reading mammograms — with correct AI advice, all experience groups sat near 80% accuracy; with wrong AI advice, the inexperienced fell to 19.8%, moderately experienced to 24.8%, the very experienced to 45.5% — "all radiologists, regardless of expertise, can be subject to automation bias." Seniority buys partial resistance only. 【single-source, verified】

Then the 2025 endoscopy diptych — this essay's money chart. The front: the gain from AI-assisted colonoscopy while in use is multi-source — a meta-analysis of 21 RCTs, 18,232 patients: ADR 44.0% vs 35.9% (RR 1.24), adenoma miss rate down 55% relative, at a cost of 0.47 minutes of extra withdrawal time. 【multi-source (RCT meta; the ADR effect graded low-certainty)】 The back: the Budzyń study from the opening — after AI's introduction, the same senior physicians' non-AI colonoscopy ADR fell from 28.4% (226/795) to 22.4% (145/648), an absolute difference of -6.0 points (95% CI -10.5 to -1.6, P=0.009). 【single-source, verified】

The methods-audit seat's full verdict must be handed to the reader intact: this is a retrospective, single-group before-after observational study; "deskilling" is the hypothesis it raises, not the mechanism it proves. Three unexcluded rival explanations: AI exposure is perfectly collinear with time (seasonality, case-mix drift, and process changes are inseparable); how the post-introduction "non-AI cases" arose is opaque (if they were the selectively remaining subset after AI took the rest, the ADR drop could be pure selection bias); and withdrawal time — the strongest modifiable predictor of ADR — was never measured. The adjusted analysis keeps the direction, but the widely quoted adjusted odds ratio was touched by a published correction (the appendix model omitted a covariate), so per our audit this essay withholds its precise value. Nineteen physicians; overall ADR below quality benchmarks; generalization limited. The contradiction-search seat found: zero direct independent replication of this clinical endpoint — it remains single-source; one nearby Kraków study points the other way (trainees who learned with AI had higher independent ADR), but with a different population, question, and likely overlapping institutions — tension, not refutation; the mechanism layer (wrong prompts misleading experts) does have independent cross-domain confirmation in radiology and pathology. ACG's framing is the fair one: a hypothesis requiring prospective confirmation. One detail adds credibility rather than subtracting: the authors are core AI-colonoscopy researchers (Mori, Bretthauer et al. — the same community behind the RR 1.24 meta). Their interests point toward "prove AI works"; they published the signal against themselves. 【direction contested (deskilling endpoint: single source + adjacent opposite-direction measurement); mechanism multi-source】

The honest reading: the in-use gain is hard; the post-withdrawal decay is a real but unconfirmed signal. Anyone making deployment decisions now keeps two ledgers: joint performance with AI present, and residual human performance with AI absent (outage, out-of-scope, migration). Nobody kept the second ledger before 2025.

Two ledgers, one technology: adenoma detection rate (ADR) in colonoscopy Ledger 1: with AI present (meta of 21 RCTs) Ledger 2: after AI removed (one observational study) 35.9%Standard 44.0%AI-assisted 28.4%Before AI era 22.4%After AI era 35.9% → 44.0% (multi-source) 28.4% → 22.4% (unconfirmed hypothesis)
Schematic: left = Hassan 2023 RCT meta (in-use gain, RR 1.24); right = Budzyń 2025 (senior physicians' non-AI colonoscopies, −6.0pp, observational, confounds unexcluded) — deployment evaluation must keep both books (Chapter 5)

6. The AI-era physical: which is evidence, which is panic

Studies on "AI makes people dumber" exploded in 2024-2026. This chapter's value is not recounting them but grading their methods — quality variance is enormous, and media amplification is nearly uncorrelated with quality.

The self-report layer (consistent direction, weakest evidence). Lee et al., CHI 2025 (Microsoft Research + CMU; 319 knowledge workers, 936 task examples): higher confidence in GenAI, less self-reported critical thinking. Three qualifiers: cross-sectional, perceived-enactment self-report, vendor-led. The paper's real value lies elsewhere — its introduction splices 1983 in verbatim: "As Bainbridge noted, a key irony of automation is that by mechanising routine tasks and leaving exception-handling to the human user, you deprive the user of the routine opportunities to practice their judgement and strengthen their cognitive musculature, leaving them atrophied and unprepared when the exceptions do arise." A process-control prophecy, quoted by a 2025 flagship-venue paper to describe GenAI — the literature grew this connection on its own. 【verified】

The panic layer (fails methods audit). MIT Media Lab's "Your Brain on ChatGPT" (EEG essay study) is the most media-amplified: unreviewed preprint, n=54 (the crucial fourth session had 18), an independent methods commentary computed the design needed ~159 participants for adequate power, the FDR correction level went unreported, and some subanalyses had 2-4 essays per group. The authors' own FAQ resists the framing: "No! Please do not use the words like 'stupid', 'dumb', 'brain rot', 'harm', 'damage'..." 【verified】 Gerlich 2025 (n=666), the "AI use correlates with worse critical thinking" study, is cross-sectional self-report; the author concedes reverse causality can't be excluded — people weak in critical thinking may simply lean on AI more. 【verified】 History supplies a direct warning: the flagship of the last "technology rots your brain" wave — Sparrow's 2011 "Google effects on memory" in Science — had its core experiment fail replication twice (the 2018 Nature Human Behaviour systematic replication project, and a dedicated 2020 replication). Cognitive-panic research has priors. 【verified】

The causal layer (a real RCT — the hardest evidence in this lane). Qazi et al. (NEJM AI 2026; single-blind, preregistered): 44 physicians who had completed a 20-hour AI literacy course, randomized to read clinical vignettes; the group receiving ChatGPT-4o advice laced with errors scored 14 points lower in diagnostic-reasoning accuracy than the group receiving error-free advice (84.9% vs 73.3%, 95% CI -18.9 to -9.1). Randomization licenses the causal reading: wrong advice caused the drop, and generic AI training did not immunize against it. 【single-source, verified (magnitude); direction multi-source — independent teams measured the same direction with bigger effects in chest X-ray (Gaube 2021) and a multicenter radiology experiment (2024, n=220, accuracy falling 55-69 points under wrong advice)】 Two audit qualifiers: 84.9→73.3 is a between-group comparison, not the same people falling; and "training doesn't help" convicts only this one 20-hour generic course — the same team's follow-up RCT showed a behavioral nudge clawing back 7.6 points. The correct sentence is "generic training alone is not enough," not "training is useless."

The coding layer (the gap itself is the finding). As of this writing, no published study has measured AI-induced skill decay in professional developers with a pre/post or longitudinal design — a May 2026 systematic meta states plainly that no controlled pre/post skill-retention studies on professionals were identified. What exists is repository traces (GitClear: copy-paste share 8.3%→12.3% 【vendor caliber】) and self-report (DORA 2024: AI adoption up, self-reported delivery stability -7.2% 【industry self-report】). The nearest controlled evidence comes as two opposite-signed footnotes: Anthropic's own RCT (52 engineers learning an unfamiliar Python library) found the AI group scored 50% on a subsequent comprehension quiz vs 67% for the manual group (d=0.738), the biggest gap in debugging — skill formation impaired 【single-source, verified; vendor-run research】; while the novice study (Kazemitabaar 2023) found kids who learned with Codex were no worse once AI was removed. A formation-harm signal, zero measurement of decay in existing skill — the single most overdue experiment in the industry; it goes into the testable claims. 【direction contested】

Radiologists' accuracy collapses under wrong AI advice (Dratsch 2023) With correct AI advice With wrong AI advice 79.7%19.8%Inexperienced81.3%24.8%Moderately exp.82.3%45.5%Very experienced
Schematic: 27 radiologists reading mammograms, AI prompts experimentally manipulated; seniority buys only partial resistance — 'all radiologists, regardless of expertise, can be subject to automation bias' (Chapter 5)

7. The opposition: there is no "Bainbridge was wrong" school — only a fight over how hard mitigation is

An honest physical needs the strongest opposition. The search result is itself a finding: forty years of literature contain no systematic "the ironies were falsified" camp. The closest thing — Dekker & Woods 2002 — attacks the "men are better at / machines are better at" function-allocation framework (MABA-MABA), and their alternative, joint human-machine systems design, is precisely "keep engineering, under a different frame." The 40th-anniversary special issue (Ergonomics 2023) reads as confirmation-plus-extension: Endsley's verdict is that "Not only are Bainbridge's original warnings still pertinent for AI, but AI's very nature and focus on cognitive tasks has introduced many new challenges." 【verified】

The real opposition runs on three finer fronts:

One: "the decay narrative is overstated." Chapter 3 listed it: the rational-sampling critique stands unanswered; "high trust → poor monitoring" lacks credible evidence; lab failure rates are inflated. Add this chapter's: the lumberjack meta-analysis (Onnasch et al. 2014, 18 experiments) confirmed the trade-off's shape — higher degrees of automation improve routine performance and worsen failure performance and situation awareness, with the penalty steepening past the "information analysis → action selection" boundary (across the 6 boundary-crossing studies, the return-to-manual correlation deepens from -.34 overall to -.90) — but it is all laboratory work, almost all student participants, the authors themselves flag low statistical power, and later critics question extrapolation to complex real settings. The direction of the trade-off is credible; the numbers must not be carried around as iron law. 【verified (meta conclusions); extrapolation limited】

Two: "human-in-the-loop is a placebo" — a camp more radical than Bainbridge. Ben Green 2022 reviewed 41 policies requiring human oversight of government algorithms: the empirical record shows people cannot perform the oversight functions the policies assume, so the policies end up legitimizing flawed algorithms and providing false comfort. The underlying experiment (Green & Chen 2019 — MTurk crowdworkers, not real judges, keep that caliber): participants underperformed the risk assessment even while seeing its predictions, and could not tell whether they or the algorithm was more accurate. The legal synthesis (Crootof, Kaminski & Price 2023) named the trap — the "MABA-MABA trap" (paraphrasing the paper's core argument): inserting a human into the loop yields not the best of both but a new entity — a hybrid system requiring its own regulation; far from marrying the strengths of humans and machines, hybrid systems can exacerbate the worst of each while adding new sources of error. Note that none of the three concludes "remove the human": Green proposes two-stage institutional oversight; Crootof et al. enumerate nine roles a human can play in the loop, then regulate the hybrid. The placebo critique targets undesigned human review, not human review. 【verified】

Three: overreliance economics — the most constructive wing of the opposition. Vasconcelos et al. 2023 (5 experiments, maze tasks) turned "will the human verify the AI" from a morality question into an economics question: overreliance is a strategic cost-benefit choice — when explanations genuinely drop verification cost to at-a-glance and doing it yourself is expensive, people verify, and overreliance falls significantly. 【single-source, verified (that setting)】 The contradiction-search seat drew the boundary sharper: the mechanism layer (verification cost as the lever) has independent same-direction support from Harvard, KIT, and UW teams 【multi-source】, but the universal version — "explanations reduce overreliance" — was contradicted with opposite sign by an independent measurement: feature-importance explanations, which do not really lower verification cost, help only on easy tasks and fail or backfire on hard ones. The macro backdrop belongs here too: a 2024 Nature Human Behaviour meta (106 studies) found human-AI combinations on decision tasks average worse than the better of human or AI alone. 【direction contested (explanation effect); multi-source (mechanism)】 This lane shares a coordinate system with the Argus mechanism in The Machine-Judge Atlas: base rate decides the false-alarm flood; verification cost decides whether the human actually verifies. Both are design variables, not constants of human nature.

8. The engineerable list: measured effects with measured costs, and three hard floors

Spread forty years of intervention evidence on the table and you get an honest scorecard — every row has a cost column:

Levers with measured effects:

Levers measured ineffective: exhortations to "please verify the AI's suggestions"; lecture-style training; naive transparency (in the 2026 Army experiment transparency significantly increased workload, and "opaque + voluntary handoff" beat "transparent + forced" — n=24, direction as caution only). 【verified】

The lever still missing its last mile: adaptive automation (dynamically rotating control) has worked for thirty years — a 10-minute manual insertion mid-session significantly revives subsequent monitoring — and has reached realistic ATC simulation and in-aircraft technology demos; but as of 2026 there is still no controlled, peer-reviewed operational deployment evaluation. Thirty years without crossing the last mile from simulator to operations is itself a data point. 【verified】

Three floors that cannot be engineered away:

  1. The physiology of vigilance. The half-hour ceiling replicates from Mackworth to today, with a modern addendum: vigilance work is high-load suffering, not easy boredom. Any workflow design where "a human watches the AI work, continuously" is betting against forty years of vigilance literature. 【verified】
  2. Unknown failures cannot be rehearsed. Strauch 2018 (retired NTSB investigator): simulators train known failure modes, and automation's new failure modes are by definition the ones designers didn't anticipate. The training lever has a principled ceiling. 【verified】
  3. Base-rate arithmetic. The better the automation, the rarer the failure; the rarer the failure, the worse the human detection (that 33% lab) and the higher the false-alarm share (the Argus arithmetic). The veto seat's value and its reliability trade off against each other by construction — structure, not attitude.

The seat-design corollary for AI workflows, four parameters: base rate (how many true anomalies does this seat see daily? below the vigilance line, don't build a pure monitoring seat — use sampling audits plus automatic backstops); verification cost (is verifying one AI output much cheaper than redoing it? if not, redesign the signal); practice dose (when did this person last do the task without AI? prescribe the dose, and audit that it happens); failure visibility (does the AI fail obviously? — 1983's "fail obviously" becomes the hardest and most valuable requirement in the LLM era, because LLMs fail fluently).

Intervention scorecard: forty years of measurements Measured effective (each with a cost) Measured ineffective Drills with experienced AI failures (not briefings)Personal accountability for overall outcomesBehavioral nudges at decision timeCutting verification cost to at-a-glanceScheduled manual dose (audit that it happens)Cognitive forcing: works at subtask level, feels harder"Please verify the AI" exhortationsLecture-style training / 20-hour coursesNaive transparency Missing its last mile Adaptive automation (rotating control): 30 years in sims; no operational evaluation
Schematic: effect sizes, boundary conditions, and costs per lever in Chapter 8; 'effective' is not universal — most effects are design-dependent

9. Conclusion: eleven testable claims

Ordered by evidence strength:

  1. Wrong advice misleading people in the moment (automation bias) is multi-source and causally established: an RCT measured -14pp (physicians, error-laced LLM advice); independent teams measured the same direction at -55 to -69pp in radiology and mammography (the very experienced still fell to 45.5%); seniority buys partial resistance only. 【multi-source + RCT】
  2. "In-use gain" and "post-withdrawal decay" are two ledgers: the in-use gain has an RCT meta (ADR RR 1.24); the post-withdrawal decay has one observational study (-6.0pp). Deployment evaluation must keep both books — and almost nobody keeps the second today. 【multi-source vs single-source, verified】
  3. Constant high reliability is the most dangerous regime: the 33% vs 82% vs 97% attention mechanism, plus the ~70%±14% dependence crossover — below the line don't rely on it, above the line fight complacency; "almost never wrong" is complacency's optimal growth medium. The lab-failure-rate extrapolation caveat is booked alongside. 【verified (laboratory)】
  4. Skill decay is dose-dependent and cognitive-first: δ≈-1.27 after a year of disuse; cognitive tasks decay ~0.4 SD faster than physical; hand-flying precision follows recent practice; the decay concentrates in the cognitive layer. Cognitive tasks are what AI automates. 【single-source, verified (each study)】
  5. The 1983 ironies' 2026 report card: skill decay — confirmed (aviation controlled measurements + meta-analysis); the vigilance ceiling — confirmed (robust in the lab); "the human is least ready when most needed" — strongly supported by the field archive (AF447 / FAA statistics); "real-time checking of a stronger system is impossible" — still an analytical claim, but the scalable-oversight issue's evidence points the same way. 【verified】
  6. The absence of an opposition school is a signal: forty years grew no "ironies falsified" literature, only "how to mitigate" literature; the 40th-anniversary verdict is that it applies to AI more, not less. 【verified】
  7. Generic training alone is not enough; what works is experienced failure, accountability, nudges, and verifiability design — each with a measured effect size, and each with a cost or boundary condition; friction is not a panacea, effects are design-dependent. 【multi-source (direction)】
  8. High override rates ≠ human dereliction: 95.6% of overrides were endorsed by expert review; audit alert specificity before auditing the human. Before convicting the veto seat, audit the system's base rate and signal quality. 【verified】
  9. The "compliance placebo" critique of human-in-the-loop holds — against undesigned human review: inserting a human produces no safety; designing and testing the human's role does. The nine-role list is the starting inventory. 【verified】
  10. Most AI-era panic studies fail methods audit (underpowered, cross-sectional, self-report — and the previous wave's "Google effect" failed replication), yet the hardest new signals all point in Bainbridge's direction (the endoscopy observation, the error-laced RCT, the skill-formation RCT). Don't trust the panic; do trust the alarm. 【verified + direction contested】
  11. Skill decay in professional developers remains unmeasured — the 2026 meta confirms the cell is empty. It is the most overdue experiment in the field: pre/post, longitudinal, with "independent performance after AI removal" as the endpoint. 【verified (the gap)】

What to watch next: whether a prospective replication of Budzyń (from the ACCEPT family or an independent team) appears and holds direction; who first runs the professional-developer pre/post; whether adaptive automation ever produces its first operational deployment evaluation; and the first measured data on "fail obviously" engineering — failure-visibility design — in LLM workflows.


The four 1983 ironies: a 2026 report card Skills decay without use; the monitor turns noviceConfirmed: aviation measurements + decay meta-analysisVigilance on low-event sources fails after ~half an hourConfirmed: vigilance studies replicate across 40 yearsThe human is least ready exactly when most neededStrong field support: AF447 / FAA accident statisticsHumans cannot real-time-check a stronger machineStill an analytical claim; oversight evidence points the same way
Schematic: 'confirmed' means the direction has independent empirical support, not literal quantification; grades and sources in the chapters and Chapter 9

Appendix: main sources

Originals & theory: Bainbridge, "Ironies of Automation", Automatica 19(6):775-779, 1983 (verbatim-checked) · Parasuraman & Riley 1997, Human Factors 39(2) · Parasuraman & Manzey 2010, Human Factors 52(3):381-410 (full text verbatim-checked; Alberdi/Bahner/Skitka/Wickens & Dixon figures via this review) · Endsley 1995, Human Factors 37(1); Endsley & Kiris 1995 (via the author's 1996 chapter) · Arthur et al. 1998, Human Performance 11(1):57-101 (original incl. Tables 3/4) · Warm, Parasuraman & Matthews 2008, Human Factors 50(3)

Aviation archive: BEA, AF447 Final Report (2012) (verbatim) · NTSB AAR-14/01 (Asiana 214) · FAA PARC/CAST, Operational Use of Flight Path Management Systems (2013) · FAA SAFO 13002 · DOT OIG AV-2016-013 · Haslbeck & Hörmann 2016, Human Factors 58(4) (DLR self-archive) · Casner et al. 2014, Human Factors 56(8) · Boeing Statistical Summary 1959-2025 (April 2026 ed.) · Airbus Commercial Aviation Accidents 1958-2025 (Feb 2026 ed.) · IATA Safety Report press release (March 2026)

Medical archive: van der Sijs et al. 2006, JAMIA 13(2) (PMC full text) · Fenton et al. 2007, NEJM 356:1399 · Lehman et al. 2015, JAMA Intern Med 175(11) (PMC full text) · Alberdi et al. 2004, Academic Radiology 11(8) · Dratsch et al. 2023, Radiology 307(4):e222176 (publisher's version verbatim-checked) · Budzyń et al. 2025, Lancet Gastroenterol Hepatol 10(10):896-903 (paywalled; figures cross-checked via ACG commentary + institutional press releases; the 2025-09-11 correction and several correspondence letters (count not individually verified) + authors' reply on record) · Hassan et al. 2023, Ann Intern Med 176(9) · Zhou, ACG Evidence-Based GI commentary (Sept 2025) · Orzeszko et al. 2025, Surgical Endoscopy (found by the contradiction seat)

AI era: Lee, Sarkar, Tankelevitch et al., CHI 2025 (full text verbatim-checked) · Qazi et al., NEJM AI 2026;3(5), DOI 10.1056/AIoa2501001 (formerly medRxiv 2025.08.23.25334280; follow-up nudge RCT NCT07328815) · Gaube et al. 2021, npj Digital Medicine · Kosmyna et al., arXiv:2506.08872 and the brainonllm.com FAQ; methods commentary arXiv:2601.00856 · Gerlich 2025, Societies 15(1):6 (correction: Societies 15(9):252) and Data 10(11):172 · Sparrow et al. 2011, Science; Camerer et al. 2018, Nature Human Behaviour; Hesselmann 2020, PeerJ 8:e10325 · Shen & Tamkin, Anthropic Research (2026-01-29) 【vendor-run】 · Yan et al., arXiv 2605.04779 (meta) · GitClear 2025 【vendor caliber】 · DORA 2024 【industry self-report】 · Kazemitabaar et al., CHI 2023

Opposition & interventions: Dekker & Woods 2002, Cognition Technology & Work 4(4) (quote second-hand-checked; original paywalled) · Read & Waterson (eds.), Ergonomics 66(11) 40th-anniversary issue; Endsley 2023, same issue · Strauch 2018, IEEE THMS 48(5) (full text) · Onnasch, Wickens, Li & Manzey 2014, Human Factors 56(3) · Buçinca, Malaya & Gajos 2021, PACM HCI 5(CSCW1) Art.188 (full text, Tables 2/3 recomputed) · Vasconcelos et al. 2023, PACM HCI 7(CSCW1) Art.129 (full text) · Zhang et al. 2024, Mensch und Computer (found by the contradiction seat) · Fok & Weld 2023 · Schemmer et al., IUI 2023 · Vaccaro, Almaatouq & Malone 2024, Nature Human Behaviour · Green 2022, CLSR 45 · Green & Chen 2019, ACM FAT* · Crootof, Kaminski & Price 2023, Vanderbilt Law Review 76(2) · Parasuraman, Mouloua & Molloy 1996, Human Factors 38(4) · USAARL-TECH-TR-2026-02 · Bahner et al. 2008 / Skitka et al. 2000 (via P&M 2010)

Research materials and all verification verdicts are archived in the research base (local ~/design/deep-research-runs/automation-irony/): 32 load-bearing claim groups × 3 fidelity votes = 96 votes, plus 5 single-source empirics × contradiction-search + methods-audit seats = 9 verdicts, all on record — including the decision to withhold the endoscopy study's adjusted coefficient, and every caliber correction.