Deep Research
Multi-agent investigation · adversarial verification · counter-evidence included
Method: ① Investigate — dozens to hundreds of parallel AI agents read primary sources → ② Adversarial check — every load-bearing claim is challenged by 3 independent verifiers (word-for-word source checks, counter-evidence search) → ③ Write — claims that fail are cut or downgraded; counter-evidence stays in.
The AI-Native Transformation of Software Organizations
Why everyone feels faster while organizations get less stable — from transaction costs, Conway/Brooks and cybernetics to innovation theory, with a dedicated chapter on high-reliability legacy organizations.
Org economicsSoftware engCyberneticsInnovationHRO / legacy
2026-07 · 166 adversarial votes · 6 testable claims
Are Junior Engineers Really Disappearing?
Employment at record highs while the 22-25 bracket falls nearly a fifth — a flow phenomenon, not a stock one. Four data gauges calibrated against each other, answering one paradox: why the technology that helps novices most shrinks novice jobs first.
Labor economicsHiring dataRCT evidenceCareer ladderAI & jobs
2026-07 · 45 adversarial votes · 7 testable claims
AI Code Review: Cure for the Verification Bottleneck, or Turtles All the Way Down?
Humans can't keep up with reviewing AI-written code, so the industry's answer is another AI to do the checking. An anatomy of the vendor benchmark wars and big tech's production funnels, asking how much independence AI-verifying-AI retains — and when it genuinely works.
Verification asymmetryBenchmark warsProduction dataHuman factorsLLM-as-judge
2026-07 · 45 adversarial votes · 8 testable claims
A Foundation Inspection of Scalable Oversight
The whole 'AI watching AI' program stands on one sentence: 'verification is easier than generation.' The founding papers qualified it; later citation used it as an axiom. Theoretical cracks, conditional empirics, and the labs' own retreat to CoT monitoring map what the foundation actually bears.
AI alignmentVerification asymmetryDebate empiricsWeak-to-strongCoT monitoring
2026-07 · 60 adversarial votes · 9 testable claims
The Evidence Hierarchy of Learning Science: What Actually Works?
Ten thousand hours, learning styles, highlighting — why does popular study advice barely overlap with the measured evidence? Five academic battles audited one by one, the common methods re-seated by evidence strength, with anti-scam rules for reading effect sizes.
Learning scienceMeta-analysesDeliberate practiceLearning stylesActive learning
2026-07 · 72 adversarial votes · 12 testable claims
The Machine-Judge Atlas: How Much Can LLMs Scale Software's Oracles?
The same AI yields a twenty-year CVE plugged into a crash detector and a slop flood dumped on human triage — the only difference is who judges. A cell-by-cell audit of software's oracle families, checking the evidence for and against each cell's 'LLM gain'.
Verification asymmetryTest oraclesFormal verificationFuzzing / differencingLLM-as-judge
2026-07 · 87 adversarial votes · 11 testable claims
The '70% of Transformations Fail' Autopsy: Does the Last Wave Predict AI?
The transformation industry's thirty-year calling card — '70% fail' — was measured by whom? Nobody, it turns out. A citation autopsy of the number, a re-reading of the measured failure-rate record, and an audit of the Agile/DevOps archives for what they do and don't predict about AI-native transformation.
Org changeZombie statisticsAgile/DevOps recordInstitutional isomorphismAI transformation
2026-07 · 90 adversarial votes · 11 testable claims
The README for Agents: Infrastructure or Cargo Cult?
Everyone says your repo needs an AGENTS.md/CLAUDE.md — but the first controlled studies fight each other, and a hostile methods audit struck down two of the most-quoted numbers. The standards war, the vendor consensus, the attack surface, and an enterprise rollout playbook — adversarially verified, contradiction-searched, claim by claim.
Agent context filesAGENTS.md standardControlled evidencePrompt injectionRollout playbook
2026-07 · 117 verdicts across 3 rounds · 10 testable claims
The Ironies of Automation: Does the Human Veto Decay as AI Improves?
The more reliable the automation, the faster the human veto seat decays — is the 1983 prophecy replaying inside AI workflows? A full physical for 'the human holds the veto': forty years of human-factors experiments, the aviation and medicine field archives, and the newest AI-era evidence with measured interventions.
Human factorsSkill decayAutomation biasHuman-in-the-loopAI workflow design
2026-07 · 96 adversarial votes · 11 testable claims
The "95% of AI Pilots Fail" Physical: Enterprise AI's Real Base Rate
The market-spooking "95%" can't stand on its own report — so how much enterprise AI actually fails? A verbatim anatomy of the viral number's birth and mutations, the vendor, consultant, and official yardsticks arranged into one ladder, and an honest answer to the base rate of transformation difficulty, built on thirty years of failure-rate archaeology and four economic theories.
Zombie statisticsEnterprise AI ROIYardstick warsProductivity J-curveShadow AI
2026-07 · 102 adversarial votes · 12 testable claims
Is This AI Capex Boom Another 1999?
Everyone cites the 1999 telecom bubble — but that bubble itself is misremembered: the disease was demand myths, accounting fraud, and debt, not "building infrastructure." This physical straightens 2026's ledger yardsticks, then tests the analogy limb by limb: where the money comes from, how long the assets live, whether the demand is real, and which two of five theories are actually operational.
AI capexCapital cycleTelecom bubbleFinancing structureDepreciation fight
2026-07 · 126 adversarial votes · 12 testable claims
AI's Hardware Shortage and Power Shortage: Real, and For How Long?
Memory sold out by spring, electricity bills climbing — but the hardware and power shortages are two different animals: one clears by price, one rations by queue. The forecasts carry two priors of overestimation, yet the constraint is already written into auction results and bills; lay out the supply-side delivery timetables and 'how long' mostly answers itself.
AI infrastructureGrid constraintsMemory supercyclePhantom demandJevons paradox
2026-07 · 36 adversarial votes + dual-seat audit · 10 testable claims
What Should You Major In? An Evidence Check for Students Choosing a Degree
Official ten-year occupational projections are trustworthy about broad direction and barely beat "assume nothing changes" at the detailed-occupation level - which is exactly where choosing a major happens. The figures most often imported into admissions advice, taken apart one caliber at a time, then a field list organized by criteria with an evidence grade in every cell.
Choosing a majorEmployment dataCaliber trapsAI and entry-levelUS-China
2026-07 · 180 adversarial votes + 24 audit seats · 10 testable claims
Screens, Teens, and Mental Health: Is Haidt Right?
The Anxious Generation propelled phone bans across half the Western world, while Nature's review ruled it 'not supported by science' — why has seven years of fighting settled nothing? The crisis, the causality, and the policies audited as three separate questions, with both camps' signature numbers sent back to their primary sources.
Teen mental healthCausal inferenceCaliber trapsPhone bansZombie statistics
2026-07 · 72 adversarial votes + 10 audit seats · 12 testable claims
Grading the Evidence on Longevity Interventions: Rapamycin, Metformin, NAD+, Fasting
The best-evidenced compound has almost no retail market while the one that failed has the biggest — and that is not a coincidence. Four longevity stars' signature numbers sent back to their primary documents, covering the three-site mouse lifespan ledger, the reversal and reproduction of metformin's epidemiology, a physical on the NAD+ major premise, the randomized evidence on caloric restriction and time-restricted eating, and how loose the surrogate endpoints under every claim really are.
Evidence-based medGeroscienceTrial methodsSupplement industryRegulation & markets
2026-07 · 102 adversarial votes · 11 testable claims
The Endgame of Indexation: Is Passive Investing Breaking Markets?
The same "passive share" differs 40-fold depending on which ruler you pick, and each camp reports only the ruler that flatters it. Every piece of evidence in this debate returned to its primary source, covering the three ways to measure the share, the warners' own words and what became of them, the six incompatible measures of price discovery, the index effect's disappearance and revival, the active-underperformance scoreboard, and the two layers where concentration is genuinely well evidenced.
Index fundsMarket microstructureCaliber trapsCorporate governanceFinancial regulation
2026-08 · 96 adversarial votes · 14 testable claims