When large models answer any question fluently, confidently, and with the appearance of evidence, an ancient question turns urgent: what makes a claim true, and what makes it trustworthy? Philosophy of science is the meta-discipline studying "what makes science science, and where reliable knowledge comes from." Today four thinkers (Popper and Kuhn from the West, Mohism and Nyāya from the East) each hand us a yardstick: Popper draws the line of demarcation — only what can be falsified is science; Kuhn tells history — science advances by whole paradigm shifts, not by linear accumulation; the Mohists give three tests of verification — root, source, use; the Nyāya school asks after knowledge's sources — perception, inference, analogy, testimony. Together they form an "epistemological toolkit" for scrutinizing both AI output and your own judgment.
Karl Popper
Western · Critical Rationalism
1902–1994 · The Logic of Scientific Discovery (Logik der Forschung, 1934); Conjectures and Refutations (1963)
Core Thesis + Primary Passage
"Ein empirisch-wissenschaftliches System muß an der Erfahrung scheitern können." — An empirical-scientific system must be capable of being refuted by experience. (The Logic of Scientific Discovery, §6)
Historical Context & Core Insight
Popper's target was the Vienna Circle, who used verifiability to distinguish science from metaphysics. Popper reversed it: induction is a myth — no number of white swans proves "all swans are white," yet a single black swan refutes it. So the mark of science is not confirmation but falsifiability: a theory that in principle no observation could ever refute (he named astrology and certain versions of psychoanalysis) is not invincible but empty of scientific content. Science advances by bold conjecture + severe refutation, forever trying to knock down its own hypotheses.
Cross-Disciplinary Link
Popper's method is engineered into machine learning: holdout sets and cross-validation are precisely the operationalization of "attempted refutation" — a model that cannot be falsified on unseen data (overfit to the training set) has no predictive value. Null-hypothesis testing shares the root: we never "prove" a hypothesis, only fail to reject it. Falsifiable = testable — this is the watershed between science and rhetoric that "can never be wrong."
Contemporary Relevance
BigCat scenario: When making a product or investment decision, flip the question from "what evidence supports me" to "what evidence would prove me wrong" — the direct antidote to confirmation bias. Make AI your red team: don't ask it to agree, order it to "find where this assumption is most likely to collapse." A judgment for which you dare state falsification conditions is the only real judgment.
In one line: the power of science lies not in being confirmable but in daring to be falsifiable — and not yet falsified.
For your single most important current judgment, what fact appearing would make you admit it's wrong? If you can't answer, is it still a judgment?
Thomas Kuhn
Western · History / Sociology of Science
1922–1996 · The Structure of Scientific Revolutions (1962)
Core Thesis + Primary Passage
"The transition between competing paradigms ... like the gestalt switch, it must occur all at once or not at all." — The shift between rival paradigms happens as a whole, all at once, or not at all. (Structure, ch. IX)
Historical Context & Core Insight
Kuhn directly challenged Popper's rationalist picture. Studying the history of science, he found that scientists mostly do "normal science" — puzzle-solving within an established paradigm, setting anomalies aside or patching them rather than discarding the theory. Only when anomalies accumulate into crisis does a "scientific revolution" erupt, one new paradigm wholly replacing the old (geocentric → heliocentric; Newton → relativity). Sharper still is incommensurability: rival paradigms share no neutral observation language to adjudicate between them — the shift is more like a conversion of belief than a compulsion of logic.
Cross-Disciplinary Link
Kuhn explicitly borrowed from Gestalt psychology: the same figure reads now as a duck, now as a rabbit — sensory data unchanged while perception restructures as a whole. A paradigm shift is the scientific community's collective gestalt flip, isomorphic with the perceptual bistability neuroscience studies (the Necker cube). It also resembles a phase transition in complex systems: past a critical threshold the whole reorganizes, and it cannot be done "half a step" at a time.
Contemporary Relevance
BigCat scenario: Large models may well be a paradigm shift in AI — mastery of the old paradigm (symbolic rules, expert systems) may be incommensurable with the new. On a personal level, distinguish: are you doing "normal science" optimization of your existing stack (which has a ceiling), or have you reached the crisis point where a paradigm jump is due? Diligence within a paradigm cannot substitute for the courage to leap between them.
In one line: facts don't overturn theories on their own; only when anomalies build into crisis does a new paradigm replace the old — whole.
In your field, how far have the current "anomalies" accumulated? Are you patching the old paradigm, or waiting for the flip?
Mohism · The Three Tests (sān biǎo)
Eastern · Mozi (pre-Qin Mohist School)
Mo Di, c. 470–391 BCE · Mozi, "Against Fatalism I" (Fēi Mìng Shàng)
Core Thesis + Primary Passage
"Every statement must have three tests. What are the three? There is the root, the source, and the use. Root it in — the deeds of the ancient sage-kings. Source it in — the evidence of the eyes and ears of the common people. Use it — enact it as law and government, and observe whether it benefits the state and people." (Mozi, "Against Fatalism I")
Historical Context & Core Insight
Among the pre-Qin schools, Mozi stood out in demanding that "statements be tested." He used the three tests to refute fatalism: who has ever seen or heard "fate" with their own eyes and ears? — an appeal to the absence of evidence. The three tests are three yardsticks: root (is there historical precedent, the deeds of sage-kings?), source (is there verifiable empirical evidence, "the eyes and ears of the people"?), and use (once enacted, does it truly benefit the state and people?). This is China's earliest systematic empiricist methodology + consequentialism.
Cross-Disciplinary Link
The "use" of the three tests — verifying a claim's truth by its actual consequences — is nearly the same line of thought as pragmatism (Day 37: Peirce's "meaning is effect," James's "cash value of truth"), by over two millennia. "Sourcing in the evidence of the eyes and ears" is a plain empirical spirit: appeal to what is observable and publicly attestable, not to solitary authority. In today's terms, the Mohists wanted us to "let the data speak and validate by results."
Contemporary Relevance
BigCat scenario: Treat the three tests as a checklist against "AI hallucination" and gut-decisions. For any conclusion AI gives, ask in order: Root — does it have a reliable basis / training source? Source — is there empirical evidence I can independently verify? Use — following it, are the actual consequences truly beneficial? Only what passes all three deserves adoption. Run the same list over your own judgments.
In one line: don't judge a statement by how pleasing it sounds — judge it by three points: has it a basis, can it be verified, and once enacted does it truly benefit?
A recent "AI suggestion" or "expert claim" you adopted — can it pass root, source, and use? Which test fails first?
Nyāya School · Pramāṇa (Theory of Knowledge-Sources)
Eastern · Indian Nyāya (Epistemology / Logic)
Akṣapāda Gautama · Nyāya Sūtra, c. 2nd century CE
Core Thesis + Primary Passage
pratyakṣa-anumāna-upamāna-śabdāḥ pramāṇāni. — Perception, inference, comparison, and testimony are the "pramāṇas" (valid means of knowledge). (Nyāya Sūtra 1.1.3)
Historical Context & Core Insight
Nyāya is India's most systematic school of epistemology and logic, its core question: how must knowledge be acquired to count as valid? It lists four "pramāṇas": perception (direct sensing), inference (reasoning, e.g. "smoke, therefore fire"), comparison (analogy from the known to the unknown), and testimony (the word of a reliable source). Unlike the West's emphasis on perception and inference, it gravely lists "reliable testimony" as an independent source — granting that we cannot witness everything, and most knowledge comes from trustworthy transmission. Its inferential schema (thesis–reason–example–application–conclusion) later deeply shaped Buddhist logic (Hetuvidyā).
Cross-Disciplinary Link
The four pramāṇas map neatly onto the epistemic predicament of large models: an LLM has almost no "perception" (it does not sense the world firsthand) and no strict "inference"; its output is overwhelmingly "testimony" + "comparison" — relaying the "testimony" of its training corpus and drawing analogies. The trouble: when that corpus-testimony is itself wrong or context-mismatched, it has no perception or inference to correct it — and hallucination arises. Nyāya warned two thousand years ago: the reliability of testimony depends on whether the speaker is credible — the very key to appraising AI output.
Contemporary Relevance
BigCat scenario: Run an "AI epistemology audit": classify each of the AI's assertions — is it perceived (almost never), inferred, drawn by analogy, or relayed authority? Different sources warrant different trust: relayed claims must be traced to the origin, inferences must have their premises re-checked. Same for yourself — knowing which pramāṇa your judgment rests on is the first line of defense against credulity and dogmatism.
In one line: knowledge must first be asked after its source — seen, reasoned, analogized, or merely heard? Confusing the source is where all credulity begins.
The thing you're most certain of today — which of the four pramāṇas is it? If it's really just "testimony," have you checked whether the speaker is reliable?
Four Views of Science in Chorus
Synthesis
Popper · Kuhn · Mohist Three Tests · Nyāya Pramāṇa
Chorus
Four thinkers answer the same question from different dimensions — what makes a claim scientific, and what makes knowledge reliable?
· Popper (demarcation): the mark of science is falsifiability, not confirmation — a theory no observation could refute has no scientific content.
· Kuhn (history): science advances by wholesale paradigm shifts, not linear accumulation; anomalies must build into crisis, and the flip is gestalt-like.
· Mohism (verification): every statement must pass three tests — root (is there a basis?), source (can it be verified?), use (does enacting it truly benefit?).
· Nyāya (source): knowledge must first be asked after its origin — perceived, inferred, analogized, or received as testimony?
The two Western thinkers press "where the boundary of science lies and how science evolves"; the two Eastern schools press "where valid knowledge comes from and how to test it." Each yardstick guards one dimension, and together they form an "operating system for knowledge": for any incoming assertion, first use Nyāya for source classification — is it perception, inference, analogy, or relayed testimony? Testimony must be traced to its origin. Then run the Mohist positive screen — basis, evidence, and consequences, gate by gate. Next apply Popper's negative stress test — demand "what evidence would prove this wrong," and downgrade any claim that cannot state its falsification conditions. Finally use Kuhn for framework self-examination — watch whether you are merely patching an old paradigm and dismissing all rival evidence as invalid. In the AI age this pipeline is daily self-defense: what large models mass-produce as "well-founded" speech is mostly testimony with no perception to backstop it (Nyāya), pleasing talk that fails the three tests (Mohism), and all-purpose rhetoric that dares state no falsifier (Popper); and the greatest hidden reef of human-AI collaboration is mistaking fluency within a paradigm for the paradigm's permanence (Kuhn).
Reflection
Next time an AI hands you a supremely confident conclusion, of the four steps — classify the source, run the three tests, hunt the falsifier, audit the paradigm — which would you do first? Which have you never done at all?
Going Deeper
1. Who is closer to the truth of science — Popper (perpetual falsification) or Kuhn (paradigm shifts)?
This is the most famous standoff in the history of philosophy of science. Popper wants scientists ever ready to overturn their theories; Kuhn objects that historically scientists rarely abandon a paradigm over one anomaly, or no theory would survive its first day. The fairer view is a division of labor: Popper describes the rational norm science ought to follow, Kuhn the social workings science actually follows. Daily research is Kuhnian puzzle-solving; the real turning points demand Popperian courage to face falsification. Norm and history — each is half right.
2. Mohist three tests (verification) vs. Popper (falsification): which suits the AI age?
The Mohists ask "is there a basis, can it be verified, is it beneficial" — positive confirmation; Popper asks "what would prove it wrong" — negative refutation. The two are complementary: for AI output, first run the three tests as a positive screen (source, evidence, consequence), then press Popper's question, "under what conditions would this advice harm me." The former keeps you from believing too little and missing out; the latter from believing too much and stepping on a mine. Sound judgment alternates both rulers.
3. Audit an LLM with Nyāya's four pramāṇas: which does its output mostly rest on, and why does it hallucinate?
An LLM lacks perception (no firsthand sensing) and rigorous inference (reasoning can break); its output is chiefly testimony (relaying corpus) plus comparison (pattern analogy). Hallucination's root lies here: when it takes unreliable or context-mismatched "testimony" as true, with no perception or inference to backstop it, it confidently fabricates. Nyāya's remedy is plain — the force of testimony depends on the speaker's credibility: so for AI's key claims, trace them to the original source and independently re-check, downgrading "it said so" to "verified, then believed."
4. Kuhn's "incommensurability": can practitioners of old and new paradigms still talk?
Incommensurable does not mean incommunicable — it means there is no neutral referee: both use the same word for different things, and each side's evidence fails within the other's frame. Historically, the old paradigm is often not persuaded but dies out as its adherents retire (Planck: science advances one funeral at a time). The lesson: don't expect cross-paradigm dialogue to be won by "presenting the facts" in one stroke — first identify the difference in premises; and if you sense yourself stuck in an old paradigm, rather than defend it to the death, go experience the "world" of the new one firsthand.