A shocking percentage of intellectual stalemates boil down to a single category error: two people evaluating the same claim using completely different evidentiary standards.
A physicist scoffs at an economics paper because its empirical identification relies on instrumental variables with wide confidence intervals rather than a 5σ particle collision detection. An economist dismisses a literary critique because its insights cannot be randomized across treatment and control groups. A philosopher rejects an empirical psychology finding because it fails to achieve formal deductive certainty.
Whenever someone asks during an argument, “Is this economics or is it physics?”, they aren’t merely classifying academic departments. They are asking fundamentally:
What counts as proof here? How many degrees of freedom does the investigator possess? How rapidly does our counterfactual certainty decay as the system evolves?
To make sense of this, we can construct the Epistemic Ladder: a hierarchy organizing human inquiry not by social prestige or philosophical purity, but by epistemic warrant, system complexity, and characteristic failure modes.
Truth by definition and syntactic validity. Completely decoupled from physical contingency.
Invariant natural constants, identical particles, and highly controlled, repeatable laboratory isolation.
High-dimensional, evolving biological organisms. No two cells, mice, or patients are identical.
Agents possess internal models and react to observation or policy. History is an N=1 path.
Phenomenological resonance, internal world consistency, and normative/evaluative meaning.
Let’s climb down the ladder rung by rung to examine what happens to evidence at each stage—and where the most common intellectual traps lie.
Rung 1: Formal Deduction & Mathematics
At the summit sits pure deductive inference. When a mathematician proves that there are infinitely many prime numbers, the claim does not depend on sampling data from the universe, running a laboratory trial, or checking if the speed of light changes on Tuesdays. The claim is true by virtue of syntactic validity within an axiomatic system.
The Machine-Verifiable Subset
Traditionally, human mathematics relied on peer review: humans reading human proofs. But human peer review is notoriously porous. Landmark proofs—from the classification of finite simple groups spanning tens of thousands of pages to subtle gaps in published algebraic topology papers—frequently harbor silent flaws for years or decades.
The gold standard of Rung 1 has now migrated to interactive theorem provers (like Lean, Coq, and Isabelle). Here, the entire proof tree is broken down into elementary inference steps checked by a tiny, formally audited micro-kernel. At this level, epistemic certainty reaches its highest possible human realization: if the kernel accepts the proof and the axioms are consistent, the claim is verified with complete machine certainty.
Does Gödel’s Incompleteness Matter in Practice?
Whenever formal deduction is discussed, someone inevitably invokes Gödel’s First and Second Incompleteness Theorems: no consistent formal system capable of doing basic arithmetic can prove its own consistency, and there will always exist true mathematical statements unprovable within the system.
This is a monumental philosophical result, but does it actually limit practical knowledge?
- When it applies: Gödel applies strictly to formal axiomatic systems powerful enough to encode arithmetic (Peano Arithmetic, ZFC). It is not a generic metaphor for “human knowledge is imperfect” or “we can never know anything for sure.”
- Why it rarely matters practically: In 99.99% of mathematics, engineering, and software verification (calculus, linear algebra, topology, cryptography, compiler proofs), working mathematicians never bump into Gödelian unprovability. Statements like the Riemann Hypothesis or P ≠ NP might be hard, but there is no evidence they are unprovable in ZFC. Most unprovable statements are specially constructed self-referential paradoxes or obscure combinatorial puzzles (like Paris–Harrington).
- Where it does matter: Gödel matters as an epistemic speed limit against foundational hubris. It proves that no system can fully bootstrap its own absolute certainty from the inside without taking axioms on trust.
Rung 2: Stationary Physical Laws
Step down one rung, and we leave pure syntax to enter the physical world. Physics deals with stationary empirical systems: every electron in the universe is indistinguishable from every other electron, gravitational fields obey invariant equations across billions of years, and experimental variables can be tightly isolated inside vacuum chambers and cryostats.
Consider the measurement of the electron’s anomalous magnetic dipole moment (g/2). Theoretical calculations using multi-loop Feynman diagrams in QED predict this value to 12 decimal places:
ae = 0.001 159 652 181 643 ...
When experimentalists trap a single electron in a Penning trap and measure its precession, the experimental value matches the theoretical calculation to within parts-per-trillion precision. This is the pinnacle of empirical confirmation.
Even in fundamental physics, human beings conduct the experiments. And human beings are prone to anchoring and confirmation bias.
In his famous 1974 Caltech commencement address on "Cargo Cult Science," Richard Feynman highlighted how Robert Millikan’s oil drop experiment for the charge of an electron was slightly off because Millikan used an incorrect value for the viscosity of air. Subsequent researchers who measured the charge got numbers slightly higher than Millikan’s, looked for reasons why their results were “wrong,” and adjusted their apparatus until their numbers came closer to Millikan’s published figure.
A nearly identical phenomenon occurred with the measurement of the speed of light (c) in the 1930s and 1940s. Following Raymond Birge's influential 1934 consensus review, measurements clustered tightly around ~299,776 km/s. Multiple independent labs published confidence intervals that were far too narrow—and which completely excluded the true modern value (299,792.458 km/s).
The epistemic lesson: Having a stationary physical system does not protect you from social confirmation bias. If investigators prune "outliers" to match existing prestige literature, their estimated confidence intervals will dramatically overstate true precision.
What If the Universe Changes? (Cosmic Phase Transitions & Evolving Laws)
The bedrock assumption of Rung 2 is stationarity—the Principle of Uniformity stating that physical laws do not drift over time or vary by location (which, via Noether’s Theorem, gives us the fundamental conservation of energy and momentum). But what if physics itself evolves as the universe expands and cools?
Theoretical cosmology and astrophysics have produced fascinating work on this exact question:
- The Laws Already Did Change (Cosmic Phase Transitions): We don't have to speculate in the abstract. In the Standard Model, the effective laws of physics were radically different in the primordial universe. At the electroweak phase transition (t ≈ 10−11 s after the Big Bang), the Higgs field dropped into a non-zero vacuum expectation value. Before this transition, W and Z bosons and elementary quarks had zero rest mass. What we call "invariant physical laws" are the frozen, low-temperature ground state of a cooled cosmos.
- Empirical Searches for Drifting Constants: Paul Dirac famously proposed the Large Numbers Hypothesis (1937), suggesting Newton's gravitational constant G decreases inversely with cosmic time (G ∝ 1/t). Modern tests look at atomic spectra from distant quasars billions of light-years away, and analyze the Oklo Natural Nuclear Reactor in Gabon—a natural 1.7-billion-year-old uranium deposit that placed an extreme upper bound on any historical drift in the fine-structure constant (|Δα / α| < 10−7 over 2 billion years).
- Cosmological Natural Selection (Lee Smolin): Physicist Lee Smolin proposed that collapsing black holes spawn "baby universes" whose physical parameters mutate slightly upon birth. Universes that produce more black holes leave more progeny, optimizing our cosmos for star formation. In The Singular Universe and the Reality of Time, Smolin and Roberto Mangabeira Unger argue that time is fundamental, and physical laws are more like evolving "habits" than timeless mathematical decrees.
The Epistemic Takeaway (The Meta-Law Escalation): Notice how physics handles non-stationarity compared to social science. When physicists suspect a constant like α or G changes, they do not descend into unstructured chaos. Instead, they promote the constant into a dynamic field (like the dilaton or inflaton) governed by a higher-order differential equation. Physics preserves its Rung 2 standing by stepping one derivative up: if the parameter varies, the meta-law governing its variation is assumed to be stationary.
Rung 3: Complex Empirical Systems (Biology & Medicine)
When we cross from physics to biology and pharmacology, the epistemic ground shifts underneath our feet. In physics, one electron is identical to all others. In biology, no two cancer cells, laboratory mice, or human immune systems are identical.
We are now dealing with high-dimensional, non-linear, evolving systems full of latent confounders, feedback loops, and biological drift.
The GRADE Framework and Why Simple Deduction Fails
In medicine, you cannot deduce whether a drug works from basic biochemical principles alone. Countless compounds that look brilliant in a petri dish fail catastrophically in vivo due to liver first-pass metabolism, off-target binding, or compensatory hormonal upregulation.
To manage this noise, medicine developed the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework:
- Low Warrant: Mechanistic plausibility, in vitro assays, and animal models (where 90%+ of positive findings fail to translate to humans).
- Moderate Warrant: Observational cohort studies and retrospective case-controls (vulnerable to healthy user bias, confounding by indication, and reverse causality).
- High Warrant: Pre-registered, double-blind, multi-center Randomized Controlled Trials (RCTs) with intention-to-treat analysis and hard clinical endpoints.
Notice what happened: because the underlying system has thousands of unmeasured degrees of freedom, the standard of proof had to shift from deductive calculation to statistical blinding and explicit randomization.
Rung 4: Reflexive & Non-Stationary Systems (Economics & Social Science)
If biology is difficult because organisms are complex, social science is difficult because the agents are looking back at the experimenter.
Social systems are reflexive (George Soros) and non-stationary. Human beings anticipate policy, alter their behavior when measured (Goodhart’s Law), and adapt strategically (the Lucas Critique). Furthermore, history offers only an N=1 realization of macroeconomic trajectories: you cannot run a parallel planet Earth where the Federal Reserve didn't raise rates in 2022.
How do you know if high-performing charter schools (like KIPP) actually improve student test scores, or if they simply attract families with unusually motivated parents (selection bias)?
Joshua Angrist and collaborators solved this using a clean natural experiment: when KIPP schools are oversubscribed, state law requires admission by randomized lottery. By comparing the lottery winners to lottery losers—two groups identical in motivation, family background, and unobserved drive—economists isolated a pristine causal effect.
This is social science at its highest epistemic rigor: the randomization breaks the feedback loop between student motivation and school quality.
In one of the most famous papers in modern political economy, Acemoglu, Johnson, and Robinson (2001) argued that contemporary economic prosperity is caused by historical institutional quality (rule of law, property rights).
To solve the endogeneity problem (rich countries can simply afford better institutions), they introduced an ingenious Instrumental Variable: historical European settler mortality in the 17th–19th centuries. The chain of logic:
Settler Mortality (1800) → Early Settlement Density → Colonial Institutions → Modern Institutions → Modern GDP
Why this is epistemically fragile: For an Instrumental Variable to be valid, it must satisfy the Exclusion Restriction: settler mortality can affect modern GDP only and exclusively through its effect on institutions.
Critics (e.g., Jeffrey Sachs, David Albouy, Edward Glaeser) immediately pointed out that high historical settler mortality was driven by malaria, yellow fever, and tropical disease ecology. Those same geographical and disease burdens directly undermine labor productivity, trade access, and human capital accumulation today—completely bypassing institutions.
Because you cannot rerun European colonialism on a randomized grid of planets, macro-historical instruments almost inevitably fail the exclusion restriction.
Rung 5: The Intersubjective & Aesthetic ("Art" / The Evaluative)
At the base of our ladder sits the vast realm of creative expression, literary analysis, visual arts, and normative philosophy.
It is tempting to treat Rung 5 as a mere wastebasket of "arbitrary personal opinion" where evidence cannot exist. But this is a mistake.
While art lacks empirical falsifiability or statistical stationarity, it possesses rigorous internal standards of validity:
- Internal Coherence (David Lewis’s Truth in Fiction): In Tolkien’s Middle-earth, it is an objective, verifiable fact that Frodo carried the One Ring to Mordor, and false that he had a laser rifle. Fictional worlds are governed by authorial stipulation and modal logic.
- Intersubjective Resonance: A great novel or musical composition is not random noise; it compresses universal structures of human psychology, grief, tension, and release. If an author writes dialogue that rings false to every human reader, the critique is not merely arbitrary—it identifies a failure of affective modeling.
- Formal Optimization: A sonnet, a fugue, or a cinematography sequence solves severe formal constraints (meter, voice-leading, compositional framing) that optimize for perceptual clarity and emotional transmission.
The error is not that art is “inferior” to physics; the error is trying to judge art by the rules of physics, or trying to settle macroeconomic monetary policy using aesthetic intuition.
The Diagnostic Matrix
Whenever you enter a technical or policy disagreement, determine where each participant is standing on the ladder:
| Rung | Core Domain | Primary Warrant | Degrees of Freedom | Signature Vulnerability |
|---|---|---|---|---|
| 1. Formal Deduction | Math, Logic, Proof Verification | Syntactic derivation, kernel checking | Zero | Axiomatic inconsistency, unnoticed assumptions |
| 2. Stationary Physics | QED, Classical Mechanics, Chemistry | 5σ detection, invariant physical constants | Minimal | Social anchoring, systematic apparatus bias |
| 3. Complex Empirical | Biology, Pharmacology, Oncology | Double-blind RCTs, GRADE meta-analyses | High | Translational failure, biological heterogeneity |
| 4. Reflexive Social | Economics, Sociology, Geopolitics | Lotteries, natural experiments, diff-in-diff | Very High | Lucas critique, broken exclusion restrictions |
| 5. Intersubjective | Art, Literature, Ethics, Philosophy | Internal coherence, affective resonance | Unbounded | Solipsism, ungrounded semantic drift |
Conclusion: Restoring Coherence to Debate
When someone claims an economic paper is “unscientific” because its findings may not hold in another country 20 years later, they are confusing Rung 4 with Rung 2. Human societies evolve; electrons do not.
Conversely, when a proponent of an unproven medical treatment insists that their personal clinical intuition outweighs a 5,000-patient randomized trial, they are trying to import Rung 5 subjective impressions into Rung 3 empirical medicine.
By making the ladder explicit, we can stop talking past one another and start asking the only question that matters: Given the noise and reflexivity of this system, what is the best evidence humanly possible?