Epistemology & Scientific Method

The Epistemic Ladder

Why We Talk Past One Another, and How to Measure What Counts as Evidence

A shocking percentage of intellectual stalemates boil down to a single category error: two people evaluating the same claim using completely different evidentiary standards.

A physicist scoffs at an economics paper because its empirical identification relies on instrumental variables with wide confidence intervals rather than a 5σ particle collision detection. An economist dismisses a literary critique because its insights cannot be randomized across treatment and control groups. A philosopher rejects an empirical psychology finding because it fails to achieve formal deductive certainty.

Whenever someone asks during an argument, “Is this economics or is it physics?”, they aren’t merely classifying academic departments. They are asking fundamentally:

What counts as proof here? How many degrees of freedom does the investigator possess? How rapidly does our counterfactual certainty decay as the system evolves?

To make sense of this, we can construct the Epistemic Ladder: a hierarchy organizing human inquiry not by social prestige or philosophical purity, but by epistemic warrant, system complexity, and characteristic failure modes.

The Ladder of Evidential Standards
Click on each rung to explore its core warrant, degrees of freedom, and signature failure mode.
▲ Higher Deductive Certainty / Lower Noise Stationary & Closed Systems ▲
Rung 1: Formal Deduction & Mathematics Axiomatic / Kernel-Checked

Truth by definition and syntactic validity. Completely decoupled from physical contingency.

Degrees of Freedom Zero (Strict Axiomatic Rules)
Primary Failure Mode Hidden Assumptions / Kernel Bugs
Benchmark Standard Formal Proof (Lean, Coq)
Rung 2: Stationary Physical Laws 5σ Detection / Closed Systems

Invariant natural constants, identical particles, and highly controlled, repeatable laboratory isolation.

Degrees of Freedom Minimal (Isolated Variables)
Primary Failure Mode Systematic Error / Anchored CIs
Benchmark Standard Precision QED (12 decimals)
Rung 3: Complex Empirical & Medicine RCTs / GRADE Framework

High-dimensional, evolving biological organisms. No two cells, mice, or patients are identical.

Degrees of Freedom High (Biological Heterogeneity)
Primary Failure Mode Confounders / Translational Gap
Benchmark Standard Double-Blind Multi-Center RCT
Rung 4: Reflexive & Non-Stationary (Econ / Social) Quasi-Experiments / IVs

Agents possess internal models and react to observation or policy. History is an N=1 path.

Degrees of Freedom Very High (Strategic Adaptation)
Primary Failure Mode Goodhart's Law / Broken Exclusion
Benchmark Standard Lottery Exploitation / Natural Exp.
Rung 5: The Intersubjective & Aesthetic ("Art") Affective & World-Building

Phenomenological resonance, internal world consistency, and normative/evaluative meaning.

Degrees of Freedom Unbounded (Subjective Qualia)
Primary Failure Mode Solipsism / Semantic Vacuity
Benchmark Standard Internal Coherence & Resonance
▼ Higher System Complexity / Noise Reflexive & Open Systems ▼

Let’s climb down the ladder rung by rung to examine what happens to evidence at each stage—and where the most common intellectual traps lie.

Rung 1: Formal Deduction & Mathematics

At the summit sits pure deductive inference. When a mathematician proves that there are infinitely many prime numbers, the claim does not depend on sampling data from the universe, running a laboratory trial, or checking if the speed of light changes on Tuesdays. The claim is true by virtue of syntactic validity within an axiomatic system.

The Machine-Verifiable Subset

Traditionally, human mathematics relied on peer review: humans reading human proofs. But human peer review is notoriously porous. Landmark proofs—from the classification of finite simple groups spanning tens of thousands of pages to subtle gaps in published algebraic topology papers—frequently harbor silent flaws for years or decades.

The gold standard of Rung 1 has now migrated to interactive theorem provers (like Lean, Coq, and Isabelle). Here, the entire proof tree is broken down into elementary inference steps checked by a tiny, formally audited micro-kernel. At this level, epistemic certainty reaches its highest possible human realization: if the kernel accepts the proof and the axioms are consistent, the claim is verified with complete machine certainty.

Does Gödel’s Incompleteness Matter in Practice?

Whenever formal deduction is discussed, someone inevitably invokes Gödel’s First and Second Incompleteness Theorems: no consistent formal system capable of doing basic arithmetic can prove its own consistency, and there will always exist true mathematical statements unprovable within the system.

This is a monumental philosophical result, but does it actually limit practical knowledge?

Rung 2: Stationary Physical Laws

Step down one rung, and we leave pure syntax to enter the physical world. Physics deals with stationary empirical systems: every electron in the universe is indistinguishable from every other electron, gravitational fields obey invariant equations across billions of years, and experimental variables can be tightly isolated inside vacuum chambers and cryostats.

✓ The Good Example: Quantum Electrodynamics (QED)

Consider the measurement of the electron’s anomalous magnetic dipole moment (g/2). Theoretical calculations using multi-loop Feynman diagrams in QED predict this value to 12 decimal places:

ae = 0.001 159 652 181 643 ...

When experimentalists trap a single electron in a Penning trap and measure its precession, the experimental value matches the theoretical calculation to within parts-per-trillion precision. This is the pinnacle of empirical confirmation.

⚠ The Warning Example: Social Anchoring on the Speed of Light

Even in fundamental physics, human beings conduct the experiments. And human beings are prone to anchoring and confirmation bias.

In his famous 1974 Caltech commencement address on "Cargo Cult Science," Richard Feynman highlighted how Robert Millikan’s oil drop experiment for the charge of an electron was slightly off because Millikan used an incorrect value for the viscosity of air. Subsequent researchers who measured the charge got numbers slightly higher than Millikan’s, looked for reasons why their results were “wrong,” and adjusted their apparatus until their numbers came closer to Millikan’s published figure.

A nearly identical phenomenon occurred with the measurement of the speed of light (c) in the 1930s and 1940s. Following Raymond Birge's influential 1934 consensus review, measurements clustered tightly around ~299,776 km/s. Multiple independent labs published confidence intervals that were far too narrow—and which completely excluded the true modern value (299,792.458 km/s).

The epistemic lesson: Having a stationary physical system does not protect you from social confirmation bias. If investigators prune "outliers" to match existing prestige literature, their estimated confidence intervals will dramatically overstate true precision.

What If the Universe Changes? (Cosmic Phase Transitions & Evolving Laws)

The bedrock assumption of Rung 2 is stationarity—the Principle of Uniformity stating that physical laws do not drift over time or vary by location (which, via Noether’s Theorem, gives us the fundamental conservation of energy and momentum). But what if physics itself evolves as the universe expands and cools?

Theoretical cosmology and astrophysics have produced fascinating work on this exact question:

The Epistemic Takeaway (The Meta-Law Escalation): Notice how physics handles non-stationarity compared to social science. When physicists suspect a constant like α or G changes, they do not descend into unstructured chaos. Instead, they promote the constant into a dynamic field (like the dilaton or inflaton) governed by a higher-order differential equation. Physics preserves its Rung 2 standing by stepping one derivative up: if the parameter varies, the meta-law governing its variation is assumed to be stationary.

Rung 3: Complex Empirical Systems (Biology & Medicine)

When we cross from physics to biology and pharmacology, the epistemic ground shifts underneath our feet. In physics, one electron is identical to all others. In biology, no two cancer cells, laboratory mice, or human immune systems are identical.

We are now dealing with high-dimensional, non-linear, evolving systems full of latent confounders, feedback loops, and biological drift.

The GRADE Framework and Why Simple Deduction Fails

In medicine, you cannot deduce whether a drug works from basic biochemical principles alone. Countless compounds that look brilliant in a petri dish fail catastrophically in vivo due to liver first-pass metabolism, off-target binding, or compensatory hormonal upregulation.

To manage this noise, medicine developed the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework:

  1. Low Warrant: Mechanistic plausibility, in vitro assays, and animal models (where 90%+ of positive findings fail to translate to humans).
  2. Moderate Warrant: Observational cohort studies and retrospective case-controls (vulnerable to healthy user bias, confounding by indication, and reverse causality).
  3. High Warrant: Pre-registered, double-blind, multi-center Randomized Controlled Trials (RCTs) with intention-to-treat analysis and hard clinical endpoints.

Notice what happened: because the underlying system has thousands of unmeasured degrees of freedom, the standard of proof had to shift from deductive calculation to statistical blinding and explicit randomization.

Rung 4: Reflexive & Non-Stationary Systems (Economics & Social Science)

If biology is difficult because organisms are complex, social science is difficult because the agents are looking back at the experimenter.

Social systems are reflexive (George Soros) and non-stationary. Human beings anticipate policy, alter their behavior when measured (Goodhart’s Law), and adapt strategically (the Lucas Critique). Furthermore, history offers only an N=1 realization of macroeconomic trajectories: you cannot run a parallel planet Earth where the Federal Reserve didn't raise rates in 2022.

✓ The Good Example: Angrist & KIPP Charter School Lotteries

How do you know if high-performing charter schools (like KIPP) actually improve student test scores, or if they simply attract families with unusually motivated parents (selection bias)?

Joshua Angrist and collaborators solved this using a clean natural experiment: when KIPP schools are oversubscribed, state law requires admission by randomized lottery. By comparing the lottery winners to lottery losers—two groups identical in motivation, family background, and unobserved drive—economists isolated a pristine causal effect.

This is social science at its highest epistemic rigor: the randomization breaks the feedback loop between student motivation and school quality.

⚠ The Debated Example: Acemoglu, Johnson & Robinson (AJR 2001) Settler Mortality

In one of the most famous papers in modern political economy, Acemoglu, Johnson, and Robinson (2001) argued that contemporary economic prosperity is caused by historical institutional quality (rule of law, property rights).

To solve the endogeneity problem (rich countries can simply afford better institutions), they introduced an ingenious Instrumental Variable: historical European settler mortality in the 17th–19th centuries. The chain of logic:

Settler Mortality (1800) → Early Settlement Density → Colonial Institutions → Modern Institutions → Modern GDP

Why this is epistemically fragile: For an Instrumental Variable to be valid, it must satisfy the Exclusion Restriction: settler mortality can affect modern GDP only and exclusively through its effect on institutions.

Critics (e.g., Jeffrey Sachs, David Albouy, Edward Glaeser) immediately pointed out that high historical settler mortality was driven by malaria, yellow fever, and tropical disease ecology. Those same geographical and disease burdens directly undermine labor productivity, trade access, and human capital accumulation today—completely bypassing institutions.

Because you cannot rerun European colonialism on a randomized grid of planets, macro-historical instruments almost inevitably fail the exclusion restriction.

Rung 5: The Intersubjective & Aesthetic ("Art" / The Evaluative)

At the base of our ladder sits the vast realm of creative expression, literary analysis, visual arts, and normative philosophy.

💡 Constructive Pushback: Is Art "Just Subjective"?

It is tempting to treat Rung 5 as a mere wastebasket of "arbitrary personal opinion" where evidence cannot exist. But this is a mistake.

While art lacks empirical falsifiability or statistical stationarity, it possesses rigorous internal standards of validity:

  • Internal Coherence (David Lewis’s Truth in Fiction): In Tolkien’s Middle-earth, it is an objective, verifiable fact that Frodo carried the One Ring to Mordor, and false that he had a laser rifle. Fictional worlds are governed by authorial stipulation and modal logic.
  • Intersubjective Resonance: A great novel or musical composition is not random noise; it compresses universal structures of human psychology, grief, tension, and release. If an author writes dialogue that rings false to every human reader, the critique is not merely arbitrary—it identifies a failure of affective modeling.
  • Formal Optimization: A sonnet, a fugue, or a cinematography sequence solves severe formal constraints (meter, voice-leading, compositional framing) that optimize for perceptual clarity and emotional transmission.

The error is not that art is “inferior” to physics; the error is trying to judge art by the rules of physics, or trying to settle macroeconomic monetary policy using aesthetic intuition.

The Diagnostic Matrix

Whenever you enter a technical or policy disagreement, determine where each participant is standing on the ladder:

Rung Core Domain Primary Warrant Degrees of Freedom Signature Vulnerability
1. Formal Deduction Math, Logic, Proof Verification Syntactic derivation, kernel checking Zero Axiomatic inconsistency, unnoticed assumptions
2. Stationary Physics QED, Classical Mechanics, Chemistry 5σ detection, invariant physical constants Minimal Social anchoring, systematic apparatus bias
3. Complex Empirical Biology, Pharmacology, Oncology Double-blind RCTs, GRADE meta-analyses High Translational failure, biological heterogeneity
4. Reflexive Social Economics, Sociology, Geopolitics Lotteries, natural experiments, diff-in-diff Very High Lucas critique, broken exclusion restrictions
5. Intersubjective Art, Literature, Ethics, Philosophy Internal coherence, affective resonance Unbounded Solipsism, ungrounded semantic drift
The Rule of Evidential Symmetry Never demand a higher epistemic standard than a domain’s degrees of freedom allow, and never accept a lower epistemic standard than its instrumentation can deliver.

Conclusion: Restoring Coherence to Debate

When someone claims an economic paper is “unscientific” because its findings may not hold in another country 20 years later, they are confusing Rung 4 with Rung 2. Human societies evolve; electrons do not.

Conversely, when a proponent of an unproven medical treatment insists that their personal clinical intuition outweighs a 5,000-patient randomized trial, they are trying to import Rung 5 subjective impressions into Rung 3 empirical medicine.

By making the ladder explicit, we can stop talking past one another and start asking the only question that matters: Given the noise and reflexivity of this system, what is the best evidence humanly possible?