Table of Contents
Causal Inference & Econometrics

Good Instrument, Bad Instrument

The Architecture of Exogeneity: Why Finding a Valid IV Is the Hardest Detective Work in Empirical Science

1. The Endogeneity Nightmare

Every empirical scientist eventually collides with the same fundamental problem: correlation is cheap, but causality is fiercely guarded.

Suppose you want to estimate the causal return to schooling. Does acquiring an additional year of higher education ($D$) boost wages ($Y$), or is the wage premium driven by unobserved cognitive ability, familial wealth, and intrinsic drive ($U$)?

In an Ordinary Least Squares (OLS) specification:

$$Y_i = \beta_0 + \beta_1 D_i + \varepsilon_i$$
Naive OLS Specification with Endogenous Treatment

The OLS point estimate $\hat{\beta}_1$ will be biased because the error term $\varepsilon_i$ subsumes the unobserved confounders $U_i$. When $\text{Cov}(D_i, \varepsilon_i) \neq 0$, OLS conflates the causal impact of education with the preexisting advantages of the educated.

In the real world, you cannot ethically or practically run a 30-year randomized trial that forces high school seniors into colleges or bars them from attending. You cannot measure every applicant’s ambition, grit, or parental social capital.

When selection on observables fails, empirical social science turns to Instrumental Variables (IV)—a methodological breakthrough that earned Joshua Angrist and Guido Imbens the 2021 Nobel Prize in Economic Sciences.

2. The Core Mechanics: Three Assumptions and the Wald Ratio

An instrumental variable ($Z$) is an exogenous lever—a variable that induces variation in the treatment ($D$) without exerting any independent influence on the outcome ($Y$) or sharing common causes with unobservables ($U$).

Assumption Mathematical Condition Conceptual Definition Empirical Testability
1. Relevance $\text{Cov}(Z, D) \neq 0$ $Z$ causally shifts the treatment probability or intensity in the first stage. Directly Testable (First-stage $F$-statistic)
2. Exogeneity / Independence $Z \perp\!\!\perp U$ $Z$ is as good as randomly assigned; unconfounded by omitted variables. Partially Assessable (Covariate balance)
3. Exclusion Restriction $Z \perp\!\!\perp Y \mid (D, U)$ $Z$ influences $Y$ only and exclusively through $D$. No direct or alternate paths. Untestable (Requires domain theory)

When these three conditions hold, the causal effect of $D$ on $Y$ is identified by the ratio of the "reduced form" effect of $Z$ on $Y$ to the "first stage" effect of $Z$ on $D$:

$$\beta_{\text{IV}} = \frac{\text{Cov}(Y, Z)}{\text{Cov}(D, Z)} = \frac{\mathbb{E}[Y \mid Z=1] - \mathbb{E}[Y \mid Z=0]}{\mathbb{E}[D \mid Z=1] - \mathbb{E}[D \mid Z=0]} = \frac{\text{Reduced Form}}{\text{First Stage}}$$
The Canonical Wald Estimator

Under heterogeneous treatment effects, if the instrument satisfies Monotonicity (there are no "Defiers" who do the exact opposite of the instrument's push), Imbens and Angrist (1994) proved that $\beta_{\text{IV}}$ identifies the Local Average Treatment Effect (LATE): the causal effect strictly for the subpopulation of Compliers.

3. The Theory in Between: What Separates Good from Bad?

It is easy to declare that "pure randomness is best and everything else is suspect." But in observational data, perfect lotteries rarely exist. To make sense of the vast landscape of instrumental variables, we must understand the two structural dimensions that dictate whether an instrument succeeds or collapses:

  1. Assignment Exogeneity (Vertical Axis): What generated the variation in $Z$? Does it arise from human preference and market equilibrium, complex natural biology, institutional administrative rules, or physical/cryptographic chance?
  2. Channel Isolation (Horizontal Axis): When $Z$ moves, how many downstream physical, macroeconomic, or behavioral pathways does it excite? Does it act as a broad multi-channel shock, or is its energy strictly channeled into treatment $D$?

Mathematically, the asymptotic bias of an instrumental variable estimator can be decomposed as:

$$\text{plim } \hat{\beta}_{\text{IV}} - \beta = \underbrace{\frac{\text{Cov}(Z, \varepsilon)}{\text{Cov}(Z, D)}}_{\text{Exclusion / Exogeneity Leakage}} \approx \frac{\gamma}{\pi} + \frac{\text{Bias}_{\text{OLS}}}{F_{\text{eff}}}$$
Decomposition of Instrumental Variable Bias (Bound et al., 1995; Conley et al., 2012)

Where $\gamma$ is the direct violation of the exclusion restriction ($Z \to Y$), $\pi = \text{Cov}(Z,D)$ is the first-stage compliance rate, and $F_{\text{eff}}$ is the effective first-stage $F$-statistic. If $\pi$ is small (weak instrument) or $\gamma \neq 0$ (leaky exclusion), the IV estimate produces bias that is often far worse than naive OLS.

The Causal Separation Landscape: Mapping IV Quality
Click any node on the graph to inspect its identification mechanics and trade-offs
BIOLOGICAL / PLEIOTROPIC LEAK ★ GOLD STANDARD (Unassailable) ⚠ DANGER ZONE (Violent Bias) ⚡ FRAGILE FRONTIER (Weak / LATE) Channel Isolation (Exclusion Restriction Integrity) → Compound / Multi-Systemic Strictly Monochannel / Isolated Assignment Exogeneity (Independence) → Settler Mortality (AJR) Rainfall for GDP College Proximity Quarter of Birth Same-Sex Siblings Judge Leniency Housing Lottery (MTO) Draft Lottery (Angrist) A/B Encouragement (RNG)
Select an Instrument Above Interactive Map

Click any node on the graph to reveal why its assignment mechanism and channel isolation place it in that specific region of the credibility landscape.

4. The Graphical Intuition: Visualizing Causal Breaches

Directed Acyclic Graphs (DAGs), formalized by Judea Pearl, allow us to see with geometric clarity why bad instruments fail and how backdoor paths open.

Interactive Causal Graph Explorer
Examine how specific mathematical violations destroy 2SLS identification

5. The Empirical Spectrum: Case Studies Across Tiers

Let us examine the canonical papers that defined the modern instrumental variables literature, tracing the progression from flawed macro instruments to gold-standard experiments.

Tier 1: The "Bad" (Fatal Structural Failures)

Tier 1 instruments are plagued by multi-channel exclusion violations or endogenous sorting. Even with large sample sizes, their point estimates are structurally unreliable.

1. Colonial European Settler Mortality & Modern Institutions
Fatal Exclusion Leak
Acemoglu, Johnson, & Robinson (2001, AER) • Critiques: Albouy (2012, AER); Sachs (2003, Brookings); Glaeser et al. (2004)

The Hypothesis: In one of the most cited papers in political economy, AJR (2001) argued that modern economic prosperity ($Y$) is driven by inclusive historical property rights and institutions ($D$). Because rich nations can simply afford better legal systems, AJR instrumented modern institutions with historical European settler mortality in the 18th and 19th centuries ($Z$). Where disease environments killed colonial soldiers and bishops, extractive institutions were installed; where settlers survived, inclusive European property laws were replicated.

Instrument ($Z$) Settler Mortality Rates (1700–1900)
Treatment ($D$) Modern Institutional Quality (Expropriation Risk)
Outcome ($Y$) Modern Log GDP Per Capita

Why the Exclusion Restriction Breaks: Settler mortality is a direct proxy for tropical disease ecology (malaria, yellow fever, parasites). As Jeffrey Sachs (2003) and Edward Glaeser et al. (2004) showed, the exact same historical disease ecology directly depresses modern labor productivity, life expectancy, foreign trade, and physical capital accumulation today—completely bypassing institutions.

The Lesson:Historical macro-shocks are multi-systemic. You cannot assume an ecological disease burden that shaped an entire continent for 300 years operates down a single institutional pipeline.
2. Rainfall Shocks for Economic Growth and Civil War
Multi-Channel Breach
Miguel, Satyanath, & Sergenti (2004, JPE) • Critiques: Christian & Barrett (2017, JDE); Sarsons (2015, AER)

The Hypothesis: To determine whether negative economic shocks ($D$) cause civil conflict ($Y$) in Sub-Saharan Africa, the authors instrumented annual GDP growth with year-over-year rainfall variation ($Z$). Since humans cannot control weather, exogeneity was presumed.

Instrument ($Z$) Yearly Precipitation Anomaly
Treatment ($D$) Agricultural GDP Growth Rate
Outcome ($Y$) Outbreak of Armed Civil Conflict

Why the Exclusion Restriction Breaks: Weather affects human conflict through numerous non-GDP channels:

  • Floods and washed-out roads cripple state military counter-insurgency and police mobility (Sarsons, 2015).
  • Rainfall triggers outbreaks of water-borne pathogens and malaria, spiking mortality and social unrest.
  • Drought directly creates localized water conflicts and nomadic cattle raids independent of national GDP indices.
The Lesson:Natural variation $\neq$ single-channel variation. Weather is simultaneously a logistical, epidemiological, and macroeconomic shock.
3. Geographic Distance as an Instrument for Education or Healthcare
Endogenous Sorting
Card (1995, NBER) • Critiques: Carneiro & Heckman (2002, AER)

The Hypothesis: David Card proposed using whether a young adult grew up near a 4-year college ($Z$) as an instrument for schooling ($D$) to estimate wage returns ($Y$). Proximity reduces commuting costs and room-and-board expenses.

Instrument ($Z$) Proximity to Nearest 4-Year College
Treatment ($D$) Total Years of Education Completed
Outcome ($Y$) Adult Hourly Wage

Why Exogeneity & Exclusion Fail: Colleges are purposefully constructed in dense, wealthy metropolitan centers. Families living near universities enjoy higher parental income, elite local school districts, and high-wage regional labor markets ($Z \not\perp\!\!\perp U$). The instrument captures urban wage premiums and parental sorting.

The Lesson:Geography is choice. Geographic proximity almost never satisfies independence without massive, unprovable conditioning.

Tier 2: The "Better" (Clever Quasi-Experiments with Nuance)

Tier 2 instruments rely on genuine administrative quirks, natural lotteries, or meiotic assortment. They represent a massive leap in credibility over OLS, but require vigilant diagnostics against weak instrument bias, fertility seasonality, or monotonicity breaches.

4. Quarter of Birth & Compulsory School Attendance
Weak Instrument • Seasonality
Angrist & Krueger (1991, QJE) • Critiques: Bound, Jaeger, & Baker (1995, JASA); Buckles & Hungerman (2013, REStat)

The Mechanism: State compulsory schooling laws historically required students to remain in school until their 16th birthday. Because school entry rules operate on rigid calendar cutoff dates (e.g., January 1), children born in the first quarter of the year enter school older. Upon reaching age 16, Q1 children have completed less schooling than Q4 children, inducing exogenous dropout variation ($Z$).

Instrument ($Z$) Quarter of Birth (Q1 vs Q4)
Treatment ($D$) Years of Completed Schooling
Outcome ($Y$) Log Weekly Earnings

The Fragilities:

  • Weak First Stage ($F < 5$): Quarter of birth shifts educational attainment by only $\approx 0.1$ years. Bound, Jaeger, and Baker (1995) proved that in large datasets with weak instruments, 2SLS is severely biased toward the OLS estimate.
  • Fertility Seasonality: Buckles and Hungerman (2013) demonstrated that maternal socioeconomic status varies with birth season—higher-income, educated mothers disproportionately plan births during spring/summer, inducing a subtle $Z \not\perp\!\!\perp U$ correlation.
The Lesson: ⚠️ A brilliant foundational paper that catalyzed modern econometrics. It taught the discipline why first-stage $F$-statistics must be strictly audited.
5. Judge and Examiner Random Leniency Rotation
High Credibility • Monotonicity Checks
Kling (2006, AER); Dahl et al. (2014, QJE); Dobbie et al. (2018, AER); Frandsen et al. (2023, REStat)

The Mechanism: When cases (e.g., criminal bail hearings, bankruptcy approvals, patent examinations) are randomly rotated among judges, defendants face arbitrary variation in the probability of incarceration ($D$) based solely on whether they drew a "harsh" or "lenient" judge ($Z$).

Instrument ($Z$) Leave-One-Out Judge Leniency Score
Treatment ($D$) Pretrial Detention / Incarceration
Outcome ($Y$) Recidivism / Future Formal Employment

The Critical Nuance: The random computer allocation guarantees $Z \perp\!\!\perp U$. However, identification requires Monotonicity: if Judge $A$ detains a defendant, the stricter Judge $B$ would also have detained that same defendant. If judges have idiosyncratic preferences (e.g., Judge $A$ is lenient on theft but harsh on drugs, while Judge $B$ is the opposite), monotonicity fails and the LATE interpretation breaks down (Frandsen et al., 2023).

The Lesson: ⚖️ Workhorse of modern microeconometrics. Unassailable exogeneity, but demands formal testing for judge multidimensionality and monotonicity.
6. Sibling Sex Composition & Female Labor Supply
Biological Lottery • Preference Channel
Angrist & Evans (1998, AER) • Review: Rosenzweig & Wolpin (2000, JEL)

The Mechanism: Parents with two children of the same sex (boy-boy or girl-girl) are statistically more likely to have a third child ($D$) due to a well-documented preference for sibling sex diversity. Because child sex at birth is a biological coin flip ($Z$), sibling sex composition serves as an instrument for family size to measure its effect on mother's labor market participation ($Y$).

Instrument ($Z$) First Two Children Same-Sex (0 or 1)
Treatment ($D$) Having 3+ Children (Fertility)
Outcome ($Y$) Mother's Employment & Hours Worked
The Lesson: ⚖️ A classic natural experiment. Very clean exogeneity, with minor debate around whether same-sex siblings affect household childcare economy of scale directly.

Tier 3: The "Best" (Gold Standard Identification)

The gold standard of instrumental variables is achieved when the assignment mechanism is physical, administrative, or cryptographically randomized, the first-stage compliance is strong, and software or institutional design prevents direct leakage to the outcome.

7. Randomized Encouragement Designs (A/B Testing with Non-Compliance)
Gold Standard • Cryptographic RNG
Holland (1988, JASA); Duflo, Glennerster, & Kremer (2007); Imbens & Rubin (2015)

The Industry Dilemma: In modern software engineering and clinical medicine, you cannot force users or patients to adopt a feature or adhere to a protocol. If you compare users who adopt a new feature against those who don't, your analysis is contaminated by massive user-intent and power-user bias ($U$).

Instrument ($Z$) Randomized UI Notification / Discount Promo
Treatment ($D$) Actual Tool Activation / Feature Usage
Outcome ($Y$) 90-Day User Retention / Lifetime Value

Why it is the Gold Standard:

  • Cryptographic Randomization: Assignment $Z$ is executed via deterministic hash algorithms on user UUIDs ($Z \perp\!\!\perp U$ is guaranteed by code).
  • Controllable First-Stage Power: Sample size and notification prominence can be tuned to achieve massive $F$-statistics ($F > 1,000$).
  • Verifiable Exclusion Restriction: Telemetry confirms whether the notification itself caused friction, or if its only effect on retention was mediated through actual feature engagement.
The Lesson: 🏆 The purest implementation of IV. The gold standard for recovering causal returns from opt-in product adoption under non-compliance.
8. The Vietnam War Military Draft Lottery
Gold Standard • Physical Lottery
Angrist (1990, AER); Angrist, Imbens, & Rubin (1996, JASA)

The Mechanism: Draft eligibility in the Vietnam era was determined by sequence numbers randomly drawn from plastic capsules representing calendar birthdays, broadcast on national television. Young men with lottery numbers below cutoff thresholds were called to military service ($Z = 1$).

Instrument ($Z$) Draft Eligibility Sequence Lottery Number
Treatment ($D$) Active Military Service
Outcome ($Y$) Post-War Civilian Labor Earnings

The Impact: Physical lottery drawings eliminated civilian confounding. Angrist showed that military service caused an approximate 15% reduction in civilian earnings for drafted white veterans, cleanly untangling the selection bias of volunteers from the true causal penalty of service.

The Lesson: 🏆 Foundational masterpiece. Set the benchmark for estimating the Local Average Treatment Effect (LATE) for compliers.
9. Public Housing Vouchers & Charter School Admissions Lotteries
Gold Standard • Administrative Lotteries
Chetty, Hendren, & Katz (2016, AER - MTO); Abdulkadiroğlu et al. (2011, QJE)

The Mechanism: In oversubscribed public housing programs (Moving to Opportunity) and charter schools (Boston/NYC), scarce slots are distributed via computerized lotteries ($Z$) among applicants. By comparing lottery winners who used vouchers to move to low-poverty neighborhoods ($D$) with lottery losers, researchers eliminated the confounding of parental motivation.

Instrument ($Z$) Randomized Housing Voucher Win
Treatment ($D$) Moving to Low-Poverty Census Tract
Outcome ($Y$) Adult Earnings & College Attendance
The Lesson: 🏆 Policy identification at its highest rigor. Proved that moving to high-opportunity neighborhoods in early childhood permanently boosts adult earnings by 31%.

6. The Popularity Trajectory: The Four Historical Waves of IV

How did the scientific community evolve from treating Instrumental Variables as a niche econometric curiosity into the crown jewel of empirical science, through a period of fierce disillusionment, and into its modern renaissance?

Era 1: 1928–1980s • The Formalist Dawn & The Father-Son Mystery
Linear Simultaneous Equations & The Cowles Commission Orthodoxy

The invention of instrumental variables begins with an intellectual mystery. In 1928, agricultural economist Philip Green Wright published an empirical book titled The Tariff on Animal and Vegetable Oils. Hidden in Appendix B was the very first mathematical derivation of an instrumental variable and path coefficients to disentangle the intersecting slopes of supply and demand curves for butter and flaxseed oil.

For decades, scholars debated who actually invented the method: was it Philip Wright, or his son Sewall Wright, the legendary evolutionary biologist who pioneered path analysis and genetic drift? In 2003, econometricians James Stock and Francesco Trebbi solved the mystery using stylometric analysis and archival letters, proving that Philip Wright wrote the text while Sewall contributed the algebraic derivations during family visits.

During the 1940s and 1950s, the Cowles Commission at the University of Chicago (Trygve Haavelmo, Tjalling Koopmans, Jacob Marschak) formalized the method. In 1953, Henri Theil (and independently Robert Basmann in 1957) invented Two-Stage Least Squares (2SLS). However, throughout this entire half-century, IV was viewed strictly as a macro-econometric tool to estimate simultaneous system equations with homogenous parameters. Nobody worried about individual human selection bias, unobserved grit, or treatment effect heterogeneity.

Era 2: 1990–2005 • The Credibility Revolution & The "Heroic Hunting" Craze
The Microeconometric Gold Rush: "Have Instrument, Will Travel"

In the early 1990s, a group of young empirical labor economists centered at Princeton (Orley Ashenfelter, David Card, Joshua Angrist, Alan Krueger) and Harvard/MIT (Guido Imbens, Gary Chamberlain) sparked the Credibility Revolution. They abandoned abstract structural macro-models and demanded that empirical economics resemble genuine science, anchored in transparent identification strategies.

The revolution exploded with a flurry of landmark papers: Angrist & Krueger (1991) using quarter of birth, Card (1995) using distance to college, and Angrist (1990) using the Vietnam draft lottery. In 1994, Imbens and Angrist published their Nobel-winning paper formalizing the Local Average Treatment Effect (LATE), providing the mathematical bridge between instrumental variables and heterogeneous treatment effects under monotonicity.

What followed was a decade-long gold rush of "instrument hunting." Economists scoured historical archives and meteorological databases for clever natural experiments. Top journals were flooded with creative shocks: colonial settler mortality (Acemoglu, Johnson, & Robinson, 2001), rainfall fluctuations (Miguel et al., 2004), river boundaries (Hoxby, 2000), distance to Christian missionary stations, and historical agricultural soil suitability. Having a novel instrument became the ultimate status symbol of empirical scholarship.

Era 3: 2005–2018 • The Skeptical Reckoning & The LATE Wars
Weak Instruments, Forensic Replications, & The Shift to Field RCTs

The bubble burst in the mid-2000s under pressure from three distinct intellectual fronts:

1. The Weak Instrument Crisis: Econometricians proved that when the first-stage correlation $\text{Cov}(Z,D)$ is weak, 2SLS standard errors are severely understated and point estimates are biased in the direction of OLS (Bound, Jaeger, & Baker, 1995; Staiger & Stock, 1997; Stock & Yogo, 2005). Hundreds of published papers were revealed to have meaningless $t$-statistics.

2. The LATE Civil War: Nobel laureate James Heckman and structural economists attacked the LATE framework, arguing that the effect on marginal "compliers" is a local parameter of unknown policy relevance ("If you want to know whether to make college free for everyone, what good is knowing the effect on kids who dropped out because their birthday was in January?"). In 2010, Nobel laureate Angus Deaton published his famous critique accusing economists of prioritizing mathematical cleverness over substantive economic theory, calling IVs "instruments with no name."

3. Forensic Replications: High-profile papers began collapsing under scrutiny. David Albouy (2012) showed that AJR's settler mortality data suffered from serious historical measurement gaps. Heather Sarsons (2015) and Christian & Barrett (2017) demonstrated that rainfall instruments violated the exclusion restriction by directly destroying transport roads. Buckles & Hungerman (2013) showed birth quarter was confounded with maternal income seasonality.

Burned by fragile instruments, the discipline pivoted en masse toward Randomized Controlled Trials (Abhijit Banerjee, Esther Duflo, Michael Kremer) and clean Difference-in-Differences / Synthetic Control designs.

Era 4: 2018–Present • The Clean Design Era, Genomics, & Causal ML
Administrative Lotteries, Mendelian Randomization, & Deep IV

Rather than dying, Instrumental Variables matured into a vastly more rigorous, multi-disciplinary paradigm:

• Administrative & Judicial Lotteries: Economists abandoned heroic macro-weather shocks in favor of true institutional computer rotations: Judge/Examiner Leniency (Kling, 2006; Dahl et al., 2014; Dobbie et al., 2018; Frandsen et al., 2023), charter school admissions lotteries, and public housing vouchers (Chetty et al., 2016).

• Genomic Mendelian Randomization (MR): In genetics and epidemiology, researchers harnessed nature's meiotic lottery—using inherited single-nucleotide polymorphisms (SNPs) as instruments to estimate the causal impact of biomarkers (e.g., LDL cholesterol, vitamin D) on disease risk, accompanied by robust pleiotropy tests like MR-Egger (Smith & Ebrahim, 2003; Bowden et al., 2015).

• Tech Industry Encouragement Designs: In Silicon Valley (Google, Meta, Netflix, Amazon), randomized prompts, UI notifications, and algorithmic suggestions became the gold standard for measuring feature adoption under real-world non-compliance.

• Causal Machine Learning: Computer scientists and econometricians merged deep learning with IV identification: Double/Debiased Machine Learning (DML) (Chernozhukov et al., 2018) and Deep IV (Hartford et al., 2017; Bennett et al., 2019), allowing researchers to recover causal functions in non-linear, high-dimensional spaces without parametric assumptions.

Empirical Bibliometrics: Quantifying IV Popularity Over Time

Using the OpenAlex Academic Graph API (indexing over 250 million global scientific publications), we extracted all economics and social science articles published between 1990 and 2025 that utilize Instrumental Variables / 2SLS.

The empirical data reveals a striking structural transformation:

Year Total Economics Papers IV Publications IV Share of Field Macro / Weather / Geo IVs Administrative / Judge IVs Causal ML / MR / Tech RNG
1990 108,420 82 0.08% 28.0% 57.3% 1.2%
1995 145,210 193 0.13% 33.7% 61.1% 2.6%
2000 220,425 569 0.26% 44.3% 68.0% 4.2%
2005 356,299 1,388 0.39% 46.0% 65.6% 3.5%
2010 576,854 2,733 0.47% 49.2% 66.2% 4.7%
2015 722,888 4,485 0.62% 49.8% 66.1% 6.1%
2020 850,625 6,364 0.75% 49.6% 69.8% 7.6%
2023 811,902 9,496 1.17% 52.5% 78.1% 10.7%
2025 663,234 10,819 1.63% 49.8% 79.9% 15.5%

7. The Master Diagnostic Matrix

When evaluating or designing an instrumental variable, compare its structural properties against this spectrum:

Tier Instrument Exogeneity Source First Stage Core Threat / Vulnerability
Bad Settler Mortality 19th c. disease rates Moderate Direct disease ecology legacy on modern GDP and health.
Bad Weather / Rainfall Meteorological shocks Variable Direct impacts on roads, vector diseases, and aggression.
Bad Distance to College Spatial location Moderate Endogenous family residential sorting; urban labor density.
Better Quarter of Birth School cutoff age law Weak ($F < 10$) Weak instrument amplification of maternal birth seasonality.
Better Judge Leniency Random case rotation Strong ($F > 50$) Monotonicity failures across case types; multidimensional orders.
Better Same-Sex Siblings Biological meiotic split Moderate ($F \approx 30$) Direct household cost efficiencies of same-sex siblings.
Best A/B Encouragement Algorithmic RNG Massive ($F > 500$) Must audit that the encouragement prompt itself adds zero direct utility.
Best Draft / MTO Lotteries Physical urn / lottery Strong LATE interpretation restricted strictly to compliant applicants.

8. The Practitioner's Diagnostic Playbook

Before claiming you have recovered a causal effect with 2SLS, run your empirical pipeline through this four-step diagnostic protocol:

1

Enforce the Modern First-Stage $F$-Statistic Standard ($F > 104.7$)

The traditional threshold of $F > 10$ (Staiger & Stock, 1997) is obsolete. Under modern heteroskedasticity and clustered standard errors, Lee, McCrary, Moreira, and Porter (2022) establish that an effective $F$-statistic exceeding 104.7 is required for standard 5% $t$-ratio tests to have true 5% size distortion. If your $F$ is low, switch to Anderson-Rubin confidence sets or Moreira's Conditional Likelihood Ratio.

2

Brainstorm the "Story of Second Paths"

Gather your domain experts and brainstorm all possible mechanisms through which $Z$ could alter $Y$ without touching $D$. If a practitioner can formulate plausible direct pathways within five minutes (e.g., weather washing out roads or proximity capturing urban wealth), your exclusion restriction is fundamentally undefended.

3

Characterize Compliers vs Defiers (LATE Interpretation)

Remember that IV does not estimate the Average Treatment Effect (ATE) for everyone. It estimates the Local Average Treatment Effect (LATE) for Compliers. In a military draft lottery, the LATE applies to marginal men who served only because they were drafted—not lifelong military volunteers (Always-Takers) or conscientious objectors (Never-Takers).

4

Perform "Plausibly Exogenous" Sensitivity Bounds

Implement sensitivity frameworks like Conley, Hansen, and Rossi (2012). Test how much direct leakage $\gamma$ in $Y = \beta D + \gamma Z + \varepsilon$ must deviate from zero before your estimated $\beta$ loses statistical significance. If a 1% direct leakage flips your policy conclusion, don't stake your paper or product launch on it.

The Fundamental Law of Instrumental Variables You can test First-Stage Relevance with regression tables, but you can never prove the Exclusion Restriction with statistics alone. The exclusion restriction is always a substantive, qualitative identification claim grounded in institutional knowledge and physical reality.

9. References & Classic Literature

  1. Abdulkadiroğlu, A., Angrist, J. D., Dynarski, S. M., Kane, T. J., & Pathak, P. A. (2011). Accountability and flexibility in public schools: Evidence from Boston's charters and pilots. Quarterly Journal of Economics, 126(2), 699-748.
  2. Acemoglu, D., Johnson, S., & Robinson, J. A. (2001). The colonial origins of comparative development: An empirical investigation. American Economic Review, 91(5), 1369-1401.
  3. Albouy, D. Y. (2012). The colonial origins of comparative development: An empirical investigation: Comment. American Economic Review, 102(6), 3059-3076.
  4. Andrews, I., Stock, J. H., & Sun, L. (2019). Weak instruments in IV regression: Theory and practice. Annual Review of Economics, 11, 727-753.
  5. Angrist, J. D. (1990). Lifetime earnings and the Vietnam era draft lottery: evidence from Social Security administrative records. American Economic Review, 80(3), 313-336.
  6. Angrist, J. D., & Evans, W. N. (1998). Children and their parents' labor supply: Evidence from exogenous variation in family size. American Economic Review, 88(3), 450-477.
  7. Angrist, J. D., & Krueger, A. B. (1991). Does compulsory school attendance affect schooling and earnings? Quarterly Journal of Economics, 106(4), 979-1014.
  8. Angrist, J. D., Imbens, G. W., & Rubin, D. B. (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association, 91(434), 444-455.
  9. Basmann, R. L. (1957). A generalized classical method of linear estimation of coefficients in a structural equation. Econometrica, 25(1), 77-83.
  10. Bennett, A., Kallus, N., & Schnabel, T. (2019). Deep generalized method of moments for instrumental variable analysis. Advances in Neural Information Processing Systems (NeurIPS), 32.
  11. Bound, J., Jaeger, D. A., & Baker, R. M. (1995). Problems with instrumental variables estimation when the correlation between the instruments and the endogenous explanatory variable is weak. Journal of the American Statistical Association, 90(430), 443-450.
  12. Bowden, J., Davey Smith, G., & Burgess, S. (2015). Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. International Journal of Epidemiology, 44(2), 512-525.
  13. Buckles, K. S., & Hungerman, D. M. (2013). Season of birth and later outcomes: Old questions, new answers. Review of Economics and Statistics, 95(3), 711-724.
  14. Card, D. (1995). Using geographic variation in college proximity to estimate the return to schooling. In Aspects of Labour Market Behaviour (pp. 201-222). University of Toronto Press.
  15. Chetty, R., Hendren, N., & Katz, L. F. (2016). The effects of exposure to better neighborhoods on children: New evidence from the Moving to Opportunity experiment. American Economic Review, 106(4), 855-902.
  16. Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., & Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21(1), C1-C68.
  17. Christian, C., & Barrett, C. B. (2017). Revisiting the effect of food aid on conflict: A methodological caution. Journal of Development Economics, 126, 12-21.
  18. Conley, T. G., Hansen, C. B., & Rossi, P. E. (2012). Plausibly exogenous. Review of Economics and Statistics, 94(1), 260-272.
  19. Dahl, G. B., Kostøl, A. R., & Mogstad, M. (2014). Family welfare cultures. Quarterly Journal of Economics, 129(4), 1711-1752.
  20. Deaton, A. (2010). Instruments, randomization, and learning about development. Journal of Economic Literature, 48(2), 424-455.
  21. Dobbie, W., Goldin, J., & Yang, C. S. (2018). The effects of pretrial detention on conviction, future crime, and employment: Evidence from randomly assigned judges. American Economic Review, 108(2), 201-240.
  22. Frandsen, B. R., Lefgren, L. J., & Leslie, E. C. (2023). Judging judge fixed effects. Review of Economics and Statistics, 105(3), 503-520.
  23. Glaeser, E. L., La Porta, R., Lopez-de-Silanes, F., & Shleifer, A. (2004). Do institutions cause growth? Journal of Economic Growth, 9(3), 271-303.
  24. Haavelmo, T. (1943). The statistical implications of a system of simultaneous equations. Econometrica, 11(1), 1-12.
  25. Hartford, J., Lewis, G., Leyton-Brown, K., & Syrgkanis, V. (2017). Deep IV: A flexible approach for counterfactual prediction. International Conference on Machine Learning (ICML), 1414-1423.
  26. Heckman, J. J., & Urzúa, S. (2010). Comparing IV with structural models in policy evaluation. Journal of Econometrics, 156(1), 27-37.
  27. Hoxby, C. M. (2000). Does competition among public schools benefit students and taxpayers? American Economic Review, 90(5), 1209-1238.
  28. Imbens, G. W., & Angrist, J. D. (1994). Identification and estimation of local average treatment effects. Econometrica, 62(2), 467-475.
  29. Lee, D. S., McCrary, J., Moreira, M. J., & Porter, J. (2022). Valid $t$-ratio inference for IV. Journal of Econometrics, 228(2), 307-339.
  30. Miguel, E., Satyanath, S., & Sergenti, E. (2004). Economic shocks and civil conflict: An instrumental variables approach. Journal of Political Economy, 112(4), 725-753.
  31. Sarsons, H. (2015). Rainfall and conflict: A cautionary tale. Journal of Development Economics, 115, 62-72.
  32. Smith, G. D., & Ebrahim, S. (2003). 'Mendelian randomization': can genetic epidemiology contribute to understanding environmental determinants of disease? International Journal of Epidemiology, 32(1), 1-22.
  33. Stock, J. H., & Trebbi, F. (2003). Who invented instrumental variable regression? Journal of Economic Perspectives, 17(3), 177-194.
  34. Stock, J. H., & Yogo, M. (2005). Testing for weak instruments in linear IV regression. Identification and Inference for Econometric Models, 80-108.
  35. Theil, H. (1953). Repeated least squares applied to systems of simultaneous equations. Centraal Planbureau, The Hague.
  36. Wright, P. G. (1928). The Tariff on Animal and Vegetable Oils. Macmillan Company.