Preprint

In one IV simulation, empirical Bayes had the lowest GMM error

Preprint: The method had the lowest error in a baseline simulation and was also applied to schooling data.

An empirical Bayes method recorded the lowest error among three approaches in a baseline simulation of a misspecified instrumental-variable model, according to a statistical preprint. The method’s root mean squared error, which combines bias and variability, was 0.242, compared with 0.254 for feasible bias correction and 0.287 for standard generalized method of moments, or GMM.

The result comes from a specific simulated design. The paper also examines one schooling-data application, where conventional estimates changed sharply across control specifications and the corrected estimates were less dispersed.

A method for related moment-condition errors

The paper studies GMM when its extra, overidentifying moment conditions contain exchangeable specification errors that shrink at an inverse-square-root rate as the sample grows. Exchangeable, in this setting, means the errors are modeled as having a comparable pattern across the moment conditions.

The paper develops estimators of specification-error mean and variance, a feasible bias-corrected estimator, an empirical Bayes estimator, and variance estimators designed for inference that allows for misspecification. Empirical Bayes is a shrinkage approach: it uses information shared across the errors when forming the estimate.

The theory applies when the number of moment conditions, m, and the sample size, n, grow together under the requirement that m²/n tends to zero. That condition excludes settings where many-instrument bias becomes a first-order problem. Within the paper’s asymptotic framework, empirical Bayes has weakly lower leading-term risk than bias correction, with a nonzero improvement under the specified covariance condition.

The simulated test

The baseline Monte Carlo study evaluated the slope in a model with 10,000 observations and 40 cell moments, using 1,000 replications. In that design, both correction methods had lower RMSE than standard GMM, and empirical Bayes was slightly lower than feasible bias correction.

The comparison also changed the picture for statistical testing. A Wald test based on standard GMM rejected 38.6% of the time. The rejection rates were 8.7% for feasible bias correction and 8.8% for feasible empirical Bayes; versions using the true, or oracle, correction had rates of 4.3% and 4.6%, respectively.

A schooling application with shifting conventional estimates

The empirical application analyzed 486,926 men from the 1940–1949 birth cohorts in the 1980 Census 5% PUMS, using 51 state-of-birth instruments. The paper compared four control specifications.

The conventional two-stage least-squares, or TSLS, estimates were 0.0424 with state of birth alone, 0.1413 with state of birth and age, 0.0360 with state of birth plus covariates, and 0.1240 with those covariates and age. The corresponding ordinary least-squares estimates were 0.0535, 0.0555, 0.0495 and 0.0513. Thus, in this application, the TSLS results moved substantially when age controls were added, while the OLS estimates stayed near 5%.

The corrected estimates were closer together across the four specifications. Bias correction produced 0.0876, 0.1108, 0.0678 and 0.0927; empirical Bayes produced 0.0846, 0.1103, 0.0769 and 0.0935. These estimates do not establish the true causal return to education, because the application is observational.

What the diagnostics found—and left open

Feature-based tests of the exchangeability assumption produced p-values of 0.90, 0.94, 0.83 and 0.92 across the four specifications. Hausman-type tests produced p-values of 0.88, 0.98, 0.69 and 0.96. The reported tests therefore provided no evidence against exchangeability in these specifications, but they cannot rule out every form of non-exchangeability.

The estimated mean specification errors were −5.39, 4.49, −3.67 and 4.75 across the four specifications. The estimated variances were 50.97, 0.00, 0.10 and 0.00. Empirical Bayes standard errors were 14% to 26% smaller than bias-corrected standard errors, including in specifications where the estimated specification-error variance was zero.

The method depends on a context-specific scaling that makes the specification errors plausibly exchangeable, and its consistency theory requires the moderately overidentified condition m²/n → 0. The finite-sample evidence comes from the specified simulated IV design and one Census application, so the results do not establish finite-sample superiority across all GMM or IV settings.

Open questions include how the procedures behave when the mean specification error is weakly identified and how robust they are to alternative moment scalings and dependence structures. The document is an arXiv version 1 preprint dated 25 August 2026. It was prepared for the 2026 Sargan lecture and acknowledges feedback and research assistance.

Paper data and sources

Original title: Repairing Locally Misspecified GMM: An Empirical Bayes Approach
Authors: Patrick Kline
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.