Guido W Imbens
Biographic Data
| ID | 1471513 |
|---|---|
| NAME | Guido W Imbens |
| GIVEN NAMES | Guido W |
| FAMILY NAME | Imbens |
| SIGNATURE | IMBENS G W |
| AFFILIATIONS | Stanford University |
| ORCID | 0000-0002-4846-7326 |
| VERIFIED | Yes |
| TOTAL WORKS | 50 |
| TOTAL CITATIONS | 176 |
| AUTHOR COUNT | 50 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1994 |
| LATEST PUBLICATION YEAR | 2025 |
| H-INDEX | 3 |
Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades after LaLonde (1986)
In 1986, Robert LaLonde published an article comparing nonexperimental estimates to experimental benchmarks (LaLonde 1986). He concluded that the nonexperimental methods at the time could not systematically replicate experimental benchmarks, casting doubt on their credibility. Following LaLonde's critical assessment, there have been significant methodological advances and practical changes, including (1) an emphasis on the unconfoundedness assump…
A Design-Based Perspective on Synthetic Control Methods
Since their introduction by Abadie and Gardeazabal, Synthetic Control (SC) methods have quickly become one of the leading methods for estimating causal effects in observational studies in settings with panel data.Formal discussions often motivate SC methods by the assumption that the potential outcomes were generated by a factor model.Here we study SC methods from a design-based perspective, assuming a model for the selection of the treated unit(…
When Should You Adjust Standard Errors for Clustering?
Clustered standard errors, with clusters defined by factors such as geography, are widespread in empirical research in economics and many other disciplines. Formally, clustered standard errors adjust for the correlations induced by sampling the outcome variable from a data-generating process with unobserved cluster-level components. However, the standard econometric framework for clustering leaves important questions unanswered: (i) Why do we adj…
Design-based analysis in Difference-In-Differences settings with staggered adoption
Synthetic Difference-in-Differences
We present a new estimator for causal effects with panel data that builds on insights behind the widely used difference-in-differences and synthetic control methods. Relative to these methods we find, both theoretically and empirically, that this “synthetic difference-in-differences” estimator has desirable robustness properties, and that it performs well in settings where the conventional estimators are commonly used in practice. We study the as…
Statistical Significance,p-Values, and the Reporting of Uncertainty
The use of statistical significance and p-values has become a matter of substantial controversy in various fields using statistical methods. This has gone as far as some journals banning the use of indicators for statistical significance, or even any reports of p-values, and, in one case, any mention of confidence intervals. I discuss three of the issues that have led to these often-heated debates. First, I argue that in many cases, p-values and …
Identification and Efficiency Bounds for the Average Match Function Under Conditionally Exogenous Matching
Consider two heterogenous populations of agents who, when matched, jointly produce an output, Y. For example, teachers and classrooms of students together produce achievement, parents raise children, whose life outcomes vary in adulthood, assembly plant managers and workers produce a certain number of cars per month, and lieutenants and their platoons vary in unit effectiveness. Let W∈W={w1,...,wJ} and X∈X={x1,...,xK} denote agent types in the tw…
External Validity in Fuzzy Regression Discontinuity Designs
Fuzzy regression discontinuity designs identify the local average treatment effect (LATE) for the subpopulation of compliers, and with forcing variable equal to the threshold. We develop methods that assess the external validity of LATE to other compliance groups at the threshold, and allow for identification away from the threshold. Specifically, we focus on the equality of outcome distributions between treated compliers and always-takers, and b…
Machine Learning Methods That Economists Should Know About
We discuss the relevance of the recent machine learning (ML) literature for economics and econometrics. First we discuss the differences in goals, methods, and settings between the ML literature and the traditional econometrics and statistics literatures. Then we discuss some specific methods from the ML literature that we view as important for empirical researchers in economics. These include supervised learning methods for regression and classi…
Optimized Regression Discontinuity Designs
The increasing popularity of regression discontinuity methods for causal inference in observational studies has led to a proliferation of different estimating strategies, most of which involve first fitting nonparametric regression models on both sides of a treatment assignment boundary and then reporting plug-in estimates for the effect of interest. In applications, however, it is often difficult to tune the nonparametric regressions in a way th…
Understanding and misunderstanding randomized controlled trials: A commentary on Deaton and Cartwright
When Should You Adjust Standard Errors for Clustering?
In empirical work in economics it is common to report standard errors that account for clustering of units.Typically, the motivation given for the clustering adjustments is that unobserved components in outcomes for units within clusters are correlated.However, because correlation may occur across more than one dimension, this motivation makes it difficult to justify why researchers use clustering in some dimensions, such as geographic, but not o…
The State of Applied Econometrics: Causality and Policy Evaluation
In this paper, we discuss recent developments in econometrics that we view as important for empirical researchers working on policy evaluation questions. We focus on three main areas, in each case, highlighting recommendations for applied work. First, we discuss new research on identification strategies in program evaluation, with particular focus on synthetic control methods, regression discontinuity, external validity, and the causal interpreta…
Redefine statistical significance
Matching on the Estimated Propensity Score
Propensity score matching estimators (Rosenbaum and Rubin (1983)) are widely used in evaluation research to estimate average treatment effects. In this article, we derive the large sample distribution of propensity score matching estimators. Our derivations take into account that the propensity score is itself estimated in a first step, prior to matching. We prove that first step estimation of the propensity score affects the large sample distrib…
Recursive partitioning for heterogeneous causal effects
In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals f…
Robust Standard Errors in Small Samples: Some Practical Advice
We study the properties of heteroskedasticity-robust confidence intervals for regression parameters. We show that confidence intervals based on a degrees-of-freedom correction suggested by Bell and McCaffrey (2002) are a natural extension of a principled approach to the Behrens-Fisher problem. We suggest a further improvement for the case with clustering. We show that these standard errors can lead to substantial improvements in coverage rates ev…
Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction
Most questions in social and biomedical sciences are causal in nature: what would happen to individuals, or to groups, if part of their environment were changed? In this groundbreaking text, two world-renowned experts present statistical methods for studying such questions. This book starts with the notion of potential outcomes, each corresponding to the outcome that would be realized if a subject were exposed to a particular treatment or regime.…
Matching Methods in Practice: Three Examples
There is a large theoretical literature on methods for estimating causal effects under unconfoundedness, exogeneity, or selection-on-observables type assumptions using matching or propensity score methods. Much of this literature is highly technical and has not made inroads into empirical practice where many researchers continue to use simple methods such as ordinary least squares regression even in settings where those methods do not have attrac…
Identification and Inference With Many Invalid Instruments
We study estimation and inference in settings where the interest is in the effect of a potentially endogenous regressor on some outcome. To address the endogeneity, we exploit the presence of additional variables. Like conventional instrumental variables, these variables are correlated with the endogenous regressor. However, unlike conventional instrumental variables, they also have direct effects on the outcome, and thus are “invalid” instrument…
Promoting Transparency in Social Science Research
Social scientists should adopt higher transparency standards to improve the quality and credibility of research.
Social Networks and the Identification of Peer Effects
There is a large and growing literature on peer effects in economics. In the current article, we focus on a Manski-type linear-in-means model that has proved to be popular in empirical work. We critically examine some aspects of the statistical model that may be restrictive in empirical analyses. Specifically, we focus on three aspects. First, we examine the endogeneity of the network or peer groups. Second, we investigate simultaneously alternat…
Optimal Bandwidth Choice for the Regression Discontinuity Estimator
We investigate the choice of the bandwidth for the regression discontinuity estimator. We focus on estimation by local linear regression, which was shown to have attractive properties (Porter, J. 2003, “Estimation in the Regression Discontinuity Model” (unpublished, Department of Economics, University of Wisconsin, Madison)). We derive the asymptotically optimal bandwidth under squared error loss. This optimal bandwidth depends on unknown functio…
Bias-Corrected Matching Estimators for Average Treatment Effects
In Abadie and Imbens (2006), it was shown that simple nearest-neighbor matching estimators include a conditional bias term that converges to zero at a rate that may be slower than N1/2. As a result, matching estimators are not N1/2-consistent in general. In this article, we propose a bias correction that renders matching estimators N1/2-consistent and asymptotically normal. To demonstrate the methods proposed in this article, we apply them to the…
Better Late Than Nothing: Some Comments on Deaton (2009) and Heckman and Urzua (2009)
Two recent papers, Deaton (2009) and Heckman and Urzua (2009), argue against what they see as an excessive and inappropriate use of experimental and quasi-experimental methods in empirical work in economics in the last decade. They specifically question the increased use of instrumental variables and natural experiments in labor economics and of randomized experiments in development economics. In these comments, I will make the case that this mov…
Redefine statistical significance
The State of Applied Econometrics: Causality and Policy Evaluation
In this paper, we discuss recent developments in econometrics that we view as important for empirical researchers working on policy evaluation questions. We focus on three main areas, in each case, highlighting recommendations for applied work. First, we discuss new research on identification strategies in program evaluation, with particular focus on synthetic control methods, regression discontinuity, external validity, and the causal interpreta…
Statistical Significance,p-Values, and the Reporting of Uncertainty
The use of statistical significance and p-values has become a matter of substantial controversy in various fields using statistical methods. This has gone as far as some journals banning the use of indicators for statistical significance, or even any reports of p-values, and, in one case, any mention of confidence intervals. I discuss three of the issues that have led to these often-heated debates. First, I argue that in many cases, p-values and …
Identification and Estimation of Local Average Treatment Effects
We investigate conditions sufficient for identification of average treatment effects using instrumental variables. First we show that the existence of valid instruments is not sufficient to identify any meaningful average treatment effect. We then establish that the combination of an instrument and a condition on the relation between the instrument and the participation status is sufficient for identification of a local average treatment effect f…
Transition Models in a Non-Stationary Environment
An alternative form of the proportional hazard model is proposed. It allows one to introduce correlation between exit rates at the same (calendar) time for different individuals. One can, in the context of this model, still allow for, and estimate, duration effects. These should be parametrized. These modifications to the original Cox model are possible by reversing the roles of duration and calendar time. It is argued that flexibility with respe…
Two-Stage Least Squares Estimation of Average Causal Effects in Models with Variable Treatment Intensity
Two-stage least squares (TSLS) is widely used in econometrics to estimate parameters in systems of linear simultaneous equations and to solve problems of omitted-variables bias in single-equation estimation. We show here that TSLS can also be used to estimate the average causal effect of variable treatments such as drug dosage, hours of exam preparation, cigarette smoking, and years of schooling. The average causal effect in which we are interest…
Evaluating the Cost of Conscription in The Netherlands
In this article we investigate the effect of military service in the Netherlands on future earnings. Estimating the cost or benefit of military service is complicated by the complex selection that determines who eventually serves in the military: On the one hand, potential conscripts have to pass medical and psychological examinations before entering the military, and on the other hand numerous (temporary) exemptions exist that can be manipulated…
Identification of Causal Effects Using Instrumental Variables
We outline a framework for causal inference in settings where assignment to a binary treatment is ignorable, but compliance with the assignment is not perfect so that the receipt of treatment is nonignorable. To address the problems associated with comparing subjects by the ignorable assignment—an “intention-to-treat analysis”—we make use of instrumental variables, which have long been used by economists in the context of regression models with c…
Imposing Moment Restrictions from Auxiliary Data by Weighting
In this paper we analyze the estimation of coefficients in regression models under moment restrictions in which the moment restrictions are derived from auxiliary data. The moment restrictions yield weights for each observation that can subsequently be used in weighted regression analysis. We discuss the interpretation of these weights under two assumptions: that the target population (from which the moments are constructed) and the sampled popul…
The role of the propensity score in estimating dose-response functions
Estimation of average treatment effects in observational studies often requires adjustment for differences in pre-treatment variables. If the number of pre-treatment variables is large, standard covariance adjustment methods are often inadequate. Rosenbaum & Rubin (1983) propose an alternative method for adjusting for pre-treatment variables for the binary treatment case based on the so-called propensity score. Here an extension of the propensity…
Estimating the Effect of Unearned Income on Labor Earnings, Savings, and Consumption: Evidence from a Survey of Lottery Players
This paper provides empirical evidence about the effect of unearned income on earnings, consumption, and savings. Using an original survey of people playing the lottery in Massachusetts in the mid-1980's, we analyze the effects of the magnitude of lottery prizes on economic behavior. The critical assumption is that among lottery winners the magnitude of the prize is randomly assigned. We find that unearned income reduces labor earnings, with a ma…
Estimation of Causal Effects using Propensity Score Weighting: An Application to Data on Right Heart Catheterization
Bias From Classical and Other Forms of Measurement Error
We consider the implications of an alternative to the classical measurement-error model, in which the observed, mismeasured data are optimal predictions of the true values, given some information set. In this model, any measurement error is uncorrelated with the reported value and, by necessity, correlated with the true value of interest. In a regression model, such measurement error in the regressor does not lead to bias, whereas measurement err…
Instrumental Variables Estimates of the Effect of Subsidized Training on the Quantiles of Trainee Earnings
The effect of government programs on the distribution of participants' earnings is important for program evaluation and welfare comparisons.This paper reports es- timates of the effects of JTPA training programs on the distribution of earnings.The estimation uses a new instrumental variable (IV) method that measures program impacts on the quantiles of outcome variables.This quantile treatment effects (QTE) estimator accommodates exogenous covaria…
Generalized Method of Moments and Empirical Likelihood
Generalized method of moments (GMM) estimation has become an important unifying framework for inference in econometrics in the last 20 years. It can be thought of as encompassing almost all of the common estimation methods, such as maximum likelihood, ordinary least squares, instrumental variables, and two-stage least squares, and nowadays is an important part of all advanced econometrics textbooks. The GMM approach links nicely to economic theor…
Sensitivity to Exogeneity Assumptions in Program Evaluation
In many empirical studies of the effect of social programs researchers assume that, conditional on a set of observed covariates, assignment to the treatment is exogenous or unconfounded (aka selection on observables). Often this assumption is not realistic, and researchers are concerned about the robustness of their results to departures from it. One approach (e.g., Charles Manski, 1990) is to entirely drop the exogeneity assumption and investiga…
Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score
We are interested in estimating the average effect of a binary treatment on a scalar outcome. If assignment to the treatment is exogenous or unconfounded, that is, independent of the potential outcomes given covariates, biases associated with simple treatment-control average comparisons can be removed by adjusting for differences in the covariates. Rosenbaum and Rubin (1983) show that adjusting solely for differences between treated and control u…
Nonparametric Applications of Bayesian Inference
This article evaluates the usefulness of a nonparametric approach to Bayesian inference by presenting two applications. Our first application considers an educational choice problem. We focus on obtaining a predictive distribution for earnings corresponding to various levels of schooling. This predictive distribution incorporates the parameter uncertainty, so that it is relevant for decision making under uncertainty in the expected utility framew…
Implementing Matching Estimators for Average Treatment Effects in Stata
This paper presents an implementation of matching estimators for average treatment effects in Stata. The nnmatch command allows you to estimate the average effect for all units or only for the treated or control units; to choose the number of matches; to specify the distance metric; to select a bias adjustment; and to use heteroskedastic-robust variance estimators.
The Propensity Score with Continuous Treatments
of the binary treatment propensity score, which we label the generalized propensity score (GPS). We demonstrate that the GPS has many of the attractive properties of the binary treatment propensity score. Just as in the binary treatment case, adjusting for this scalar function of the covariates removes all biases associated with dierences in the covariates. The GPS also has certain balancing properties that can be used to assess the adequacy of p…
Confidence Intervals for Partially Identified Parameters
this paper, we study the use of these intervals as CIs for the partially identified parameter f(P,#). Our most basic finding is Lemma 2.1: Lemma 2.1 Let CN0 0, CN1 0, # #, and P #P
Nonparametric Estimation of Average Treatment Effects Under Exogeneity: A Review
Recently there has been a surge in econometric work focusing on estimating average treatment effects under various sets of assumptions. One strand of this literature has developed methods for estimating average treatment effects for a binary treatment under assumptions variously described as exogeneity, unconfoundedness, or selection on observables. The implication of these assumptions is that systematic (for example, average or distributional) d…
Large Sample Properties of Matching Estimators for Average Treatment Effects
Matching estimators for average treatment effects are widely used in evaluation research despite the fact that their large sample properties have not been established in many cases. The absence of formal results in this area may be partly due to the fact that standard asymptotic expansions do not apply to matching estimators with a fixed number of matches because such estimators are highly nonsmooth functionals of the data. In this article we dev…
Identification and Inference in Nonlinear Difference-in-Differences Models
This paper develops a generalization of the widely used difference-in-differences method for evaluating the effects of policy changes. We propose a model that allows the control and treatment groups to have different average benefits from the treatment. The assumptions of the proposed model are invariant to the scaling of the outcome. We provide conditions under which the model is nonparametrically identified and propose an estimator that can be …
Regression discontinuity designs: A guide to practice
Nonparametric Tests for Treatment Effect Heterogeneity
In this paper we develop two nonparametric tests of treatment effect heterogeneity. The first test is for the null hypothesis that the treatment has a zero average effect for all subpopulations defined by covariates. The second test is for the null hypothesis that the average effect conditional on the covariates is identical for all subpopulations, that is, that there is no heterogeneity in average treatment effects by covariates. We derive tests…
Dealing with limited overlap in estimation of average treatment effects
Estimation of average treatment effects under unconfounded or ignorable treatment assignment is often hampered by lack of overlap in the covariate distributions between treatment groups. This lack of overlap can lead to imprecise estimates, and can make commonly used estimators sensitive to the choice of specification. In such cases researchers have often used ad hoc methods for trimming the sample. We develop a systematic approach to addressing …
Recent Developments in the Econometrics of Program Evaluation
Many empirical questions in economics and other social sciences depend on causal effects of programs or policies. In the last two decades, much research has been done on the econometric and statistical analysis of such causal effects. This recent theoretical literature has built on, and combined features of, earlier work in both the statistics and econometrics literatures. It has by now reached a level of maturity that makes it an important tool …
Econometrics (39 works) · Mathematics (37 works) · Statistics (36 works) · Advanced Causal Inference Techniques (33 works) · Computer Science (30 works) · Economics (20 works) · Statistical Methods and Inference (18 works) · Estimator (15 works) · Propensity score matching (10 works) · Regression (10 works)