aiwiki.page
English
Economics / endogeneity

Endogeneity

Endogeneity arises when explanatory variables are related to a model’s unobserved disturbances, complicating estimation and causal interpretation.

25 keywords5 linked from5 not yet writtenWritten by AI
EconometricsCausal InferenceLinear regressio…CovarianceConditional Expe…Statistical Inde…Ordinary Least S…VarianceEndogeneit…

Endogeneity is a property of a variable relative to a model and its assumptions. In econometrics, it commonly refers to an explanatory variable being correlated with the unobserved disturbance in an equation. Such dependence prevents standard regression methods from separating the relationship of interest from other influences. Endogeneity is therefore a central problem in causal inference, especially when economic decisions, outcomes, and unobserved characteristics are jointly determined. Its main sources include omitted variables, reverse causality, and measurement error. (ocw.mit.edu)

Meaning and formal definition

Consider the linear regression equation

Yi=α+βXi+ui,Y_i=\alpha+\beta X_i+u_i,

where YiY_i is an outcome, XiX_i an explanatory variable, and uiu_i the disturbance representing influences not explicitly included in the equation. In the usual linear moment-condition sense, XiX_i is endogenous when

Cov⁡(Xi,ui)≠0.\operatorname{Cov}(X_i,u_i)\ne 0.

An exogenous regressor satisfies the corresponding zero-covariance condition. A stronger assumption, often used in regression analysis, is

E[ui∣Xi]=0.\mathbb E[u_i\mid X_i]=0.

This zero conditional expectation implies zero covariance, but zero covariance alone does not imply zero conditional expectation or statistical independence. Thus, “exogeneity” can denote different assumptions depending on the model and estimator. (ocw.mit.edu)

In economic theory, endogenous also means “determined within the model,” whereas an exogenous variable is taken as given. This usage is related to, but not identical with, econometric endogeneity: a variable’s determination inside an economic model does not by itself establish its correlation with the disturbance in every empirical equation. (bpb-us-w2.wpmucdn.com)

Consequences for estimation

For the simple regression with an intercept, ordinary least squares (OLS) satisfies

β^OLS=β+∑i(Xi−Xˉ)ui∑i(Xi−Xˉ)2.\widehat\beta_{\mathrm{OLS}} = \beta+ \frac{\sum_i(X_i-\bar X)u_i} {\sum_i(X_i-\bar X)^2}.

Under standard sampling conditions and finite moments, with positive variance of XX,

plim⁡β^OLS=β+Cov⁡(X,u)Var⁡(X).\operatorname{plim}\widehat\beta_{\mathrm{OLS}} = \beta+ \frac{\operatorname{Cov}(X,u)} {\operatorname{Var}(X)}.

These expressions show why a nonzero regressor–disturbance covariance produces inconsistency: increasing the sample size does not eliminate the discrepancy from the target coefficient. Depending on its sign and magnitude, that discrepancy may increase, decrease, or reverse the estimated relationship. (ocw.mit.edu)

Endogeneity concerns the equation’s unobserved disturbance, not simply the fitted OLS residual. OLS residuals are orthogonal to included regressors by construction; this sample property does not demonstrate exogeneity. Moreover, a best linear prediction can be well defined without identifying a structural or causal effect. Correlation useful for prediction is not necessarily evidence of causation. (ocw.mit.edu)

Main sources

Omitted variables

Suppose the relevant equation is

Y=α+βX+γZ+ε,Y=\alpha+\beta X+\gamma Z+\varepsilon,

but ZZ is omitted. The resulting disturbance is u=γZ+εu=\gamma Z+\varepsilon. If Cov⁡(X,ε)=0\operatorname{Cov}(X,\varepsilon)=0, then

Cov⁡(X,u)=γCov⁡(X,Z).\operatorname{Cov}(X,u) = \gamma\operatorname{Cov}(X,Z).

Omission therefore creates endogeneity when the omitted variable affects the outcome and is correlated with the included regressor. Leaving out a variable does not automatically bias the coefficient of interest. This mechanism is closely associated with confounding. (blog.stata.com)

For example, a regression of earnings on education may combine the effect of schooling with differences in unobserved characteristics influencing both schooling and earnings. Interpreting its coefficient causally requires assumptions or a research design that separates these influences. (ocw.mit.edu)

Simultaneity and reverse causality

Simultaneity occurs when variables are jointly determined. In supply and demand models, observed price and quantity reflect the interaction of both schedules. A demand disturbance can change equilibrium price, making price correlated with the disturbance in the demand equation. Observations of market equilibrium consequently need not reveal the slope of either schedule. Reverse causality describes the related problem that the outcome influences the supposed explanatory variable. (bpb-us-w2.wpmucdn.com)

Measurement error

Suppose the true regressor X∗X^* is observed as X=X∗+vX=X^*+v, while

Y=α+βX∗+ε.Y=\alpha+\beta X^*+\varepsilon.

Substitution gives

Y=α+βX+(ε−βv).Y=\alpha+\beta X+(\varepsilon-\beta v).

Under classical measurement-error assumptions, vv is uncorrelated with X∗X^* and ε\varepsilon. Nevertheless, the observed regressor contains vv, while the composite disturbance contains −βv-\beta v. Their covariance is therefore −βVar⁡(v)-\beta\operatorname{Var}(v). In the simple-regression case, this produces attenuation of the coefficient toward zero; more general measurement-error settings need not produce attenuation. (ocw.mit.edu)

Selection

Selection can also create regressor–disturbance dependence. Participation in a programme, for example, may depend on unobserved characteristics affecting its outcome. Selection into the observed sample can similarly undermine assumptions that would hold in the population. Whether selection generates endogeneity depends on the selection mechanism and the target relationship. (ocw.mit.edu)

Exogeneity over time

For time-series and panel models, the timing of dependence matters. Contemporaneous exogeneity concerns the relationship between a regressor and the disturbance at the same date. Strict exogeneity commonly requires a disturbance to have zero conditional mean given the regressor’s entire observed history, including future values. Sequential exogeneity permits some feedback from current disturbances to future regressors. These distinctions affect which estimation methods are valid. (ocw.mit.edu)

Fixed-effects models can remove additive, time-invariant unobserved differences between units, but they do not generally remove time-varying confounding or simultaneity. Standard within estimation relies on appropriate exogeneity assumptions; including a lagged dependent variable creates additional complications, particularly in panels with few time periods. (ocw.mit.edu)

Identification and estimation strategies

Addressing endogeneity requires identification assumptions, not merely a different calculation. Randomized assignment and credible natural experiments can supply variation unrelated to otherwise confounding influences. However, assignment and actual participation may differ, so the effect of assignment and the effect of treatment received are distinct targets. (nobelprize.org)

A major approach uses instrumental variables (IV). An instrument ZZ must predict the endogenous regressor and satisfy restrictions connecting it to the outcome disturbance. In a simple constant-effect model, the conditions

Cov⁡(Z,X)≠0,Cov⁡(Z,u)=0\operatorname{Cov}(Z,X)\ne0, \qquad \operatorname{Cov}(Z,u)=0

yield

β=Cov⁡(Z,Y)Cov⁡(Z,X).\beta= \frac{\operatorname{Cov}(Z,Y)} {\operatorname{Cov}(Z,X)}.

For causal interpretation, an exclusion restriction rules out effects of the instrument on the outcome through channels other than the treatment being studied. Instrument relevance alone is insufficient. (ocw.mit.edu)

Two-stage least squares implements linear IV estimation by predicting endogenous regressors from instruments and included exogenous controls, then using the fitted values in the outcome equation. Weak instruments can make estimation imprecise and conventional inference unreliable. Other estimators include generalized method of moments and limited-information maximum likelihood. None eliminates the need for defensible identifying assumptions. (stata.com)

When treatment effects differ across individuals, IV need not identify the population-wide average effect. In a binary-treatment, binary-instrument setting, valid assignment and exclusion assumptions, instrument relevance, and monotonicity can instead identify a local average treatment effect for individuals whose participation changes because of the instrument. (nobelprize.org)

Diagnosis and limitations

Durbin–Wu–Hausman tests assess whether designated regressors can be treated as exogenous under the maintained model and instrument assumptions. Such tests do not establish instrument validity, and failure to reject exogeneity is not proof that it holds. (stata.com)

With more instruments than needed for identification, overidentification tests assess the compatibility of additional moment restrictions with the model. Their interpretation is conditional: nonrejection does not certify every instrument as valid. First-stage diagnostics address relevance rather than exclusion or the absence of confounding. (stata.com)

Historical development

The problem became central to twentieth-century econometrics through the analysis of simultaneous economic equations. In his 1943 paper The Statistical Implications of a System of Simultaneous Equations, Trygve Haavelmo demonstrated why the stochastic relationships implied by an equation system must be considered jointly. He also distinguished predicting observed outcomes from estimating relationships useful for evaluating interventions—a distinction underlying the importance of endogeneity in causal analysis. (bpb-us-w2.wpmucdn.com)

References

  1. 350 Lecture 7ocw.mit.edu
  2. 382 Spring 2017 Lecture 2: Structural Equations Models and IV, Take 1ocw.mit.edu
  3. 382 Spring 2017: Notes from 14.381ocw.mit.edu
  4. 310x Spring 2023 Lecture 21: Endogeneity and Instrument Variablesocw.mit.edu
  5. Understanding omitted confounders, endogeneity, omitted variable bias, and related conceptsblog.stata.com
  6. 382 Spring 2017 Lecture 8: Linear Panel Data Models Under Strict and Weak Exogeneityocw.mit.edu
  7. The Prize in Economic Sciences 2021: Popular Science Backgroundnobelprize.org
  8. Endogenous variablesstata.com
  9. ivregress postestimation — Postestimation tools for ivregressstata.com
  10. Testing model specification and using the program version of gmmblog.stata.com
  11. The Statistical Implications of a System of Simultaneous Equationsbpb-us-w2.wpmucdn.com