Omitted variable bias
1. The Data Generating Process
Simulations are incredibly powerful tools to understand estimators, such as ordinary least squares (OLS).
In a simulation, we first pretend we control the world and know exactly how the data is produced; this is usually called a data generating process (DGP). We then pretend that we do not know this DGP, and instead just see the data it produces. Our task is then to see how well our estimator reflects the real effect of the DGP. If our estimator fails to work under these ‘ideal’ settings, then it probably won’t work on real data.
We can also test how robust our estimator is to misspecification, which is what we will do here.
Suppose the following relationship described the real world: a student ’s test score is completely determined by the hours they study, the hours they sleep, and some normally distributed noise:
DGP details: how the data are drawn
For each student, we draw two independent uniform variables and an independent noise term. Draws are independent across students:
Download the R DGP file to experiment. Edit the settings at the top, then run it in R or RStudio to generate data, fit both regressions, and plot the results. Extension: how robust is a regression estimator to other distributions of noise?
We prespecify how correlated sleep and study hours are. Play around with the following two sliders and see how this changes the following plot. What do you notice?
The green plane shows the true relationship before adding noise. Each dot is one student’s observed score, including noise.
2. Hitting the Jackpot: Guessing the Correct DGP
Now pretend that we are economists who are either very smart or very lucky and propose the following regression model:
We use ordinary least squares (OLS) to estimate the model’s coefficients using our sample of 500 students:
Regression table
This is a standard regression table. Try to put the coefficients, standard error, and in sentences. Are standard errors always positive?
| (1) | |
|---|---|
| Dependent variable: | Test score |
Notes: Standard errors in parentheses. OLS estimates include a constant; standard errors assume independent errors with constant variance.
Study and sleep are hours per day. Test scores are points.
Let’s look at one specific person.
Ask yourself, does the exogeneity condition hold in this regression?
3. A model that ignores sleep
Let’s now pretend we are economists who wake up without any insight into how the real world works. We think that test scores depend only on the number of hours studied. We propose the following model:
On the same sample of students, we estimate this model by OLS:
Regression table: both specifications
This regression table contains the results from the regression in the previous section in column (1) and this section in column (2).
| (1) | (2) | |
|---|---|---|
| Dependent variable: | Test score | Test score |
| With sleep | Without sleep |
Notes: Standard errors in parentheses. Both specifications use the same students and include a constant. Standard errors assume independent errors with constant variance.
The brown line is the regression estimated using study alone. The dashed green line shows the true study effect of 5, holding sleep at its sample average. Omitting sleep can make the estimated slope differ from this true effect.
4. Does the exogeneity condition hold in the misspecified regression?
Since we know the DGP, we can check whether the study-only regression recovers the causal effect of study. Keep the true study coefficient of 5 and collect sleep and noise in a composite error, :
Now check the exogeneity condition for this causal effect, :
5. Where does the bias come from?
In class, we spent some time on intuition of the direction of bias. Here, I will provide a slightly more programmatic way to figure that out with two simple questions. As a review, remember we are trying to understand the direction of bias on the study hours coefficient in the short regression (Section 3). Recall that
The direction of bias comes from two parts: the effect of sleep on test scores () and the sign of the correlation between sleep hours and study hours. These two parts let us quickly understand which direction bias will go in:
Does sleep cause better test scores (apart from study hours’ own effect)?
Are sleep hours and study hours positively correlated?
Exercises
For questions 1 and 2, use this DGP and let be study hours and be test score.
-
Does omitting a variable that directly affects but is uncorrelated with violate unconfoundedness?
Does omitting a variable that is correlated with but has no direct effect on , holding fixed, violate unconfoundedness?
Can the OLS slope estimator be unbiased for the causal study effect even when conditional exogeneity fails?