Heteroskedasticity and Autocorrelation
In Lesson 4, Priya's regression was sabotaged by two columns that said the same thing. This lesson is a subtler kind of trouble: the columns are fine, the slope is fine, and yet the number printed next to it is a lie.
When Priya fits lm(cups ~ temp), R hands her a slope and, right beside it, a standard error: the "give or take" that turns into her p-value and confidence interval. That standard error is computed under two quiet promises about the errors, that they all have the same spread and that they do not lean on each other. Break either promise and the slope stays correct, but the standard error, and every p-value and interval built on it, quietly goes wrong. Usually it comes out too small, so you feel far more certain than you have any right to.
This lesson teaches you to catch both broken promises and to repair the standard error without touching the (perfectly good) slope. The fanning plot below is the first villain: watch the spread grow as you move right.
By the end of this lesson you will be able to:
- Explain which part of a regression heteroskedasticity and autocorrelation corrupt, and which part they leave alone
- Test for non-constant variance with the Breusch-Pagan test and fix it with robust (White) standard errors
- Test for correlated errors with the Durbin-Watson statistic and fix it with Newey-West (HAC) standard errors
- Choose the right correction for a given dataset
Prerequisites: Lessons 1 to 4 (you can fit a line with lm(), and read its coefficients, standard errors and p-values). Lesson 2 introduced the residual funnel and runs-in-time as shapes; here we make them precise, test them, and correct them. Every new term is defined as it appears.
What a standard error is quietly assuming
Priya has expanded her cart and now has 60 trading days of records. A fresh R session starts empty, so we build that table right here.
Now fit the line and read the full table, not just the slope:
The slope is 1.931: each extra degree buys about two more cups, exactly the relationship Priya has trusted all course. The column that matters this lesson is Std. Error, the 0.291 next to the slope. It is the estimated give-or-take on that 1.931, and everything downstream, the t value of 6.6 and the tiny p-value, is just the slope divided by its standard error. Formally, for the whole vector of coefficients \(\hat\beta\),
\[ \operatorname{Var}(\hat\beta) = \sigma^2 (X^\top X)^{-1} \]
where \(X\) is the design matrix (a column of 1s and the temp column), \(X^\top X\) is its cross-product, and \(\sigma^2\) is one single number: the variance (the typical squared size) of the error on every day. That one-number-for-all-days assumption is the promise. When Priya's misses are the same size on cool days and hot days, a healthy flat band of residuals like the one below, the formula is right and her standard error of 0.291 is honest.