Lesson 4 of 8

Multicollinearity in Regression

In Lesson 3 you watched a single row hijack a regression: one freak heatwave day halved Priya's slope. This lesson is a quieter sabotage from a different direction, two columns that secretly say the same thing.

Priya, still running her iced-coffee cart, wants a sharper model. Her weather app shows two numbers every morning: the actual temperature and the "feels like" temperature. "Feels like is what customers really experience," she reasons, "so let me add it." She drops both into her regression, and the slope on temperature, a rock-solid +1.9 cups per degree since Lesson 1, suddenly turns negative. The model now claims hotter days sell fewer cups. Nothing about her data changed. She just added one well-meaning column.

By the end of this lesson you will be able to:

  • Explain what multicollinearity is, and why it makes coefficients unstable while leaving predictions intact
  • Detect it with a correlation matrix, and know the one case that matrix misses
  • Compute the variance inflation factor (VIF) in R and read it against the usual cutoffs
  • Decide what to do about it, including when the honest answer is "nothing"

Prerequisites: Lessons 1 to 3 (you can fit a line with lm(), read its coefficients and standard errors, and you know what R-squared means). Every new term is defined as it appears.

The scatter below is the whole problem in one picture: each dot is one day, actual temperature on the bottom, "feels like" up the side. They fall almost perfectly on a line. Two columns carrying one piece of information.

The setup

A second thermometer

Priya logged four weeks of trading: each day's high temperature in Celsius, the "feels like" reading off her weather app, and cups sold. A fresh R session starts empty, so we build that table right here.

RInteractive R
set.seed(22) n <- 28 temp <- round(runif(n, 16, 34), 1) # the day's high, in Celsius feels_like <- round(temp + 3 + rnorm(n, 0, 0.9), 1) # the weather-app "feels like" cups <- round(8 + 1.9 * temp + rnorm(n, 0, 4)) # sales, driven by the REAL temperature coffee <- data.frame(temp, feels_like, cups) head(coffee) #> temp feels_like cups #> 1 21.5 23.8 52 #> 2 24.5 26.7 62 #> 3 33.9 37.7 76 #> 4 25.4 30.2 69 #> 5 31.2 35.0 67 #> 6 29.0 30.5 64

  

Look at how we built it, because it is the whole reason the trouble appears. Cups depend on the real temperature. feels_like is just that same temperature plus about three degrees of "mugginess" and a sprinkle of noise. It carries no information of its own about sales; it is a near-copy of temp wearing a different label.

First, the model Priya trusted from Lesson 1, temperature alone:

RInteractive R
round(summary(lm(cups ~ temp, data = coffee))$coef, 3) #> Estimate Std. Error t value Pr(>|t|) #> (Intercept) 9.388 5.232 1.794 0.084 #> temp 1.902 0.196 9.710 0.000

  

Each extra degree buys her about 1.9 more cups. The small standard error (0.196) and the large t-value (9.7) say that estimate is rock solid: this is exactly the relationship she has counted on for three lessons. This is the baseline we are about to break.