Lesson 1 of 8

OLS Regression from Scratch

Priya runs a small iced-coffee cart outside a train station. On warm days she sells out; on cool days she carries stock home. For twelve days she wrote down two numbers: the day's high temperature and how many iced coffees she sold. She wants one honest rule, a straight line, that turns tomorrow's forecast into a guess for the day's sales.

There are infinitely many lines she could draw through her data. This lesson is about the single method that picks the best one: ordinary least squares, or OLS. By the end you will be able to:

  • Say exactly what makes one line "best", and define a residual
  • Compute that line two ways by hand (a formula, then the matrix normal equations) and confirm R's lm() gets the identical answer
  • Read the line's slope and intercept, and judge the fit with R-squared

Prerequisites: you can run an R code block, read a scatterplot, and know what an average is. No calculus or matrix algebra assumed; we build those pieces up as we reach them.

The widget below is the whole lesson in miniature, running on Priya's real twelve days. Drag the slope and intercept: every point drops a red square onto the line, and the square's area is that day's error. Your goal, and OLS's goal, is to make the total red area as small as it can be. Press "Snap to least squares" to see where the math lands.

The rule

A line is a prediction machine

Here are Priya's twelve days. Each dot is one day: its temperature along the bottom, the cups she sold up the side. The dots climb from lower-left to upper-right, so warmer really does seem to mean more cups.

A straight line through that cloud is a rule for turning any temperature into a predicted number of cups. We write it

\[ \hat{y} = b_0 + b_1 x \]

Reading it in Priya's words: \(x\) is the day's temperature (the predictor), \(\hat{y}\) (said "y-hat") is the line's predicted cups, \(b_0\) is the intercept (the predicted cups at a temperature of zero, where the line crosses the vertical axis), and \(b_1\) is the slope (how many extra cups the line adds for each one-degree rise). The hat on \(\hat{y}\) matters: it is the line's guess, not the actual cups Priya sold.

Every choice of \(b_0\) and \(b_1\) is a different line, a different rule. Our whole job is to choose those two numbers well. First, let us keep Priya's data in R so we can work with it for the rest of the lesson. Each lesson starts a fresh R session, so we build the twelve days right here (run this first).

RInteractive R
# Priya's cart: the day's high temperature (deg C) and iced coffees sold. coffee <- data.frame( temp = c(15, 17, 18, 20, 21, 23, 24, 26, 27, 29, 30, 31), cups = c(30, 36, 33, 42, 40, 47, 44, 52, 55, 56, 61, 60) ) head(coffee) #> temp cups #> 1 15 30 #> 2 17 36 #> 3 18 33 #> 4 20 42 #> 5 21 40 #> 6 23 47