Inference and Prediction in Regression
In Lesson 5 you learned to make a standard error honest even when the errors misbehave. Now you get to spend it. A trustworthy standard error is the raw material for the two questions every regression is really asked: is this relationship real, and what will happen next.
We come back to Priya's original iced-coffee cart from Lesson 1, the clean 12-day log where the line's assumptions hold, so the numbers lm() prints are trustworthy as they stand. On that honest fit you will answer both of Priya's questions with numbers, not hand-waving.
By the end of this lesson you will be able to:
- Wrap a coefficient in a confidence interval and run the t-test that asks whether it is really non-zero
- Predict a brand-new day two different ways: a confidence interval for the average, and a prediction interval for one specific day
- Explain why a prediction interval is always wider, and why explaining a relationship and predicting a value are two different jobs
Prerequisites: Lessons 1 to 5. You can fit a line with lm() and read its coefficients, standard errors and p-values. Every new term is defined as it appears.
Two jobs, two questions
A fitted line quietly does two jobs at once, and confusing them is the single most common regression mistake. Priya has one model, lm(cups ~ temp), but two very different things she wants from it:
| Priya asks... | The job | The tool |
|---|---|---|
| Is warmer weather really driving my sales, and how sure can I be? | Explain the relationship | a test and a confidence interval on the slope |
| How many cups will tomorrow's forecast 25-degree day bring? | Predict a new value | a prediction interval for one new day |
The first question is about the slope: a number describing the whole relationship. The second is about a single future day: one dot that has not happened yet. They use different intervals, and mixing them up leads either to false confidence or to pointless caution.