Decision Boundaries and Model Geometry
Lesson 4 ended with a decision tree carving the plane into rectangular boxes, one label per box. Lesson 3 drew smooth Gaussian curves; Lesson 1 followed a jagged trail of nearest neighbours. Five different machines, five different pictures. This lesson reveals that they are all the same idea wearing different clothes.
Here is the running example for the whole lesson. A sleep-tracking watch labels each night restful or restless from just two numbers: your average overnight heart rate in beats per minute, and your movements per hour. Once the watch has trained, it has an opinion about every possible night, so it has effectively painted the entire heart-rate-by-movements plane with two colours. The border between the colours is the decision boundary, and its shape, straight or curved or wiggly or boxy, is a fingerprint of the model that drew it.
By the end of this lesson you will be able to:
- Read any classifier as a decision boundary that splits feature space into one region per class, and tell a linear boundary from a nonlinear one
- Explain generative versus discriminative, a second axis that is independent of the boundary's shape
- Predict which of kNN, logistic regression, LDA, QDA and a tree draws a straight, curved, wiggly or boxy boundary, and why
- Fit all of them on one dataset and draw and compare their boundaries in R
Prerequisites: you can run R and read its output, and you have met this course's classifiers (kNN in Lesson 1, Naive Bayes in Lesson 2, LDA and QDA in Lesson 3, the decision tree in Lesson 4) and the idea of bias versus variance.
The picture above is one such boundary, a tree's, with a dial for how flexible it is. Slide it later; for now, just notice that the model has an answer everywhere, and a visible fence between the two answers. That fence is what this whole lesson is about.
Every classifier is a boundary
Think about what your sleep watch actually does. You hand it a night, a single point on the plane like "heart rate 57, movements 15", and it returns one word. Do that for every point on the plane and the whole surface fills in with two colours: a restful region and a restless region. The decision boundary is simply the line where those two regions meet, the set of nights the model finds a perfect toss-up.
We can say that precisely. Write \(x\) for a night's two numbers, \(x = (\text{heart rate}, \text{movements})\), and write \(P(\text{restless} \mid x)\) for the probability the model assigns to "restless" given that night. The boundary is exactly the set of nights where the two verdicts are tied at fifty-fifty:
\[ \{\, x \;:\; P(\text{restless} \mid x) = P(\text{restful} \mid x) = 0.5 \,\} \]
Everything on one side is painted restless, everything on the other restful. So a classifier is its boundary: to know the model is to know the shape of that fence. Different models draw the fence differently, and that is the story of this lesson.