Decision Trees for Classification
Lesson 3 drew smooth, curved boundaries by modelling each class as a Gaussian cloud. A decision tree throws all of that out. It asks a short series of plain yes/no questions, "is the petal shorter than 2.5 cm?", and carves the feature space into rectangular boxes, one label per box. It is the most readable classifier there is: the whole model is a flowchart you can follow by hand.
By the end of this lesson you will be able to:
- Explain how a tree splits data into axis-aligned rectangles with a sequence of yes/no questions, and read a fitted tree
- Define node impurity (Gini and entropy) and how a tree picks the split that lowers it most
- Grow, read and prune a classification tree in R, and explain why an unpruned tree overfits
Prerequisites: you can run R and read its output, you know what a training set and a classifier are, and you have met the earlier classifiers in this course (kNN, Naive Bayes, LDA/QDA).
A tree is a flowchart of questions
Picture identifying a wild iris with a ruler. You measure the flower and walk down a field guide: "Is the petal shorter than 2.5 cm? If yes, it is a setosa. If no, is the petal narrower than 1.75 cm? If yes, versicolor; if no, virginica." That field guide IS a decision tree. Each question is a split, each endpoint is a leaf, and the label a leaf predicts is simply the majority class of the training flowers that land there.
We will use a real, famous dataset: 150 irises measured by the botanist Edgar Anderson, 50 each of three species, with four measurements per flower in centimetres. It is built into R, so a fresh session already has it.
The tree below is the flowchart we will actually grow from this data in a moment. The whole job is choosing the right question at each split, which is the next step.