Lesson 4 of 6

Decision Trees for Classification

Lesson 3 drew smooth, curved boundaries by modelling each class as a Gaussian cloud. A decision tree throws all of that out. It asks a short series of plain yes/no questions, "is the petal shorter than 2.5 cm?", and carves the feature space into rectangular boxes, one label per box. It is the most readable classifier there is: the whole model is a flowchart you can follow by hand.

By the end of this lesson you will be able to:

  • Explain how a tree splits data into axis-aligned rectangles with a sequence of yes/no questions, and read a fitted tree
  • Define node impurity (Gini and entropy) and how a tree picks the split that lowers it most
  • Grow, read and prune a classification tree in R, and explain why an unpruned tree overfits

Prerequisites: you can run R and read its output, you know what a training set and a classifier are, and you have met the earlier classifiers in this course (kNN, Naive Bayes, LDA/QDA).

The idea

A tree is a flowchart of questions

Picture identifying a wild iris with a ruler. You measure the flower and walk down a field guide: "Is the petal shorter than 2.5 cm? If yes, it is a setosa. If no, is the petal narrower than 1.75 cm? If yes, versicolor; if no, virginica." That field guide IS a decision tree. Each question is a split, each endpoint is a leaf, and the label a leaf predicts is simply the majority class of the training flowers that land there.

We will use a real, famous dataset: 150 irises measured by the botanist Edgar Anderson, 50 each of three species, with four measurements per flower in centimetres. It is built into R, so a fresh session already has it.

RInteractive R
# 150 real irises: 3 species, 4 measurements (cm). Built into R. data(iris) head(iris, 4) #> Sepal.Length Sepal.Width Petal.Length Petal.Width Species #> 1 5.1 3.5 1.4 0.2 setosa #> 2 4.9 3.0 1.4 0.2 setosa #> 3 4.7 3.2 1.3 0.2 setosa #> 4 4.6 3.1 1.5 0.2 setosa table(iris$Species) #> #> setosa versicolor virginica #> 50 50 50

  

The tree below is the flowchart we will actually grow from this data in a moment. The whole job is choosing the right question at each split, which is the next step.