Cluster Validation and Stability
Dr. Nadia, an entomologist, measured 150 beetles. For each one she wrote down two numbers with calipers: its body length and its antenna length, both in millimetres. She ran k-means, asked for three groups, and got three tidy clusters back. She is about to write "our beetles fall into three species" in her paper.
But here is the uncomfortable question this lesson answers: k-means would have handed her three tidy groups even if the beetles had no species at all. Finding clusters is easy; the algorithm always finds them. Knowing whether they are real is the hard part, and it is a separate job.
By the end of this lesson you will be able to:
- Explain why a tidy clustering is not evidence that the groups are real
- Score cohesion and separation with the silhouette, and use it to pick the number of clusters
- Test for structure at all with the gap statistic, which can vote "there are no clusters"
- Measure stability by resampling, and combine all three into an honest verdict
Prerequisites: the earlier lessons in this course, where you learned to run a clustering (k-means in Lesson 3, hierarchical and density in Lesson 4, Gaussian mixtures in Lesson 5). You should be comfortable running R and reading its output, and know what a mean, a standard deviation, a Euclidean distance, and a scatter plot are. No linear algebra is assumed; every symbol is defined as it appears.
The panel above previews the two curves you will learn to read: an elbow that bends at the right number of clusters, and silhouette bars that peak there. Move the slider to get a feel for them, then let us build the real thing on Nadia's beetles.
Every method hands you clusters
Over the last three lessons you clustered data three different ways: k-means split it into round groups (Lesson 3), hierarchical and density methods found groups of other shapes (Lesson 4), and Gaussian mixtures gave soft, probabilistic memberships (Lesson 5). They disagree on the how, but they share one habit: every one of them returns clusters, whatever you feed it. Ask for three groups and you get three groups. The algorithm never says "actually, there are no groups here."
So the real question, and the whole subject of this lesson, is: once you have a clustering, how do you know it reflects real structure in the data rather than lines the algorithm drew through a shapeless cloud?
Let us meet Nadia's data. We simulate the beetles as three species so that, at the very end, we can check the verdict against the truth. Real data never comes with an answer key, which is exactly why the validation tools in this lesson exist.
Because there are only two measurements, we can simply plot every beetle and look. How many groups do you see?