Gaussian Mixture Models
In Lesson 4, hierarchical clustering and DBSCAN gave Maria clean neighbourhoods on her coffee-shop map. But every method so far, k-means included, has handed back a hard answer: each customer belongs to exactly one group, full stop.
Meet Dan. He comes into Maria's shop about 17 times a month, right in the gap between her regulars (who average around 13 visits) and her devotees (around 20). k-means had to shove Dan fully into one tier, even though he honestly sits on the fence. A Gaussian mixture refuses to pretend. It can say Dan is, say, 60% regular and 40% devotee, and that fraction is a real, computed probability, not a hand-wave. The panel below is that idea in miniature: flip it to soft and watch the fence-sitters take an in-between colour.
By the end of this lesson you will be able to:
- Read a soft assignment: what it means to say a customer is 60% one group and 40% another
- Write down the model behind it (data as a weighted sum of bell curves) and compute a membership probability yourself in R
- Run the EM algorithm that fits a mixture, know why it never gets worse but can get stuck, and fit one for real with
mclust
Prerequisites: you can run R and read its output, and you have done Lesson 3 on k-means (clustering, distance, scaling, and why you run several random starts) and Lesson 4 on hierarchical and density clustering. A bell curve just means the normal distribution; every other symbol is defined as it appears.
Hard labels throw away what you know
Toggle the panel between hard and soft. In hard mode every dot is painted fully blue or fully gold, exactly what k-means does: each point is assigned to its single nearest group. In soft mode the dots in the overlap turn an in-between colour, because the model reports a probability of belonging to each group instead of forcing a choice.
That probability has a name. For each customer and each group, the model gives a responsibility: a number between 0 and 1 saying how much that group "owns" the customer, with the responsibilities across all groups adding up to 1. A responsibility of 0.95 means "almost certainly this tier"; 0.55 versus 0.45 means "leaning one way, but genuinely close."
In fact, hard assignment is a mixture model with the confidence deleted: k-means is what you get if you force every responsibility to be exactly 0 or 1. So a Gaussian mixture is not a rival to k-means so much as the honest, probabilistic version of the same idea.