k-Means and Choosing k
Maria runs a coffee shop. For each of her 200 loyalty-card customers she has two numbers from the past year: how many times they visited, and how much they spent per visit on average. She is sure there are a few natural kinds of customer hiding in there, but nobody labeled them. k-means is the algorithm that finds those groups, and this lesson is about how it works and how to decide how many groups to look for.
By the end you will be able to:
- Explain what k-means does and the quantity it is trying to make small
- Step through Lloyd's algorithm, the two moves that actually do the clustering
- Run k-means in R, standardize your features first, and choose the number of clusters with the elbow and the silhouette
Prerequisites: you can run R, and you have met the idea (from PCA and Factor Analysis) that unsupervised methods find structure with no labels to guide them.
Finding groups nobody labeled
The last two lessons worked on the columns of your data: PCA and factor analysis compressed many correlated variables into a few underlying ones. Clustering turns the other way and works on the rows. Given a table of customers, which ones belong together?
That is exactly Maria's question. Let us build her 200-customer table so we have something concrete to cluster. We plant three real segments (occasional visitors, regulars, and big spenders) only so we can check later whether k-means rediscovers them; the algorithm itself never sees these labels.
Each row is a dot in a two-dimensional space: visits along one axis, spend along the other. Clustering means finding the clumps of dots.