Lesson 3 of 8

k-Means and Choosing k

Maria runs a coffee shop. For each of her 200 loyalty-card customers she has two numbers from the past year: how many times they visited, and how much they spent per visit on average. She is sure there are a few natural kinds of customer hiding in there, but nobody labeled them. k-means is the algorithm that finds those groups, and this lesson is about how it works and how to decide how many groups to look for.

By the end you will be able to:

  • Explain what k-means does and the quantity it is trying to make small
  • Step through Lloyd's algorithm, the two moves that actually do the clustering
  • Run k-means in R, standardize your features first, and choose the number of clusters with the elbow and the silhouette

Prerequisites: you can run R, and you have met the idea (from PCA and Factor Analysis) that unsupervised methods find structure with no labels to guide them.

Where we are

Finding groups nobody labeled

The last two lessons worked on the columns of your data: PCA and factor analysis compressed many correlated variables into a few underlying ones. Clustering turns the other way and works on the rows. Given a table of customers, which ones belong together?

That is exactly Maria's question. Let us build her 200-customer table so we have something concrete to cluster. We plant three real segments (occasional visitors, regulars, and big spenders) only so we can check later whether k-means rediscovers them; the algorithm itself never sees these labels.

RInteractive R
set.seed(1) # 200 loyalty-card customers. Two numbers each: # visits = store visits per year, spend = average spend per visit (dollars) occasional <- data.frame(visits = rnorm(70, 45, 10), spend = rnorm(70, 4.6, 0.7)) regulars <- data.frame(visits = rnorm(70, 220, 22), spend = rnorm(70, 5.4, 0.9)) big_spender <- data.frame(visits = rnorm(60, 70, 14), spend = rnorm(60, 13.0, 1.4)) coffee <- rbind(occasional, regulars, big_spender) coffee$visits <- round(pmax(6, coffee$visits)) # whole visits, floor at 6 coffee$spend <- round(pmax(1, coffee$spend), 1) # dollars, one decimal truth <- rep(c("occasional", "regular", "big_spender"), c(70, 70, 60)) head(coffee, 4) #> visits spend #> 1 39 4.9 #> 2 47 4.1 #> 3 37 5.0 #> 4 61 3.9

  

Each row is a dot in a two-dimensional space: visits along one axis, spend along the other. Clustering means finding the clumps of dots.