Lesson 7 of 8

t-SNE and UMAP

Meet Priya, a data analyst at a music-streaming service. Her library holds thousands of songs, and for each one the audio engine has measured eight numbers: how energetic it is, its tempo, how danceable it is, how acoustic, how loud, how upbeat (its valence), how instrumental, and how full of spoken words. Eight numbers per song is far too many to eyeball. Priya wishes she could drop every song as a single dot on a flat map, where songs that SOUND alike sit close together, so the genres reveal themselves.

In Lesson 6 you learned to check whether clusters are real. Now you will learn to SEE them. The map below is the kind PCA drew back in Lesson 1 (here, of flowers): a flat 2D picture where similar items sit close. It is a fine start, but PCA can only bend in straight lines, and that turns out to be a real limit. This lesson builds sharper, nonlinear maps, t-SNE and UMAP, and, just as importantly, teaches you the traps in reading them.

By the end of this lesson you will be able to:

  • Explain why a nonlinear map can reveal cluster structure that a linear PCA map hides
  • Say, in plain words, how t-SNE turns distances into a 2D map, and what perplexity controls
  • Run t-SNE in R, meet UMAP, and choose between them
  • Read one of these maps without being fooled by its most seductive traps

Prerequisites: Lesson 1 (PCA in R) for the idea of a 2D map and variance explained, plus the clustering lessons before this. You should be comfortable running R and reading its output, and know what a mean, a standard deviation, a Euclidean distance, and a scatter plot are. No new linear algebra is assumed; every symbol is defined as it appears.

The flat map above is PCA, the linear kind. By the end of this lesson you will build maps that pull apart structure PCA leaves in a smear.

The problem

Eight numbers per song

Let us give Priya some data to work with. Each lesson runs in a fresh R session, so we build her song library right here (run this once). We invent six genres and, for each, a characteristic "sound": a couple of audio features it scores high on. Then we scatter individual songs around that signature.

RInteractive R
set.seed(1) genres <- c("lo-fi", "classical", "metal", "EDM", "folk", "hip-hop") n_per <- c(90, 45, 45, 30, 60, 30) # how many songs in each genre spread <- c(0.35, 0.50, 0.45, 0.65, 0.45, 0.50) # how tightly each genre clusters names(n_per) <- names(spread) <- genres feat <- c("energy", "tempo", "dance", "acoustic", "loud", "valence", "instr", "speech") # Each genre is "high" in a couple of characteristic audio features (its sound). signature <- rbind( "lo-fi" = c(0, -1.3, 0, 0.6, 0, 0, 1.6, 0), "classical" = c(0, 0.0, 0, 1.7, 0, 0, 1.4, 0), "metal" = c(1.8, 0.0, 0, 0.0, 1.7, 0, 0, 0), "EDM" = c(1.3, 1.6, 1.5, 0, 0, 0, 0, 0), "folk" = c(0, 0.0, 0, 1.5, 0, 1.6, 0, 0), "hip-hop" = c(0, 0.0, 1.4, 0, 0, 0, 0, 1.9) ) colnames(signature) <- feat # scatter each song around its genre's signature songs <- do.call(rbind, lapply(genres, function(g) matrix(rnorm(n_per[g] * 8, 0, spread[g]), n_per[g], 8) + matrix(signature[g, ], n_per[g], 8, byrow = TRUE))) colnames(songs) <- feat genre <- factor(rep(genres, times = n_per), levels = genres) dim(songs) #> [1] 300 8 table(genre) #> genre #> lo-fi classical metal EDM folk hip-hop #> 90 45 45 30 60 30

  

We have 300 songs, each a row of eight numbers, and we happen to know the true genre of every one (that is our ground truth, handy for judging the maps later). The averages confirm each genre really does have its own sound:

RInteractive R
# the average of each feature within each genre round(t(sapply(genres, function(g) colMeans(songs[genre == g, ]))), 2) #> energy tempo dance acoustic loud valence instr speech #> lo-fi 0.04 -1.30 -0.02 0.64 -0.03 0.00 1.60 -0.08 #> classical -0.03 -0.01 0.04 1.71 0.02 0.02 1.29 0.02 #> metal 1.81 -0.16 0.05 0.07 1.88 -0.09 -0.10 0.16 #> EDM 1.20 1.54 1.54 0.07 -0.04 -0.06 -0.02 -0.14 #> folk -0.01 -0.08 0.14 1.39 0.00 1.58 -0.08 -0.01 #> hip-hop 0.02 -0.02 1.46 -0.05 0.06 0.04 0.11 1.93

  

metal is loud and high-energy; classical is acoustic and instrumental; hip-hop is danceable and full of words. Each genre is a cluster somewhere in this eight-dimensional space. The catch: you cannot plot eight dimensions. Priya can view any two features at a time, but the genres overlap in most pairs, and there are 28 pairs to check. She wants ONE picture.