Lesson 2 of 6

Naive Bayes for Tabular and Text

Lesson 1 classified by distance: to label a new point, look at who it sits next to. That instinct drowned once we piled on features, because in high dimensions everything is roughly equidistant. This lesson takes the opposite route. Instead of measuring distance, it reasons with probability, and it shrugs off the very high-dimensional data that sank kNN.

Picture the spam filter guarding your inbox. A new email arrives: "free money, claim now". The filter has read thousands of past emails, each already stamped spam or real (in the jargon, real mail is called ham). From those labels alone it computes: given these exact words, how likely is spam versus ham, and files the message under whichever wins. That filter is Naive Bayes, and by the end of this lesson you will have built one from scratch, twice: once on the words of an email, once on numeric features in a table.

By the end you will be able to:

  • Use Bayes' rule to turn a base rate and some evidence into the probability that an email is spam
  • Apply the naive independence assumption, add Laplace smoothing, and classify in log space, all in plain R
  • Fit Gaussian Naive Bayes on numeric features, and explain why an assumption this crude still works so well

Prerequisites: you can run R and read its output, you know what a training set and a classifier are (the ML Workflow course), and Lesson 1 (kNN). Basic probability is enough; conditional probability is defined the moment it appears.

Those three moves are the whole method. The rest of this lesson is where each number comes from.

The core idea

Bayes' rule flips the question

The question we want answered is awkward: given these words, what is the probability the email is spam? We have no direct way to know that. But we can easily measure the reverse from labeled history: among known spam, how often do these words show up? Bayes' rule is the bridge that flips one into the other.

First, one piece of notation. \(P(A \mid B)\) is read "the probability of \(A\) given \(B\)", meaning the probability that \(A\) is true once we already know \(B\) is true. So \(P(\text{spam} \mid \text{"free"})\) is the chance an email is spam given that it contains the word "free". Bayes' rule says:

\[ P(y = k \mid x) = \frac{P(y = k)\,P(x \mid y = k)}{P(x)} \]

Read it left to right. \(y\) is the label and \(k\) is one class (spam or ham); \(x\) is the evidence (the words). The three named pieces:

  • \(P(y = k)\) is the prior: how common class \(k\) is before we look at the email at all (the base rate).
  • \(P(x \mid y = k)\) is the likelihood: how typical this evidence is for class \(k\).
  • \(P(y = k \mid x)\) is the posterior: what we actually want, the probability of the class after seeing the evidence.

Let us make it concrete with one word. Suppose that of the last 100 emails, 40 were spam, the word "free" appeared in 30 of those 40 spam emails, and in just 6 of the 60 real ones. What is the probability an email is spam given it says "free"?

RInteractive R
# Priors: 40 of the last 100 emails were spam. p_spam <- 0.40 p_ham <- 0.60 # Likelihoods: how often "free" shows up in each class. p_free_spam <- 30 / 40 # 0.75 p_free_ham <- 6 / 60 # 0.10 # Bayes' rule. The numerator for each class is prior * likelihood; # the denominator P("free") is just the two numerators added up. num_spam <- p_spam * p_free_spam num_ham <- p_ham * p_free_ham c(num_spam = num_spam, num_ham = num_ham) #> num_spam num_ham #> 0.30 0.06 p_spam_given_free <- num_spam / (num_spam + num_ham) round(p_spam_given_free, 3) #> [1] 0.833

  

One word swings the odds from a 40% base rate to 83% spam. Notice the denominator \(P(x)\) is the same for every class (it is just the two numerators summed), so it never changes which class is larger. For the decision we can drop it entirely and compare only the numerators:

\[ P(y = k \mid x) \;\propto\; P(y = k)\,P(x \mid y = k) \]

The symbol \(\propto\) means "proportional to". The rule for classifying is now one sentence: compute prior times likelihood for each class, and pick the bigger one.

Key Insight
Bayes' rule lets us answer a question we cannot measure directly (is this spam, given the words?) using two things we can count from labeled history (how common spam is, and how typical these words are for spam).