Feature Selection and Spotting Leakage
For six lessons you have been manufacturing features: encoding, transforming, splitting dates, imputing gaps. Now you have more columns than you need, and a new pair of questions. Which features actually earn their place in the model? And is any column secretly cheating?
Picture the model you built to predict which free-trial users of a project-management app convert to a paid plan. You have 600 trials and eight behavioral columns about each one, from weekly logins to support tickets. Some genuinely predict conversion. Some are noise dressed up as signal. And one column your data team joined in later will hand you a stunning test score that collapses the moment you ship. This lesson is about telling all three apart.
By the end you will be able to:
- Rank features by a fast filter, and tell filter, wrapper and embedded selection apart
- Choose the right selection method for the job
- Define target leakage, spot its tell-tale signs, and remove a leak before it fools you
Prerequisites: you can fit and read a logistic regression, you know a train/test split and what overfitting is, and you have met data leakage once already (Train/Validation/Test and Data Leakage). Comfortable with dplyr helps. Every new idea here is taught from scratch.
More features is not more signal
It is tempting to throw every column you have at a model and let it sort them out. It rarely works out well. Each irrelevant feature is one more chance for the model to fit noise instead of signal, so it nudges variance up and generalization down, the overfitting you met in the imputation lesson. Extra columns also cost money to collect and store, slow training and scoring, and bury the two or three features that actually matter under a pile that does not.
So the goal of feature selection is simple to state: keep the smallest set of features that predicts well, and drop the rest. First, the data. Each lesson runs in a fresh R session, so we build the 600 trials right here (run this once). Only the first three columns truly drive conversion; the other five are along for the ride.
Eight candidate features, one target, and a secret we built in: only three columns carry signal. A good selection method should rediscover that from the data alone. There are three families of methods that try, and they differ in one thing: how much they let the model itself do the choosing.