Scaling and Transformations
In Lesson 2 you turned Maya's high-cardinality neighborhood column into a leak-free number. Every column in her home-price table is a number now. That is not the finish line, because two numeric columns can be numerically honest and still trip a model up in two very different ways.
The first is scale. Maya's sqft runs into the thousands while beds runs from 1 to 5, and a model that measures distance will hear only the loud feature. The second is shape: a lot_size column with a long right tail distorts the very models that are sensitive to scale. This lesson fixes both, carefully, and keeps the fix leak-free.
By the end you will be able to:
- Explain why unequal feature scales let a big-unit feature drown out the others, and why tree models shrug it off
- Standardize and normalize a feature in R, and read what the numbers become
- Reshape a skewed feature with log, Box-Cox, or Yeo-Johnson, and say what each one requires
- Fit every scaler and transform on the training data alone, so none of it leaks
Prerequisites: you can run R and read its output, and you know what a train/test split is and why leakage matters (from Target Encoding Without Leakage and Train, Validation, Test, and Data Leakage).
Drag the toggle below to feel the "shape" half of the lesson: a lopsided feature pulled toward symmetry.
Everything is a number, on wildly different scales
Each lesson runs in a fresh R session, so let us rebuild Maya's home data right here and look at the numbers we now have to work with.
Look at how different the columns are. sqft is measured in the thousands (it averages about 1,971, with a typical spread of 554), beds lives between 1 and 5 (average 2.85, spread about 1.35), age sits in the tens, and lot_size stretches from 400 all the way to 27,516 with most homes down at the low end. Same table, six numeric columns, and no two of them speak in the same units. The next step shows why that is a problem.