Monitoring and drift
Dev's meal-kit cancellation model launched three months ago at about 79% accuracy, and everyone moved on. This morning the retention team notices something odd: the discount emails the model targets are landing on subscribers who were never going to leave, while the ones quietly cancelling get nothing. Nobody changed a line of code. The model did not break. The world it was trained on did.
That is drift, and it is the quietest failure in machine learning: no error, no crash, just decisions getting worse while every dashboard stays green. This lesson is about catching it before your users do.
By the end of this lesson you will be able to:
- Explain why a model that was accurate at launch decays even though its code never changes
- Detect data drift by comparing live inputs to what the model trained on, and put a number on it with the population stability index (PSI)
- Tell data drift from concept drift, and decide when a monitoring signal means it is time to retrain
Prerequisites: Lesson 4 (predictions flowing in batch or real time), and you can [fit a model and read predict output](Your-First-End-to-End-Model-in-R.html).
A model is a snapshot of a world that keeps moving
When Dev trained the cancellation model, it learned the patterns in one specific slice of time: last quarter's subscribers, last quarter's prices, last quarter's reasons for leaving. Training freezes those patterns into fixed coefficients. From that moment the model is a photograph, and photographs do not update themselves.
Then the world moves. Two months after launch a cheaper rival, FreshBox, opens for business. Boxes get discounted to compete, so the typical price a subscriber pays drifts down. And the reason people cancel changes: where idle weeks used to be the main signal, now even loyal, long-tenure customers leave, lured by the cheaper option. The model never saw any of this. It keeps applying last quarter's rules to this quarter's world, and quietly gets them wrong.
We do not have to take this on faith. Let us rebuild Dev's model from scratch, then score it on a fresh batch from the launch-era world and on a batch from the world three months later, and read the two accuracies side by side.
Same model, same code, fourteen points of accuracy gone. Nothing alerted, because from the code's point of view nothing happened. The only way to catch this is to watch for it on purpose. The rest of the lesson is how.