Diagnosing a drop in forecast accuracy
Today let's understand how to tell apart the different reasons a forecast's accuracy can suddenly get worse.
Carrow Outdoor Supply is a regional garden-and-hardware chain. It forecasts how many of its best-selling patio umbrellas it will sell each week, using nothing but that week's price. The forecast is a straight line, fit once on two years of past data, and for months at a stretch it stays accurate to within about 3 to 5 umbrellas a week.
Below is that forecast's rolling 4-week average error, week by week, for the weeks after it went live.
One of the four bumps towers so far above the rest that it flattens the other three into what looks like a flat line beside it. Each of the four happens for a different reason, and this lesson works through them in the order they are cheapest to check.
Why "retrain the model" is usually the wrong first move
When a forecast's accuracy suddenly drops, the instinct is to retrain the model right away. That is usually the wrong first move, and here is why.
There are four different things that can make a forecast go wrong, and they are not equally likely or equally cheap to check. In the order worth checking them:
- A pipeline bug: something broke in how the data reaches the model, and the model itself was never the problem.
- A distribution shift in an input: one of the model's inputs moved outside the range it was trained on.
- A structural break: something in the real world changed the series itself, permanently.
- Model decay: the relationship between the inputs and the outcome has genuinely drifted, little by little.
Each cause down this list is more real, and more expensive to be wrong about, than the one before it. A pipeline bug costs you a five-minute look at a log file. Retraining a model to chase what was actually a data feed mistake costs you a broken model, and the same bug still sitting there, waiting to break the next one too.
So before touching the model, Carrow's forecast needs a number to compare everything else against: how accurate does it run normally?
Carrow's forecast predicts weekly sales of its best-selling patio umbrella from price alone, using a straight-line regression fit once on 104 weeks, two years, of past sales. Press Run to fit that line and check how far off it usually runs.
The fitted line says predicted sales = 113.96 - 1.61 x price: every extra dollar on the umbrella's price costs Carrow about 1.6 sales a week. Over two ordinary stretches of weeks, its mean absolute error, or MAE, the average size of its miss regardless of direction, comes out at 2.86 umbrellas in one stretch and 4.57 in the other, against sales that average around 82 a week.
Call that Carrow's usual error: about 3 to 5 umbrellas a week. Every cause in this lesson is a way that number can blow up, and the shape of the blow-up is what tells you which cause it is.