Scaled errors and what the M competitions found
Today let's understand how to judge a forecast's accuracy fairly, even when the things you are forecasting sell at wildly different volumes, and what decades of forecasting competitions found about which methods actually win.
Thompson Supply runs an online catalog with 5 products: Packing Tape, Office Chair, Desk Lamp, Whiteboard Marker Set, and Laser Pointer. Here is how many units of each one it sold a week, on average, over its last 30 weeks.
Packing Tape sells a little over 80 times as many units a week as Laser Pointer does. Whatever measure Thompson Supply uses to judge its forecasts, that measure has to work fairly for both of them, not just for whichever one happens to sell the most.
One portfolio, one forecast, and the raw errors
Thompson Supply forecasts all 5 products the same simple way: the mean method. It looks at a product's last 30 weeks of sales, its training period, averages them into one number, and repeats that number as the forecast for every one of the next 10 weeks, its holdout period. The holdout weeks are held back from training so the forecast can be judged against sales it never got to see.
Build all 5 products, split each into training and holdout weeks, then score the mean method on Office Chair and Laser Pointer with 2 of the standard error measures.
MAE, the mean absolute error, averages the absolute size of every miss in the holdout: how far off the forecast was, ignoring whether it was too high or too low. RMSE, the root mean squared error, does something close, but it squares each miss before averaging and only undoes the squaring with a square root at the end. Squaring first means one unusually large miss counts for more than several small ones.
Office Chair sells about 102 units a week on average, and its mean-method forecast misses by 13.6 units a week on MAE and 17.44 on RMSE. Laser Pointer sells only about 3 units a week, and its forecast misses by just 1.54 units on MAE and 1.735 on RMSE.
Read only the raw MAE and you would say Office Chair's forecast is the worse one, 13.6 against 1.54. But that gap is mostly about how many units each product sells, not about how good either forecast actually is. A product selling 100 units a week will almost always carry a bigger raw error than one selling 3, even when both forecasts are equally good relative to what each product normally does. Thompson Supply needs an error measure that does not just rank products by their sales volume.