Forecasting thousands of series with fable
Today let's understand what changes when a forecasting job stops being about one series and becomes about a few thousand of them at once.
Australia's Bureau of Statistics publishes monthly retail turnover for every industry in every state. Below are 2 real years of that turnover, in millions of Australian dollars, for 8 Victorian retail industries: cafes and restaurants, takeaway food services on its own, food retailing, household goods, clothing and footwear, liquor, newspapers and books, and department stores.
Overlaid on one chart, the 8 lines tangle into one mess. Food retailing and household goods run in the thousands while newspapers and books runs in the tens, so on a shared scale the smaller industries flatten into near-straight lines near the bottom and you cannot read any of their shape. Toggle to small multiples and each industry gets its own panel with its own scale, and now every one of the 8 reads on its own.
One model() call, a model for every key
In a tsibble, the key is the column that says which rows belong to the same series. The lesson's data comes from tsibbledata::aus_retail, Australia's real monthly retail turnover by state and industry, and its key is Industry: every row that shares an industry name is one series.
Build the batch this lesson uses throughout: 8 real retail industries in Victoria, from January 2010 to December 2018. That is 108 months for each industry, 864 rows in total.
key() confirms the series are split by Industry, n_keys() says there are 8 of them, and nrow() confirms 864 rows: 8 industries times 108 months.
Now fit one ETS model to every one of those 8 industries, with a single call to model().
fabletools calls this result a mable, short for "model table." It has one row per key and one column per model you asked for, here just ets. Look down that column: 6 different ETS specifications turn up across the 8 industries, because ETS() searches for the best error, trend and season combination separately for each series, from that series' own data.
That is the whole mechanism of fitting at scale: one call to model(), one fit per key, and the result comes back as one row per key. Nothing about the code changes whether there are 8 keys or 8,000.