Calendar and event features for production forecasts

Today let's understand how a forecast can quietly read the wrong holiday calendar, and how to build one calendar table that fixes it for good.

Northbay Supply sells online in both the US and India. In November 2023, its India forecast had to cover Diwali, the busiest shopping day of the year there. The forecast expected an ordinary Sunday, somewhere around 300 orders. The real count landed close to 900.

Below are four days from that same week on Northbay's India calendar. Toggle the table to see what each date actually carries once you look past the bare date.

One of those four dates is not like the others. Once you add its weekday and its holiday name, 2023-11-12 stands out as Diwali, the date a forecast has to catch.

Why a calendar feature breaks between training and serving

Let's work out exactly what went wrong with Northbay's Diwali forecast.

Northbay's forecasting model does not only use past order counts. It also uses a holiday flag: a column that is TRUE on a date that is a holiday for a given country, and FALSE otherwise. That flag is a feature, the same way lag values or a day-of-week indicator are features, and it is built from a list of known holiday dates.

A model like this is trained once, usually offline, and then used again and again to produce new forecasts as new dates arrive. The problem is that Diwali's date is not fixed. It depends on the lunar calendar and moves every year, so the list used to flag it has to be kept current.

Here is what happened at Northbay. The team that trained the model knew Diwali fell on 2023-11-12 that year, so the training data correctly flagged that date. But the forecasting service, running separately at serving time, was still relying on a hardcoded list someone had written the year before. That list still carried 2022's Diwali date, 2022-10-24, not 2023's.

Run the code below to see the two flags side by side, for the exact same date.

RInteractive R
# Compare the holiday flag a training pipeline computed against what a serving pipeline computed, for the same date training_holidays <- as.Date("2023-11-12") # Diwali's real 2023 date, known when the model was trained serving_holidays_stale <- as.Date("2022-10-24") # last year's Diwali date, cached at serving time and never refreshed check_date <- as.Date("2023-11-12") check_date %in% training_holidays #> [1] TRUE check_date %in% serving_holidays_stale #> [1] FALSE

  

Same date, two different answers. Training says 2023-11-12 is a holiday. Serving, reading from a list nobody updated, says it is not.

This gap between what a model learned at training time and what it is given at serving time has a name: training-serving skew. It happens whenever the same feature is computed one way for training and a different way, or from different data, for serving. Northbay's Diwali miss is a direct case of it. At serving time, the model never saw a TRUE for 2023-11-12, so it forecast that date like an ordinary Sunday instead of like Diwali.