Checking prediction interval coverage

Today let's work out whether a forecast's prediction interval actually delivers what its percentage claims.

Covewood Teas is a small online tea subscription seller. Here is how many new orders it picked up each week over the last 104 weeks, two full years.

Look at how tight those early week-to-week swings are, and how loose they get by the end. A forecasting model fit to a series like this will not just hand back one number for next week. It will also hand back a range, tagged with a confidence level such as 95%.

The question this whole lesson answers is this: does that range actually contain the real outcome 95% of the time, or is 95% just a label the model prints without anyone checking whether it holds up?

Defining nominal coverage

Let's fit a model to Covewood's history and see what it says about next week.

The model is ETS(A,A,N), the exponential smoothing model with additive error and additive trend and no seasonal part. Fit it to all 104 weeks, then forecast four weeks ahead.

RInteractive R
# Simulate Covewood's 104 weeks of orders, then fit ETS and forecast 4 weeks ahead library(tsibble) library(fable) library(fabletools) library(dplyr) set.seed(2026) week <- 1:104 orders <- round(30 + 0.7 * week + rnorm(104, 0, 4 + 0.1 * week)) orders <- pmax(orders, 1) dat <- tsibble(week = week, orders = orders, index = week) fit_full <- dat |> model(ets = ETS(orders ~ error("A") + trend("A") + season("N"))) fc_full <- fit_full |> forecast(h = 4) hilo_full <- fc_full |> hilo(level = c(80, 95)) forecast_table <- data.frame( week = hilo_full$week, point = round(hilo_full$.mean, 1), lo80 = round(hilo_full$`80%`$lower, 1), hi80 = round(hilo_full$`80%`$upper, 1), lo95 = round(hilo_full$`95%`$lower, 1), hi95 = round(hilo_full$`95%`$upper, 1) ) forecast_table #> week point lo80 hi80 lo95 hi95 #> 1 105 102.4 89.4 115.4 82.6 122.3 #> 2 106 103.1 90.1 116.1 83.3 123.0 #> 3 107 103.8 90.8 116.8 84.0 123.7 #> 4 108 104.5 91.5 117.5 84.7 124.4

  

Read the point column first. It climbs from 102.4 to 104.5, which just continues the climb you already saw in the series. The lo80/hi80 and lo95/hi95 columns are two separate claims wrapped around that same point forecast.

Take week 105. The model's 95% interval runs from 82.6 to 122.3. That 95 is called the interval's nominal level. Nominal here means "as stated" or "as claimed": it comes straight out of the model's own formula, built from how spread out the fitted model's past errors were, not from checking this particular interval against any real outcome yet.

So what does "95% nominal" actually mean? It states a long-run frequency: if you ran this exact forecast over and over, on data generated the same way, the real outcome should land inside the band 95 times out of 100. One forecast cannot test that claim on its own. You need to actually run it over and over, for real, and count.