The Forecast Is Not Accurate to 3.4 Percent.
It is accurate to about 5.6, and the difference is not a modeling error. It comes from having tried thirty-two models against a single year, and that year happening to be the flattest in our history.
Recommendation
Re-size next year's plan against 5.5 percent rather than 3.4, quote December as a range of roughly 7,300 to 9,600 cases rather than a single figure, and change how we choose a forecasting model. The model we are using is not a bad one. The number attached to it was never achievable.
Where the 3.4 percent came from
Last year's bake-off did nothing unusual. Thirty-two candidate models, each scored on the most recent twelve months, lowest error wins. The winner came in at 3.34 percent and that figure went into the production plan, the hiring plan and the raw-material contracts.
We re-ran the exercise with three years held back that were used to choose nothing. The model the bake-off picked delivered 5.57 percent on those three years. Its true error was two-thirds larger than the number we planned against, and it never once reached 3.4 percent in any of the three years.
Two things went wrong, and neither was a mistake
We tested on the flattest year we have ever had. The twelve months the bake-off scored on grew by 0.0 percent. Ten of the seventeen twelve-month windows available in our history grew faster than five percent. Six of the eight models that topped that leaderboard assume the business does not grow at all, and the other two damp their growth toward zero. The test rewarded models that expect a flat year, and we then used one to plan a year that grew 7.9 percent, where it posted 6.93 percent error against 4.53 for the alternative.
We reported the best of thirty-two as though it were the accuracy of one. When you try many models against a single year and keep the winner, some of what you are seeing is skill and some is luck, and the luck does not come with you into next year.

The effect is measurable and it scales with how hard we search. At one model there is nothing to select on and the reported figure is honest. By thirty-two we are overstating by 2.24 percentage points. This matters more, not less, as we automate: forecasting packages routinely try several hundred configurations, and every one of them widens this gap.
What honest selection produced instead
Scoring the same thirty-two candidates from seventeen different starting points rather than one picked a different model. It claimed a worse number, 5.47 percent, and then delivered 4.62 percent on the sealed years. It promised less and produced more, which is the behavior we want from a planning input.
| What it claimed | What it delivered | |
|---|---|---|
| Last year's bake-off | 3.34% | 5.57% |
| Selection across many origins | 5.47% | 4.62% |
| Repeating last year, no model at all | not claimed | 6.76% |
The last row is worth keeping in view. Simply repeating last year month for month, which costs nothing and requires no maintenance, comes within about two points of our maintained model. That is the floor. It is also the cheapest monitoring we will ever have: the month our model stops beating it is the month something has broken.
The planning range was also wrong
The range we publish around the forecast is supposed to contain the actual month 95 times out of 100. Tested properly it contained it 68 times out of 100, and only 57 times at twelve months out. We have been giving the plant a tolerance that is roughly half as reliable as advertised, and worst exactly where the commitments are largest.
Rebuilt from what our forecasts have actually done, month by month of lead time, the range covers 94 percent. It is a little wider than the old one at one month and 2.6 times wider at twelve. Next December is 7,285 to 9,631 cases, not a point.
What we are asking for
One. Choose the model across many starting points, not one year, and expect the accuracy claim to get worse when we do. That is the claim getting honest, not the forecast getting worse.
Two. Hold back a period that nothing is allowed to touch, and open it once a year to check what we actually delivered against what we said.
Three. Publish ranges rather than points for anything beyond a quarter out, and widen them with the horizon.