Back to blogData and analytics

Demand forecasting: evaluate the planning decision, not just the score

Evaluate demand forecasts at the real planning horizon, using historically available inputs, stockout-aware targets, relevant error measures, and policy tests.

A forecast can beat another model on an average error measure and still lead to a worse inventory decision. The gain may come from high-volume products while intermittent items become less reliable. The evaluation may use information that was unavailable when planners had to act. Or the forecast may target tomorrow’s sales when purchasing decisions require a much longer horizon.

Useful forecasting and analytics start with the decision the prediction will inform. Evaluation must reproduce the information, timing, and constraints of that decision closely enough to reveal whether the forecast adds value. Selecting the best score on a convenient dataset is a different task.

Consider a hypothetical distributor replenishing products with varying supplier lead times. Some items sell steadily; others sell infrequently but are important to particular customers. Inventory shortages can suppress recorded sales, and planners sometimes override forecasts for known promotions. These conditions make a single aggregate accuracy score an incomplete guide to procurement.

Define the planning action and its horizon

Identify when the organization commits to an action. A replenishment order placed before a supplier lead time has elapsed cannot use information learned afterward. If planners order weekly for delivery several weeks later, evaluating only next-day forecasts does not test the relevant capability. The horizon should match the commitment, including review intervals where they matter.

Specify the target at the grain used for the decision. Product-level demand may not be enough if inventory is allocated by location. Location-level forecasts may be too noisy for a supplier commitment made across the whole network. Decide which level controls the action and whether forecasts at different levels must reconcile.

Distinguish prediction from the policy that uses it. A point forecast estimates an outcome; a purchasing rule combines that estimate with uncertainty, existing inventory, outstanding orders, and service requirements. A better point estimate does not automatically produce the best stock policy. Keep those responsibilities separate so improvements can be attributed meaningfully.

Set the comparison around an actual alternative. The distributor might already use a seasonal rule or planner estimates. A sophisticated model should improve on a relevant baseline, not merely outperform a deliberately weak one. Document the current method, its available information, and the operating effort it requires.

Reconstruct what was knowable at each decision

Build evaluation datasets around historical decision dates. At each date, use only information that would have been available then. Finalized promotion calendars, corrected inventory records, or later supplier updates can leak future knowledge into a historical test. Their timestamps and revision history matter as much as the date of the event they describe.

Rolling-origin evaluation repeatedly trains or fits using earlier observations and tests later periods. The forecasting textbook listed in the references explains this method and its extension to multiple horizons. For this distributor, the implementation also needs to recreate the forecast horizon and data availability of the purchasing decision.

Avoid preprocessing with the full historical dataset when that process learns from future observations. Imputation, feature selection, and learned transformations can introduce leakage even when the model itself uses a chronological split. Fit those operations within the appropriate training boundary and preserve their versions for reproducibility.

Keep a final evaluation period separate from repeated model selection where practical. Tuning against the same historical windows can overfit the evaluation process. An untouched period does not remove every uncertainty, but it provides a stronger check on whether the chosen method generalizes beyond the comparisons that selected it.

Distinguish observed sales from unconstrained demand

Recorded sales may be limited by stock availability. A product that could not be purchased does not necessarily have zero demand. Treating its sales history as complete demand can teach a model that shortages are normal low-interest periods, potentially reinforcing the problem when the forecast drives future replenishment.

Inspect inventory status and ordering behavior alongside sales. The team may have backorders, attempted orders, or substitute purchases that help interpret shortages. Each signal has limitations: abandoned demand may remain invisible, and customers may change behavior when availability is unreliable. Avoid presenting a reconstructed demand estimate as a directly observed fact.

Define the treatment of censored or uncertain periods. Options include excluding them from particular comparisons, using documented estimates, or evaluating scenarios. The right choice depends on data quality and the planning consequence. Whatever method is selected, report its effect on the population so a cleaner score does not hide exclusion of the hardest cases.

This is a data readiness issue as much as a modeling issue. Assessing data for AI use requires checking whether the target represents the intended decision. A model can learn a technically consistent sales measure that is commercially unsuitable for estimating demand under constrained availability.

Evaluate a replenishment decision: set the planning horizon (match the timing of procurement); recreate known inputs (exclude information learned later); assess errors by segment (keep bias and sparse demand visible); test the stock policy (include constraints and uncertainty).
A lower forecast error is useful only when it improves the replenishment decision under realistic constraints. View full-size graphic

Compare error measures by their decision consequences

No single measure captures every planning concern. Absolute error is interpretable in units but can be dominated by high-volume products. Percentage errors can become unstable or undefined near zero actual demand. Squared error gives larger mistakes more weight. Choose measures with an explicit reason rather than adopting the one a tool displays first.

Inspect bias separately. Persistent underprediction can create a different operating problem from equally sized errors alternating above and below demand. Segment the analysis by volume, intermittency, product importance, location, and horizon where these dimensions affect procurement. Aggregate improvements should not conceal deterioration in a consequential group.

Evaluation viewQuestion it answersLimitation to keep visible
Absolute errorHow large are misses in units?High-volume items can dominate
BiasAre forecasts persistently high or low?Offsetting segment biases can disappear
Horizon-level resultsDoes performance hold at the planning lead time?Short-horizon gains may not transfer
Inventory simulationWhat could the policy change operationally?Results depend on policy and simulation assumptions

Use distributional forecasts or intervals when the planning policy requires uncertainty. Assess their calibration and usefulness over the relevant populations. A narrow interval is not better if it excludes actual outcomes too often. Forecast uncertainty should support a decision about exposure, not decorate a point forecast with an unexplained band.

Evaluate the inventory policy as a separate experiment

Translate candidate forecasts into the replenishment policy used by planners. Include supplier minimums, order cycles, outstanding stock, capacity constraints, and the treatment of lead-time uncertainty where relevant. A simulation that assumes instant replenishment evaluates a different business from the distributor’s real operating environment.

Compare service outcomes and inventory consequences together. Reduced stock may increase missed demand; higher availability may require more working inventory. The organization must decide which tradeoff is acceptable for different product groups. There is no universal model setting that resolves that commercial choice for every item.

Test sensitivity to assumptions. If the new forecast looks attractive only when lost demand is estimated one particular way, present that dependence. If uncertain supplier lead times dominate the result, improving forecast accuracy may have less value than addressing supply variability. This analysis helps prioritize investment beyond the model itself.

Do not claim causal savings from a historical simulation alone. It can establish a reason to test the approach and reveal potential failure modes. Real users, suppliers, and customers may respond differently when the policy changes. Treat live validation as a separate source of evidence with a defined scope and stopping condition.

Make planner overrides visible and evaluable

Planners often know about events absent from the dataset. An override can add useful information about a promotion, customer commitment, or supply disruption. It can also introduce optimism or repeated adjustments that degrade performance. Preserve the original forecast, adjusted forecast, reason, and timing so the contribution can be examined rather than assumed.

Evaluate overrides using the information available at the time. An adjustment made after partial demand has arrived is not directly comparable with the earlier model forecast. Group overrides by reason and consequence instead of judging the entire human contribution from one average. Some intervention types may help while others add noise.

Design the interface to show assumptions and uncertainty without overwhelming users. Explain the relevant horizon, recent demand context, and any data limitations. A planner should understand what the model has considered and what it has not. Blind acceptance and indiscriminate override are both symptoms of a weak decision interface.

Maintain consistent targets through metric governance. If demand definitions change, compare forecasts and actuals under an aligned interpretation. Otherwise, planners can be penalized or models replaced because the evaluation target changed silently rather than because forecasting performance deteriorated.

Release the capability with continuing evaluation

Start with a controlled population where the decision, data, and operating owners are understood. Shadow forecasts can expose timing and integration issues before they influence orders. A limited live trial can then test whether planners can use the output and whether the expected tradeoff holds under actual procurement constraints.

Monitor forecast availability and data completeness as well as statistical accuracy. A strong model that arrives after the ordering deadline is not useful to that cycle. A fallback method should be defined for missing or delayed output, and users should know which method produced the recommendation they received.

Review performance when the business changes. New products, revised supplier terms, assortment changes, and unfamiliar promotions can alter the validity of past comparisons. Diagnose the source of deterioration before retraining automatically. The issue may be a changed target, missing input, or purchasing constraint rather than an obsolete model.

For the distributor, the acceptance question is whether the candidate forecast improves a defined replenishment decision under realistic information and timing. Accuracy remains important, but it gains meaning through the horizon, population, uncertainty, and policy it serves. That is the evidence needed before a promising prediction becomes a dependable planning capability.

References and further reading

Research and editorial perspectives relevant to this topic. The project guidance and illustrative examples are Ayterate editorial analysis. Some research may require registration or a subscription.

Next steps

Discuss the implications for your project.

Share your current environment and the decision you need to make. We can help assess the relevant service scope.

Get in touch

Let’s explore this for your business

Tell us how this topic relates to your plans and what you want to achieve. We’ll help you identify the next step.

Fields marked * are required

We’ll use your details to respond to your enquiry. Read our privacy policy.