ERP / Source date:

Demand Forecasting in ERP: Statistics Beat Gut Feel

Basic statistical models consistently outperformed manual planner adjustments on service levels.

Staged demand planner comparing baseline, adjusted and actual columns beside a lead-time horizon card, without forecast results.

Ask a planning team how accurate their forecast is and you will usually get one of two answers. Either nobody measures it, or somebody quotes a percentage that turns out to be measured at an aggregation level and a time horizon chosen because it produces a comfortable number. Demand forecasting in ERP has been a standard module for two decades, and in a large share of mid-market implementations it is switched on, ignored, and overwritten every month by a planner with a spreadsheet and an opinion. The uncomfortable finding, repeated across enough organisations to be a pattern rather than an anecdote, is that a plain statistical baseline usually beats the adjusted number. Not always, and not for every item, but often enough that the adjustment process itself deserves to be audited rather than assumed.

Measure whether the humans are helping

The single most useful thing a planning function can do this quarter costs nothing and needs no software. Record three numbers every cycle for every item: what the statistical model produced, what the final approved forecast was after adjustment, and what actually happened. That comparison answers a question most organisations have never asked, which is whether the adjustment step adds value or subtracts it. The typical result is that adjustments help for a small number of items where the planner genuinely knows something the history cannot contain, a new customer, a lost contract, a promotion, a tender win, and hurt everywhere else, because adjusting is what planners are paid to do and an unchanged forecast feels like an absence of work. This is not an argument for removing people. It is an argument for pointing them at the items where their knowledge is the input, and leaving the long tail to arithmetic.

Measure whether overrides add valueQualitative measurement sequence from the article, not evidence of improved accuracy or service levels.
  1. Retain the baseline

    Keep the statistical forecast for each item and planning cycle.

  2. Record the adjustment

    Retain the approved number and the information behind the override.

  3. Compare with actual demand

    Evaluate both forecasts at the same decision level and lead-time horizon.

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Get the measurement right before the model

Accuracy figures are easy to manipulate without meaning to. Aggregation level. A forecast is almost always accurate at the national annual level and almost always poor at the item-location-week level. Measure at the level at which decisions are actually made, which is usually the level at which you place a purchase order. Horizon. Measure at the lead time that matters. Forecasting next week accurately is not useful if the replenishment lead time is eleven weeks. Bias separately from error. Persistent over-forecasting and random error are different diseases with different treatments. Bias is organisational, error is statistical, and averaging them together hides both. Weighting. Percentage error treats a slow-moving spare part and your highest-volume line as equals. Weight by volume or value, or the number will be dominated by items nobody cares about. Demand, not sales. Sales history is censored. If you were out of stock, the zero in the history is not demand; it is an absence of supply, and a model trained on it will helpfully forecast continued failure. Capturing lost sales and substitutions, even crudely, is worth more than a better algorithm.

Match the method to the demand pattern

Most ERP forecasting failures are a single model applied to items that behave nothing alike. High-volume, regularly moving items with seasonality respond well to exponential smoothing with trend and seasonal components, which is unglamorous and hard to beat. Intermittent and lumpy items, which in a spare parts or industrial business can be the majority of the catalogue, need methods built for intermittency, and applying standard smoothing to them produces confident nonsense. Items driven by a handful of known customers or projects should not be statistically forecast at all; they should be planned from the pipeline. New items have no history and need a similar-item profile and an explicit review date. Segmenting the catalogue by volume and by variability, then assigning a method and a review frequency to each segment, produces more improvement than any algorithm upgrade. It also tells you where to spend planner attention, which is the scarce resource.

The machine learning question

Every planning vendor is now selling demand sensing and machine learning, and the claims deserve the same scepticism as any other. Learned models do genuinely beat classical methods where there are causal drivers with signal, promotions, price changes, weather, events, and where there is enough clean history to learn from. They do not rescue a business whose history is three years long, whose item master is inconsistent, and whose stockouts are unrecorded. The order of operations is unchanged: fix the demand history, segment the catalogue, get a statistical baseline running and measured, then evaluate whether a more sophisticated model adds anything on top. Organisations that skip to the end buy a model that learns from bad data and produces a more expensive version of the same error.

Practical Guidance for a Forecasting Capability Assessment

  • Start recording baseline, adjusted and actual every cycle. Three columns. Within two quarters you will know whether your adjustment process helps, which is the finding that changes behaviour.
  • Measure at order level and lead-time horizon. Any accuracy number quoted at a higher aggregation or a shorter horizon than the decision it supports is decoration.
  • Separate bias from error and report both. Chronic bias is usually incentives, not statistics, and no model will fix it.
  • Capture lost sales, even approximately. A flag on the order line when stock was unavailable is a small change that improves every future forecast.
  • Segment by volume and variability, then assign methods. Fast and stable to smoothing, intermittent to intermittent methods, project-driven to pipeline planning, new items to profiles with review dates.
  • Separate the forecast from the stocking policy. Safety stock, service level and reorder rules are separate decisions that people routinely change by editing the forecast, which corrupts both.
  • Give planners the exceptions, not the catalogue. Alert on the items where the model and reality have diverged, and leave the rest alone. Attention spread evenly is attention wasted.
  • Review the calendar assumptions annually. Seasonal profiles, holidays and event effects drift, and a model carrying a seasonality learned three years ago will confidently repeat an event that no longer happens.

The Regional Angle

The most common forecasting failure in this region is calendar. Standard seasonal models learn an annual pattern indexed to the Gregorian calendar, and the largest demand events here do not sit still in it. Ramadan and the two Eids move by roughly eleven days each year, so a model that learned a March peak two years ago will place it in March again while the actual demand has walked into February. For grocery, food service, retail, hospitality, logistics and anything consumer-facing, that single error is larger than everything the algorithm choice contributes. The fix is not a better model; it is treating the Hijri calendar as an explicit regressor or building the seasonal profile against it, and then layering the summer exodus, school terms and national holidays on top. The second regional distortion is in the history itself. The introduction of value added tax at the start of last year produced a pull-forward of purchasing in the final weeks of 2017 and a corresponding trough in the first quarter of 2018, and that spike and dip now sit in the training data of every model in the Gulf. Left unadjusted, the model will helpfully forecast a January collapse every year. Anyone building or retraining a forecast this year needs to treat that period as an outlier and say so in writing, because the next person to look at the model will not know. Third, the demand signal is frequently not yours. Much of the region's distribution runs through dealers, resellers and trading partners, so what the ERP records is sell-in rather than sell-out. Forecasting from sell-in means forecasting your partners' stocking decisions, which are driven by their credit terms and your quarter-end incentives rather than by end demand. Even partial sell-out data from the largest partners changes the quality of the forecast more than any software purchase. Fourth, the supply side makes accuracy worth more here than in shorter-lead-time markets. Sea freight from the Far East, consolidation through Jebel Ali, customs clearance and onward movement into Saudi Arabia or across the Gulf mean replenishment horizons measured in months rather than weeks. A long lead time is exactly the condition under which forecasting pays, because there is no possibility of reacting instead. The re-export business complicates this further: goods transshipped onward are demand from a warehouse's perspective and not consumption anywhere nearby, and mixing the two produces a forecast that describes nothing. Finally, contracting and project-driven businesses, which are a large share of the regional economy, should stop pretending their demand is statistical. Government and developer project awards are lumpy, politically timed and knowable in advance through the pipeline. Those items belong in a planning process that reads the tender book, not in a smoothing model.

The objection worth taking seriously

The objection is that forecasting is an expensive way to be wrong. Demand in a volatile market is not a signal with noise around it; it is genuinely unpredictable at the item level, and effort spent chasing accuracy would be better spent on responsiveness. Shorten lead times, hold buffer where it is cheap, build supplier flexibility, and the forecast stops mattering. There is a serious literature behind this position and a lot of practical experience supporting it. The harder version is about what happens when accuracy does improve. Most organisations that raise forecast accuracy see no inventory reduction, because safety stock parameters were set years ago by a different team and are never revisited. The improvement is real and the benefit is never collected. Worse, in organisations where a forecast miss is treated as a personal failure, planners protect themselves by biasing high, and every accuracy programme becomes an exercise in producing defensible numbers rather than correct ones. Both arguments are right about where the value sits, and neither argues for leaving the module switched off. Responsiveness is the better investment wherever the response time is shorter than the lead time, and in this region it frequently is not, which is precisely when forecasting earns its keep. The honest position is narrow and defensible: forecast where lead time exceeds your ability to react, measure whether the process adds value, connect accuracy improvements to stocking parameters so the benefit is actually taken, and stop grading planners on a number that rewards hedging.

Common Questions

What is a good forecast accuracy figure?

There is no useful universal benchmark, because the achievable number depends entirely on demand variability. A stable consumer line and an intermittent spare part are not comparable. The meaningful benchmark is your own statistical baseline: if the adjusted forecast does not beat it, the adjustment process is the problem.

Should planners be allowed to override the model?

Yes, where they hold information the history cannot contain, and with the override recorded so its value can be measured. Unrecorded, unmeasured overrides are the most common reason a well-configured forecasting module produces worse results than doing nothing.

Do we need a separate planning system?

Usually not at mid-market scale. The forecasting capability in a mainstream ERP is adequate for most catalogues once the data and the segmentation are right. Specialist systems earn their place at high item counts, complex multi-echelon networks or heavy promotional activity.

What should we expect over the next twelve months?

Expect every planning vendor to put machine learning at the front of the proposal, and expect the honest ones to concede that the gains come from causal data rather than from the algorithm. Expect demand sensing claims to be tested against short regional histories and found wanting. And expect the organisations that spent this year fixing demand history, stockout capture and calendar effects to quietly outperform the ones that bought a model.


Forecasting Capability Assessment — we measure whether your planners are improving the number or moving it, fix the calendar and history problems underneath, and point attention at the items where judgement actually pays.

Continue reading

Talk to OPS

Start with the operating problem.