Machine learning · ML.NET · E-commerce

Demand forecasting for e-commerce with ML.NET: calendar, region and weather

Annual averages hide the patterns that run an online store. A forecasting model built on your own order history can see them, and it does not need a new platform.

Most online retailers plan from last year’s numbers. Last year’s totals are a fine starting point, but they average away the things that actually decide whether you are over- or under-stocked: the weekly rhythm, the sale peaks, the regions that behave differently, and the seasonal effects nobody has measured. The order history to do better is already in your database. This is how we turn it into a forecast.

Why ML.NET

If your store runs on .NET and SQL Server, as most nopCommerce and custom .NET stores do, ML.NET lets you train and run forecasting models in the same stack. No Python service to host, no data pipeline to a separate platform, no customer data leaving your environment. The models are not the most exotic available, and for demand forecasting they do not need to be. What matters is the data preparation and the decomposition, and those you control.

Preparing the series

Forecasting models want clean time series. Order history is not one. Before any modelling:

  • Aggregate deliberately. Decide the grain: units per SKU per day per region, or revenue per category per week. Finer is not always better; sparse series forecast badly.
  • Net out noise. Remove cancellations and returns, test orders and bulk corrections. A refund batch in March is not negative demand.
  • Reconcile the catalogue. SKUs get renamed, merged and replaced. Map old identifiers to current ones so a product does not appear to die and be reborn.
  • Fill the calendar. Days with no orders are zeros, not missing values.

Decompose instead of black-boxing

You could feed the whole series to a single model and take whatever comes out. We do not, for two reasons: the business cannot read the result, and sale events wreck it. Instead we model demand as a sum of components:

  • Baseline trend: the slow growth or decline of the business.
  • Weekly cycle: the day-of-week pattern, which differs between B2C and B2B stores.
  • Yearly cycle: the ordinary seasonality of the category.
  • Sale events: Black Friday, holiday sales, the retailer’s own promotions. These are modelled as explicit effects with a ramp before and a dip after, because customers pull purchases forward.
  • Regional climate seasons: by region, the effect of the local season on the category. This is where the surprises live.

ML.NET’s time-series components (singular spectrum analysis for the cyclical parts, regression for the event and regional effects) handle each of these well. The result is a forecast a planner can decompose: this week is high because of the trend, the weekly pattern and a promotion, in that proportion.

Regions are different businesses

A national store is several regional stores sharing a website. Shipping zones, climate, housing stock and local events all move demand. Fit the regional components separately, with the national trend shared, and you get forecasts that are both stable (the shared part) and local (the regional part). In one of our engagements this is how a rise in sales during hurricane season in one region went from an anomaly to a planned peak.

Backtest, or do not bother

The only honest measure of a forecasting model is how it would have done in the past. Hold out the last twelve months, train on everything before, forecast the held-out period and compare. Report the error per region and per month, not just one overall number, because an average error of ten percent can hide a forty percent miss in December. Re-run this every time you change the model, and show the client the chart. If the backtest does not beat last-year-plus-growth, you are not done.

Putting it into operations

A forecast nobody uses is an experiment. The parts that make it operational:

  • Scheduled retraining from SQL Server, weekly or monthly, with the backtest re-run automatically.
  • Output tables the planning spreadsheets and reports already read, so adoption needs no new tool.
  • Component reporting, so a planner can see why a number is what it is and override it when they know something the model does not.
  • Alerts when actuals drift from the forecast, which is often the earliest sign of a supply problem or a competitor move.

How long this takes

With order history already in SQL Server and a clear grain, a first model with backtests is a matter of weeks, not quarters. The engagement we describe in our forecasting case study took two months with one engineer, inside an existing relationship. The hard part is rarely the model. It is agreeing what to forecast, cleaning the history, and building the habit of planning against the output.

Have a similar problem?

Tell us what you are building. Hiten reads every enquiry and replies within one working day.

Let’s talk