Retail, Demand and Supply Chain

Turning a forecast into an order quantity

A forecast gives you a likely number, but the right order quantity depends on which mistake costs more — buying too much or running out.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A forecast tells you the likely demand. It does not tell you how much to order — that depends on which mistake costs you more.

Think of a street vendor deciding how many samosas to fry each morning. Fry too many, and the leftovers are thrown away by evening — wasted oil, wasted flour, wasted time. Fry too few, and customers who wanted one walk away — a sale lost, and maybe a customer who does not come back tomorrow.

Both mistakes cost money. They are almost never the same amount of money.

Why it exists

Imagine the vendor's forecast says "120 samosas will sell today, on average." Ordering exactly 120 sounds sensible. It is usually wrong.

Frying an extra samosa that goes unsold costs ₹8. Losing a customer who wanted one costs ₹20 in lost profit and goodwill — so running out is the far more expensive mistake. It makes sense to fry a bit more than the average forecast, on purpose, to protect against the costlier error.

This problem is famous enough in operations research to have its own name: the newsvendor problem. It is named after the classic example of a newspaper seller who cannot return unsold papers. The answer is not "order the average forecast." It is "order the amount that balances the two different costs of being wrong."

How it works

Forecast says:            demand will average around 120 today

Cost if you order too many:   Rs 8 per leftover samosa
Cost if you order too few:    Rs 20 per lost sale

Since running out costs more than wasting a few samosas,
the right order is HIGHER than the average forecast — maybe 136, not 120.

The exact amount above the average depends on a ratio between the two costs, called the critical ratio. A higher cost of running out, relative to the cost of overstocking, pushes the order quantity up. A higher cost of waste pushes it down.

Where you have already seen it

  • A bakery ordering extra bread before a long weekend, betting that a little waste is cheaper than turning away customers.
  • An event caterer ordering slightly more food than the guest count, because running short at a wedding is far worse than a little leftover food.
  • A pharmacy keeping some buffer stock of common medicines, since a patient unable to get medicine matters more than a slightly bigger inventory bill.

Remember this

  • A forecast gives you a likely number. An order quantity needs the cost of being wrong in each direction, too.
  • When running out costs more than overstocking, the right order is above the average forecast, not equal to it.
  • Getting this balance right, even by a small amount, saves real money every single day it is applied.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

We simulate 200 past days of demand for a samosa stall, set the two costs, and compare ordering the average forecast against the newsvendor-optimal quantity.

newsvendor.py
import numpy as np

rng = np.random.default_rng(4)

# 200 past days of demand for a samosa stall, from the forecast model
daily_demand = rng.normal(loc=120, scale=25, size=200).round().astype(int)
daily_demand = np.clip(daily_demand, 0, None)

# Cost of frying one samosa that does NOT sell (ingredients, oil, wasted time)
cost_overstock = 8
# Profit lost on one customer who wanted a samosa but none was left
cost_understock = 20

critical_ratio = cost_understock / (cost_understock + cost_overstock)
print(f"critical ratio: {critical_ratio:.3f}")

optimal_order = np.quantile(daily_demand, critical_ratio)
mean_order = daily_demand.mean()
print(f"order at the mean forecast:        {mean_order:.0f} samosas")
print(f"newsvendor-optimal order quantity: {optimal_order:.0f} samosas")
print()


def expected_cost(order_qty, demand_samples, c_over, c_under):
    leftover = np.maximum(order_qty - demand_samples, 0)
    shortage = np.maximum(demand_samples - order_qty, 0)
    return (c_over * leftover + c_under * shortage).mean()


cost_at_mean = expected_cost(mean_order, daily_demand, cost_overstock, cost_understock)
cost_at_optimal = expected_cost(optimal_order, daily_demand, cost_overstock, cost_understock)
print(f"expected daily cost, ordering the mean:    Rs.{cost_at_mean:.0f}")
print(f"expected daily cost, ordering the optimum: Rs.{cost_at_optimal:.0f}")
print(f"savings per day from ordering correctly:   Rs.{cost_at_mean - cost_at_optimal:.0f}")
Output
critical ratio: 0.714
order at the mean forecast:        121 samosas
newsvendor-optimal order quantity: 136 samosas

expected daily cost, ordering the mean:    Rs.273
expected daily cost, ordering the optimum: Rs.237
savings per day from ordering correctly:   Rs.36

What actually happened

critical_ratio = cost_understock / (cost_understock + cost_overstock) is the entire idea in one line. Here it comes out to 0.714, meaning: order enough to cover demand on 71.4% of days, not 50%.

np.quantile(daily_demand, critical_ratio) finds exactly that point in the historical demand data — the samosa count that was met or exceeded on 71.4% of past days. That is 136, noticeably above the mean of 121.

expected_cost simulates what actually happens, day by day, under each ordering rule: pay cost_overstock for every leftover unit, pay cost_understock for every unit of unmet demand. The optimal quantity saves ₹36 a day in this small example over ordering the plain average. At real retail scale, across thousands of products and stores, that is where a lot of quiet, unglamorous money either gets made or gets left on the table.

Common mistakes

Ordering exactly the mean forecast. This is only correct when overstocking and understocking cost exactly the same — which is rare in practice.

Estimating the critical ratio quantile from too little history. With 200 days here the 71.4th percentile is reasonably stable. With 20 days of history, that same quantile estimate becomes noisy and unreliable.

Ignoring that the demand distribution itself might have moved. This model assumes daily_demand reflects current conditions. Feed it stale, pre-price-change history and the quantile will be wrong in a way that has nothing to do with this formula.

Forgetting the censoring problem from two lessons ago. If daily_demand was recorded sales rather than true demand, and some of those days were stockouts, this whole calculation inherits that bias. Fix that first.

Try it yourself

Swap the two costs: set cost_overstock = 20 and cost_understock = 8, imagining a product that spoils fast and costs little to lose a sale on. Re-run, and watch the optimal order drop below the mean instead of rising above it.

What to learn next

Researcher — Mathematics and papers.

The newsvendor model, formally

Let demand D be a random variable with cumulative distribution function F. Let c_o be the per-unit overage cost and c_u the per-unit underage cost. Order quantity Q incurs expected cost:

C(Q) = c_o * E[max(Q - D, 0)] + c_u * E[max(D - Q, 0)]

Differentiating with respect to Q and setting the result to zero gives the classical result:

Q* = F^{-1}( c_u / (c_u + c_o) )
  • F^{-1} — the inverse CDF (quantile function) of the demand distribution
  • c_u / (c_u + c_o) — the critical ratio, also called the critical fractile

This is a direct application of quantile regression's loss-minimising property: the quantile F^{-1}(tau) is exactly the value minimising an asymmetric ("pinball") loss with weight tau on underprediction and 1 - tau on overprediction. The newsvendor solution and the quantile-regression optimum are the same object viewed from operations research and statistics respectively.

Extending beyond a single period

The classical newsvendor assumes unsold inventory has zero salvage value and unmet demand is lost outright — a single-period, no-carryover model. Real retail inventory carries over between periods, which turns this into a dynamic inventory control problem: a base-stock or (s, S) policy, where s is a reorder point and S a target stock level, computed from the same critical-ratio logic applied to the demand distribution over the replenishment lead time rather than a single day. Zipkin (2000), Foundations of Inventory Management, is the standard reference for this extension.

Where the demand distribution itself comes from

Using empirical quantiles, as in the code above, implicitly assumes the historical sample is a good estimate of the future distribution — reasonable when demand is stable, weak when it has trend, seasonality, or has been distorted by past stockouts (see You measure sales, not demand). Modern practice increasingly replaces the empirical-quantile step with a quantile regression model — LightGBM with a pinball loss objective, or a neural forecaster trained to output multiple quantiles directly — conditioned on price, promotions and seasonality, rather than reading a single unconditional quantile off raw history.

Cost misspecification is the dominant real-world failure mode

c_o and c_u are rarely known exactly. c_o typically includes holding cost, spoilage risk and markdown risk; c_u typically includes lost margin plus a harder-to-quantify customer-goodwill cost. Sensitivity analysis — sweeping the critical ratio across a plausible range of cost assumptions and checking how much Q* moves — is standard practice before committing an ordering policy to production, since an overconfident point estimate of c_u is a more common source of poor ordering decisions than an inaccurate demand forecast.

Key references

  • Arrow, K., Harris, T. & Marschak, J. (1951). Optimal Inventory Policy. Econometrica 19(3). The original newsvendor formulation.
  • Zipkin, P. (2000). Foundations of Inventory Management. McGraw-Hill.
  • Koenker, R. & Bassett, G. (1978). Regression Quantiles. Econometrica 46(1). The statistical theory connecting quantiles to asymmetric loss minimisation.

What to learn next