Retail, Demand and Supply Chain
You measure sales, not demand
A stockout hides the true demand behind whatever was on the shelf, and averaging recorded sales alone quietly teaches a model to order too little, forever.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A stockout hides the real demand. Your sales number stops the moment the shelf goes empty.
Think about buying tickets for a big cricket match or a popular movie on release day. The counter shows "500 tickets sold" and then closes. That number is not how many people wanted a ticket. It is how many tickets existed. Hundreds more people wanted one and were turned away, uncounted.
A shop works the same way. If ten customers wanted milk but only five packets were on the shelf, the till only records five sales. The other five customers walked out, and nothing captured that they ever wanted milk at all.
Why it exists
Every point-of-sale system records units_sold. Almost none of them record units_wanted.
That gap would not matter if shelves were always full. They are not. Stores run lean on purpose, because unsold stock costs money to store and can go stale. So stockouts happen often, especially for popular items right before they get restocked.
Here is the trap. If you feed raw sales history straight into a forecast, the forecast learns "this product sells about 5 units a day." That stays true even on days it could have sold 10, if only there had been more stock. The forecast then orders enough for 5. The shop stocks out again. The forecast sees another day capped at 5, and gets even more confident that 5 is correct.
This is called demand censoring — censoring means the true number is cut off and hidden, not that it is unrecorded by accident. A forecast trained on censored sales quietly teaches itself to always under-order.
How it works
What really happened: 12 customers wanted milk today
What the shelf could hold: 8 packets
What the till recorded: "sold: 8" <- this is what your data has
What a naive average learns: "this product sells about 8 a day"
(wrong — it undercounts every time this happens)The record says 8. The truth was 12. Nothing in the raw data tells you those are different numbers, unless you also know whether the shelf ran empty that day.
Where you have already seen it
- A popular kurta size selling out online within a day of a sale starting. The site's "8 sold" counter is real, but demand was plainly higher — the size was gone by evening.
- LPG cylinder booking slots filling up in minutes during a shortage. The booking count hits a hard ceiling that has nothing to do with how many households actually needed a refill.
- A restaurant's most popular dish marked "sold out" by 9pm. Every night it happens, the kitchen's own numbers say demand was lower than it really was.
Remember this
- Recorded sales are capped at whatever stock was available — they are not the same thing as demand.
- A model trained only on capped numbers will learn to order too little, and the shortage repeats itself.
- The fix starts with tracking whether a stockout happened, not only how much was sold.
What to learn next
- Intermittent demand and mostly-zero series — the next way real sales data misleads a forecast.
- Turning a forecast into an order quantity — why getting this right changes how much stock you actually order.
- Anomaly detection — spotting the days something unusual, like a stockout, happened.
Developer — Code and libraries.
Setup
pip install numpy pandas scipyMinimal runnable code
We simulate a slow-moving product with real demand nobody directly observes, and a shelf that only holds 5 units. Then we compare three ways of estimating "true" average demand from the sales record alone.
import numpy as np
import pandas as pd
from scipy import stats
from scipy.optimize import minimize_scalar
rng = np.random.default_rng(7)
n_days = 200
true_demand = rng.poisson(lam=6, size=n_days) # nobody in real life observes this column
shelf_stock = 5
units_sold = np.minimum(true_demand, shelf_stock)
stocked_out = units_sold >= shelf_stock
df = pd.DataFrame({"true_demand": true_demand, "units_sold": units_sold, "stocked_out": stocked_out})
print(df.head(8).to_string(index=False))
print()
true_mean = df["true_demand"].mean()
naive_mean = df["units_sold"].mean()
stockout_rate = df["stocked_out"].mean()
print(f"true average demand: {true_mean:.2f} units/day")
print(f"naive average of recorded sales: {naive_mean:.2f} units/day")
print(f"share of days that sold out: {stockout_rate:.1%}")
dropped_mean = df.loc[~df["stocked_out"], "units_sold"].mean()
print(f"average using only non-stockout days: {dropped_mean:.2f} units/day <- WORSE, not better")
print()
def neg_log_likelihood(lam, sold, censored, limit):
n_censored = censored.sum()
uncensored_ll = stats.poisson.logpmf(sold[~censored], lam).sum()
# censored days only tell us true demand was >= the shelf limit
censored_ll = n_censored * stats.poisson.logsf(limit - 1, lam)
return -(uncensored_ll + censored_ll)
result = minimize_scalar(
neg_log_likelihood,
bounds=(0.5, 30),
method="bounded",
args=(df["units_sold"].to_numpy(), df["stocked_out"].to_numpy(), shelf_stock),
)
print(f"censored-MLE estimate of demand: {result.x:.2f} units/day <- close to {true_mean:.2f}") true_demand units_sold stocked_out
6 5 True
7 5 True
8 5 True
6 5 True
2 2 False
3 3 False
8 5 True
6 5 True
true average demand: 5.90 units/day
naive average of recorded sales: 4.50 units/day
share of days that sold out: 71.0%
average using only non-stockout days: 3.28 units/day <- WORSE, not better
censored-MLE estimate of demand: 6.01 units/day <- close to 5.90What actually happened
units_sold = min(true_demand, shelf_stock) is the whole problem in one line. This is exactly what a real till does — it cannot record a sale for stock that was not there.
The naive average, 4.50, undershoots the true 5.90. Expected, since 71% of days were capped.
The surprising result is the "drop the censored days" row. It looks like a sensible fix — "the capped days are lying to me, so ignore them" — and it makes things worse, not better: 3.28, further from the truth than the naive average was. Dropping censored days keeps only the days demand happened to fall below the shelf limit, which are exactly the low-demand days. That is a selection bias, not a fix.
The last line is a proper fix: a censored likelihood. For an ordinary day we know the exact demand. For a stocked-out day we only know demand was at least 5 — so instead of pretending we saw the number 5, the model asks "what value of average demand makes both kinds of days most probable?" scipy.optimize.minimize_scalar searches for that value directly.
Common mistakes
Dropping censored rows entirely, shown above, biases the estimate downward, not toward the truth.
Forgetting to track whether a stockout happened at all. If your data pipeline never logs stocked_out, none of this is fixable after the fact. Log it from day one.
Applying this fix to a single lucky value instead of a distribution. The MLE above estimates an average rate. Real ordering decisions need the shape of the whole distribution, covered in Turning a forecast into an order quantity.
Assuming every zero is a stockout. A zero can also mean genuinely zero demand. Confusing the two directions of this problem is the subject of the next lesson.
Try it yourself
Change shelf_stock to 3 and re-run. Watch the stockout rate climb and the naive average fall further behind the truth — while the censored-MLE estimate keeps tracking it.
What to learn next
Researcher — Mathematics and papers.
The formal problem: demand unconstraining
Retail calls this demand unconstraining: recovering an estimate of unconstrained demand D from sales S that are censored by available inventory I.
S = min(D, I)
c = 1[S = I] (the censoring indicator: did the day hit its ceiling?)This is right-censored data in the classical survival-analysis sense, with inventory playing the role of the censoring time. The likelihood for a parametric demand model with density f and CDF F is
L(theta) = product over uncensored days f(s_i ; theta)
x product over censored days (1 - F(i_i - 1; theta))theta— parameters of the assumed demand distribution (e.g. the Poisson ratelambda, or a negative binomial's mean and dispersion)f,F— the probability mass/density function and cumulative distribution function of the assumed demand modeli_i— the inventory ceiling on dayi
Maximising the log of this likelihood is exactly the MLE computed above. It is the standard approach in the operations-research literature going back to Nahmias (1994), Demand Estimation in Lost Sales Inventory Systems, Naval Research Logistics.
Why Poisson is often the wrong assumption in practice
Poisson assumes the mean equals the variance. Retail demand is usually overdispersed — variance exceeds the mean — because demand is itself a mixture of a "normal day" rate and occasional spikes (payday, a local event, a competitor's stockout). The standard practical fix is a negative binomial demand model, which adds a dispersion parameter and nests Poisson as a special case as that parameter goes to zero. Refitting the same censored-likelihood approach with scipy.stats.nbinom requires only swapping the density and survival functions.
EM as an alternative to direct MLE
Where the demand model is complex (e.g. it depends on price, promotions and weather through a regression), direct likelihood maximisation over a high-dimensional censored objective is harder to fit reliably. The standard alternative is an expectation-maximisation (EM) loop: given current parameters, impute the expected demand on censored days as E[D | D >= I]; refit the regression on the completed data; repeat until the imputed values stop changing. This converges to the same MLE under correct model specification and is easier to implement inside an existing forecasting pipeline that already fits a regression.
What "lost sales" costs, and why this matters economically
Under-forecasting demand from censored sales feeds directly into the reorder quantity. See the newsvendor model in the next lesson: an underestimated demand mean shifts the optimal order quantity down, which increases the future stockout probability, which produces more censored days, which reinforces the underestimate. Left uncorrected, this feedback loop is self-sustaining — a structurally under-ordered product rarely fixes itself through more data alone.
What to learn next
- Turning a forecast into an order quantity — where the corrected demand estimate is actually used.
- Intermittent demand and mostly-zero series — censoring compounds with sparsity for slow movers.
- Reliability diagrams and calibration error — checking whether a corrected demand model is honest about its own uncertainty.