Retail, Demand and Supply Chain
Retail demand data
Retail demand data is a giant table of what sold, where and when, and its shape decides everything you can build on top of it.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Retail demand data is a giant table of what sold, where, and when — one row per product, per shop, per day.
Think about the kirana store near your house. The owner does not guess how much milk to stock. She remembers Mondays sell less than Saturdays. She remembers that wedding season means more ghee. She learned all of this by watching her own shelves, day after day, for years.
She is running a forecasting system in her head. Retail demand data is that same memory, written down as a table instead of kept in one person's head.
Why it exists
A single shop is manageable by memory. A chain with two thousand stores and forty thousand products is not.
Nobody can hold "milk in store 101 on a rainy Tuesday" and "ghee in store 900 during Diwali week" in their head at the same time. Once a business has more than a handful of products and locations, memory stops working. The only way to plan stock is to write every sale down and let a computer find the pattern.
That table becomes the raw material for every forecast, every reorder, and every price change the business makes.
How it works
A row of retail demand data usually looks like this:
date store sku units_sold price on_promo holiday
2026-08-01 store_101 milk_1L 59 58 false false
2026-08-01 store_101 bread_400g 34 45 false false
2026-08-01 store_102 ghee_1kg 3 620 true falseOne product, in one store, on one day, is called a series — short for time series, a sequence of numbers ordered by date. A big retailer might track hundreds of thousands of these series at once, not only one.
That is the first thing that makes retail different from a textbook forecasting example. You are rarely forecasting one number. You are forecasting a huge grid of them, and they all interact. A milk shortage in store 101 might push some of those customers to buy milk in store 102 instead.
Where you have already seen it
- A supermarket's "out of stock" shelf tags. Behind them is a forecast that decided how much to order and got it wrong.
- Blinkit or Zepto's ten-minute delivery promise. It only works because a forecast pre-positioned stock in a nearby dark store before you even opened the app.
- Amazon's "usually dispatched in 1 day" label. That depends on a warehouse already holding the item, based on a forecast made days earlier.
- A shop running out of your favourite biscuit right before a festival. That is a forecast that failed, not bad luck.
The part that is different here
Most forecasting lessons show you one clean line going up and down over time. Retail demand data rarely looks like that.
Real store-level sales are lumpy. A slow-moving product might sell zero units on most days and then eight units on a Saturday. Price changes and promotions punch sudden holes and spikes into the pattern. Store openings and closures start and end series without warning.
Every later lesson in this section deals with one piece of that mess. This lesson is only about knowing what the raw table looks like before you touch it.
Remember this
- Retail demand data is one row per product, per store, per day — not one tidy line, but thousands of short, messy ones at once.
- The scale is the challenge. A chain forecasts hundreds of thousands of series, not a handful.
- Everything downstream — ordering stock, setting prices, planning delivery — starts from this table being reasonably clean.
What to learn next
- You measure sales, not demand — why this table already lies to you a little.
- Intermittent demand and mostly-zero series — what to do when most days show zero.
- What is time series data? — the general idea this section keeps applying to shelves and stores.
Developer — Code and libraries.
Setup
pip install pandas numpyMinimal runnable code
We will build a small slice of a retail table by hand — two stores, three products, two weeks — and look at the shape every real retail dataset shares.
import pandas as pd
import numpy as np
rng = np.random.default_rng(0)
# Two stores, three SKUs, 14 days -- a tiny slice of a real retail table.
# SKU means "stock keeping unit": one specific, sellable product variant.
stores = ["store_101", "store_102"]
skus = ["milk_1L", "bread_400g", "ghee_1kg"]
dates = pd.date_range("2026-08-01", periods=14, freq="D")
rows = []
for store in stores:
for sku in skus:
base = {"milk_1L": 40, "bread_400g": 25, "ghee_1kg": 4}[sku]
for date in dates:
weekend_lift = 1.4 if date.dayofweek >= 5 else 1.0
units = rng.poisson(base * weekend_lift)
rows.append({
"date": date.date().isoformat(),
"store": store,
"sku": sku,
"units_sold": units,
"on_promo": bool(rng.random() < 0.1),
})
demand = pd.DataFrame(rows)
print(demand.head(6).to_string(index=False))
print()
print("rows:", len(demand), " unique store x sku series:", demand.groupby(["store", "sku"]).ngroups)
print()
zero_share = (demand["units_sold"] == 0).mean()
print(f"share of rows with zero units sold: {zero_share:.2%}")
print()
wide = demand.pivot_table(index="date", columns=["store", "sku"], values="units_sold")
print(wide.iloc[:3, :3])date store sku units_sold on_promo 2026-08-01 store_101 milk_1L 59 True 2026-08-02 store_101 milk_1L 68 False 2026-08-03 store_101 milk_1L 41 False 2026-08-04 store_101 milk_1L 41 False 2026-08-05 store_101 milk_1L 43 False 2026-08-06 store_101 milk_1L 43 False rows: 84 unique store x sku series: 6 share of rows with zero units sold: 0.00% store store_101 sku bread_400g ghee_1kg milk_1L date 2026-08-01 34.0 4.0 59.0 2026-08-02 22.0 3.0 68.0 2026-08-03 17.0 4.0 41.0
What actually happened
The data is stored long — one row per (date, store, sku) combination — because that is how a point-of-sale system actually logs a sale. Nobody's checkout scanner writes a wide table.
groupby(["store", "sku"]).ngroups counts the distinct series. Six series here. A real chain runs the same code and gets a number with six or seven digits.
pivot_table reshapes long into wide — one column per series, one row per date. Models that treat every series independently want the long format. Models that look at many series together, like the reconciliation methods in a later lesson, want it wide.
The zero share came out at 0.00% for this sample. Milk and bread sell every single day here, and even ghee_1kg — with its much smaller base of 4 units — did not hit zero in these 14 days either. A slow mover does not need many bad days before it does. Lower the base and check it yourself:
ghee = demand[demand["sku"] == "ghee_1kg"]
print((ghee["units_sold"] == 0).mean())That gap between a fast mover like milk and a slow mover like ghee is the entire subject of the next lesson.
Common mistakes
Treating the whole table as one time series. Summing units_sold across every store and SKU hides the pattern models actually need — a chain-wide total can look flat while individual stores swing wildly in opposite directions.
Forgetting that a missing row is not a zero. If store 102 has no row for bread_400g on a date, that could mean zero sales, a stockout, or a data pipeline failure. Confusing these three is a recurring source of forecast errors, covered directly in the next lesson.
Building one giant wide table for millions of series. A wide pivot with a hundred thousand columns is slow and memory-hungry. Production systems keep data long and only pivot small slices for modelling or inspection.
Try it yourself
Change base for ghee_1kg to 1 instead of 4, and re-run. Watch the zero share jump. That single change turns this from an easy forecasting problem into the hard one covered next.
What to learn next
Researcher — Mathematics and papers.
The panel structure
Retail demand is a panel: a collection of N time series observed over the same T time steps, indexed by keys like store and SKU rather than by one entity alone.
y_{i,t} = units sold for series i (a store x sku pair) at time t, i = 1..N, t = 1..TN is typically 10^4 to 10^6 for a national chain. T is a few hundred to a few thousand days of history. Two modelling philosophies follow directly from this shape.
Local models fit one model per series — classical exponential smoothing or ARIMA, covered in the time series section. They ignore cross-series structure entirely.
Global models pool every series into one model, usually a gradient-boosted tree or a sequence network, with the series identifier (store, SKU, category) fed in as a feature. Montero-Manso and Hyndman (2021), Principles and Algorithms for Forecasting Groups of Time Series, showed that pooling almost always beats per-series models once N is large, because slow-moving series borrow statistical strength from related ones.
Benchmarks that shaped the field
The M5 competition (Makridakis, Spiliotis and Assimakopoulos, 2022, The M5 competition: Background, organization, and implementation, International Journal of Forecasting) is the standard reference dataset here. It released three years of real Walmart sales across 10 stores and roughly 3,000 products, hierarchically organised by item, department, category, store and state. Every winning method was a global model, and the top entries were LightGBM variants, not classical per-series statistics.
Data quality dimensions specific to retail panels
- Structural zeros vs missing rows. A SKU discontinued mid-panel should not silently vanish; it should be represented so a model does not learn a false demand collapse. See You measure sales, not demand.
- Assortment churn. SKUs enter and exit constantly. A global model needs a strategy for genuinely new items, covered in Forecasting a product with no history.
- Calendar effects that are not simple weekly seasonality. Moving holidays (Diwali, Eid, Easter) shift by the lunar or lunisolar calendar year to year, so a fixed day-of-year feature systematically misdates them. Retailers typically maintain an explicit holiday-and-events table joined onto the panel rather than deriving it from the date alone.
Storage and scale
At N = 10^5 series and T = 1000 days, the long-format table has 10^8 rows. Column-oriented formats (Parquet) with per-series partitioning are standard; a naive dense wide matrix of that size (10^5 x 10^3 floats) is about 800 MB per feature before any modelling begins, which is why production pipelines keep data long until the last possible step.
What to learn next
- Streaming feature aggregation — building this table reliably at scale.
- Multivariate forecasting — modelling series that influence each other directly.
- Making store, region and national forecasts add up