Retail, Demand and Supply Chain
Predicting arrival times
The average delivery time is the wrong number to promise a customer, because promising the average means running late on roughly half of all deliveries.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Predicting an arrival time is not the same as predicting the average travel time — because "average" means late half the time, by definition.
Think about telling your parents "I'll be home by 8." You are not naming the time you will typically arrive. You are naming a time you are fairly confident you will beat, even if traffic is a little worse than usual. If you named your average arrival time instead, you would break that promise roughly every second day.
Delivery apps face exactly this choice, at huge scale, every single order.
Why it exists
An ETA, short for estimated time of arrival, sounds like a single number question: how long will this delivery take? It is really two different questions wearing the same name.
"How long does this delivery usually take?" is one question — the average. "How long should I tell the customer, so I am rarely wrong?" is a different question entirely. It is the one that actually matters to someone waiting for their order.
Traffic, weather, and how busy a restaurant's kitchen is all add unpredictable delay on top of the basic travel time. That delay is not symmetric — a delivery can be held up by a lot, but it is rarely finished a lot earlier than expected. Quoting the average time ignores that lopsidedness completely, and it shows up as broken promises.
How it works
Distance: 7.3 km, during rush hour
Predicted AVERAGE time: 50 minutes <- quoting this: late on ~1 in 3 deliveries
Predicted SAFE (90%) time: 65 minutes <- quoting this: late on ~1 in 14 deliveries
Actual time taken: 52 minutesBoth numbers come from the same underlying model of how long deliveries take. The difference is which part of the range of possible outcomes gets quoted to the customer. A responsible ETA system quotes a number that is comfortably likely to be met, not the number that happens to be typical.
Where you have already seen it
- Swiggy, Zomato or Blinkit's delivery countdown, which widens visibly during rain or a cricket match, because the range of likely outcomes widened, not only the average.
- Google Maps' arrival time, which shifts as you drive, updating as new information about the road ahead arrives.
- A courier company's "delivery by end of day" promise. It stays deliberately vague, so it can almost always be kept, unlike a precise time that would often be missed.
Remember this
- Quoting the average delivery time means being late roughly half the time — that is what "average" means.
- A useful ETA quotes a time from the safer end of the range of possible outcomes, not the typical one.
- Uncertainty itself is not constant — rush hour and bad weather widen the range of possible arrival times, not only the average.
What to learn next
- Snapping noisy GPS to roads — the raw signal an ETA system's location data comes from.
- Predict then optimise — why the "safest" prediction depends on what is actually at stake.
- Prediction intervals for regression — the general technique behind quoting a range instead of one number.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
We simulate 800 deliveries with a distance and a rush-hour flag, where rush hour adds both a bigger average delay and much more unpredictable traffic. We train two models on the same data: one predicting the average time, one predicting a "safe" 90th-percentile time.
import numpy as np
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.linear_model import LinearRegression
rng = np.random.default_rng(6)
n = 800
distance_km = rng.uniform(1, 12, n)
rush_hour = rng.integers(0, 2, n)
# Traffic adds a variable delay on top of a base travel time -- and that
# delay is far more unpredictable during rush hour than outside it.
base_minutes = 6 + distance_km * 3.2
traffic_noise = rng.exponential(scale=np.where(rush_hour == 1, 12, 3))
true_minutes = base_minutes + rush_hour * 8 + traffic_noise
X = np.column_stack([distance_km, rush_hour])
y = true_minutes
X_train, X_test = X[:600], X[600:]
y_train, y_test = y[:600], y[600:]
mean_model = LinearRegression().fit(X_train, y_train)
p90_model = GradientBoostingRegressor(loss="quantile", alpha=0.9, n_estimators=150,
max_depth=2, random_state=0).fit(X_train, y_train)
mean_pred = mean_model.predict(X_test)
p90_pred = p90_model.predict(X_test)
late_if_mean_quoted = (y_test > mean_pred).mean()
late_if_p90_quoted = (y_test > p90_pred).mean()
print(f"test deliveries: {len(y_test)}")
print(f"average predicted ETA (mean model): {mean_pred.mean():.1f} minutes")
print(f"average predicted ETA (p90 model): {p90_pred.mean():.1f} minutes")
print()
print(f"share of deliveries LATER than the quoted mean ETA: {late_if_mean_quoted:.1%}")
print(f"share of deliveries LATER than the quoted p90 ETA: {late_if_p90_quoted:.1%}")
print()
i = 3
print(f"example delivery: {X_test[i,0]:.1f} km, {'rush hour' if X_test[i,1] else 'normal hour'}")
print(f" mean-model ETA quoted: {mean_pred[i]:.0f} min")
print(f" p90-model ETA quoted: {p90_pred[i]:.0f} min")
print(f" actual time taken: {y_test[i]:.0f} min")test deliveries: 200 average predicted ETA (mean model): 38.1 minutes average predicted ETA (p90 model): 48.1 minutes share of deliveries LATER than the quoted mean ETA: 34.5% share of deliveries LATER than the quoted p90 ETA: 7.0% example delivery: 7.3 km, rush hour mean-model ETA quoted: 50 min p90-model ETA quoted: 65 min actual time taken: 52 min
What actually happened
LinearRegression predicts the conditional average — the typical delivery time given a distance and a rush-hour flag. Quoting that number produced deliveries running later than promised 34.5% of the time on the held-out test set. That is not a bug; it is what predicting the average means.
GradientBoostingRegressor(loss="quantile", alpha=0.9) predicts a different target entirely: a value that the actual outcome exceeds only 10% of the time. Quoting that number instead dropped the late rate to 7.0%, close to the 10% the model was built to aim for.
The one example delivery makes the trade-off concrete. The mean model quoted 50 minutes; the delivery took 52 — technically late. The p90 model quoted 65 minutes for the same delivery, comfortably covering the actual 52.
Common mistakes
Reporting model accuracy with MAE or RMSE alone and calling it done. Those metrics judge how close the mean prediction is to the truth, on average. They say nothing about how often a specific promised number gets broken, which is the metric that actually matters to a customer.
Using one fixed buffer added to the mean prediction, everywhere. Rush-hour deliveries need a much bigger buffer than off-peak ones, because rush hour is not only slower on average — it is far less predictable, as the wider traffic_noise in this example shows.
Quoting a p90 or higher percentile always, everywhere, "to be safe." An unnecessarily wide ETA that is always beaten easily trains customers to distrust the number, or to think the service is slow when it is actually being cautious.
Forgetting the quantile model needs its own evaluation. A quantile model is well-calibrated when the actual late-rate matches its target percentile — here, close to 10% for a p90 model. Check that number on real held-out data before trusting it.
Try it yourself
Change alpha=0.9 to alpha=0.95 in the GradientBoostingRegressor, aiming for an even safer, rarer-to-miss ETA. Re-run and check how much the average quoted time grows, and how close the actual late-rate lands to 5%.
What to learn next
Researcher — Mathematics and papers.
ETA as conditional quantile estimation
The developer example trains a model to approximate Q_tau(Y | X), the tau-quantile of travel time Y given features X, by minimising the pinball loss:
L_tau(y, y_hat) = (y - y_hat) * tau if y >= y_hat
= (y_hat - y) * (1 - tau) if y < y_hattau— the target quantile,0.9in the exampley— the realised travel timey_hat— the predicted quantile
This loss is asymmetric by construction: underprediction (y > y_hat) is penalised tau / (1 - tau) times more heavily than overprediction when tau > 0.5. Minimising its expectation yields exactly the tau-quantile of the conditional distribution, which is why GradientBoostingRegressor(loss="quantile", alpha=tau) targets a specific percentile rather than the mean. This is the identical mathematical object as the critical-ratio quantile in Turning a forecast into an order quantity; ETA prediction and inventory ordering are, formally, the same asymmetric-loss problem applied to different variables.
Production ETA systems are graph-structured, not point-feature models
Real ETA systems rarely reduce a route to a straight-line distance and a rush-hour flag. Google's ETA system, described by Derrow-Pinion et al. (2021), ETA Prediction with Graph Neural Networks in Google Maps, KDD, represents the road network as a graph, with a graph neural network propagating predicted per-segment travel time (and its uncertainty) across the route graph, explicitly modelling how congestion on one segment correlates with neighbouring segments. This captures spatial correlation in delay that a per-trip feature model, like the one above, cannot: two deliveries sharing a congested stretch of road are not independent, even if their reported distance and time-of-day features look identical.
Calibration, not only sharpness, is the correct target
A quantile forecast can be "sharp" — narrow, confident-looking — without being well calibrated — matching its stated confidence level in practice. Gneiting, Balabdaoui and Raftery (2007), Probabilistic Forecasts, Calibration and Sharpness, JRSS-B, formalise the standard evaluation principle: maximise sharpness subject to calibration, never the reverse. For an ETA system this means: first verify the p90 model's actual late-rate is close to 10% across relevant segments (by distance band, time of day, city), and only then optimise for how tight the interval can be made. See Reliability diagrams and calibration error for the general machinery.
Multi-stop routes compound the uncertainty
A single-leg ETA, as modelled above, is the building block. A route with several stops (a bike delivering to five addresses in sequence) needs the sum of several correlated random delays, not several independent ones added together — treating each leg's delay as independent understates the tail risk of the whole route being late, since the same traffic or weather event tends to delay every leg on a route simultaneously. Production routing-and-ETA systems typically model route-level delay directly, or apply an explicit correlation structure across legs, rather than summing single-leg quantile predictions.
Key references
- Derrow-Pinion, A. et al. (2021). ETA Prediction with Graph Neural Networks in Google Maps. CIKM.
- Koenker, R. & Bassett, G. (1978). Regression Quantiles. Econometrica 46(1).
- Gneiting, T., Balabdaoui, F. & Raftery, A. (2007). Probabilistic Forecasts, Calibration and Sharpness. Journal of the Royal Statistical Society: Series B 69(2).