Feature stores
A feature store keeps one definition of each feature for both training and live serving, which is the only reliable cure for a model that scores well offline and badly in production.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A feature store is one shared place where each model input is defined once, for both training and live serving.
Think about two cooks in one kitchen working from the same handwritten recipe. One reads "one cup of rice" as a small steel cup. The other uses a large plastic mug.
Both follow the recipe. Both are careful. The two dishes come out different, and neither cook did anything wrong.
That is what happens when two teams each write their own version of the same model input.
What a feature is
A feature is one number the model looks at. Age. Amount spent last month. How many days since the last order.
Most are not stored anywhere. They are calculated from raw records, and the calculation is where the trouble starts.
Take "average spend in the last 30 days". Thirty days from when? Does today count? Do refunds count? Four reasonable people write four different answers.
Why it goes wrong
TRAINING LIVE SERVING
a data scientist writes an engineer writes
a query over the whole code that looks at
history table the last 30 days
average = 732 average = 667
\ /
\ /
------> the same model <-------
trained on 732,
asked about 667The model was taught what 732 means. It is handed 667 for the same customer. Nobody wrote a bug and the answers are wrong anyway.
This has a name: training and serving skew. It is the gap between the numbers a model learned from and the numbers it later receives.
The second, subtler trap
Imagine predicting whether a customer will stop shopping with you. You build a feature: how much they spent, averaged over their whole history.
For a customer who left in March, that average includes April, May and June. Those months had not happened when the decision was needed.
Your model looks brilliant in testing and useless in production. It was reading the future.
The rule is: every feature must use only what was known at the moment the decision was made.
What a feature store does
Three things, and none of them is exotic:
- Keeps one definition of each feature, in one place, used by both sides.
- Serves the same value fast when a live request arrives.
- Can rebuild a feature as it was on any past date, so training never reads the future.
Somewhere you have seen this
A food delivery app estimating your arrival time uses "how busy this restaurant is right now". That number has to mean exactly the same thing in the model's training data as it does at 8pm on a Saturday.
What is honestly hard here
Feature stores are heavy. They add a database, a service and a deployment to your system.
For a small team with one model, a shared Python function achieves most of the benefit. It costs you nothing. Reach for the full machinery when many models share features, or when several teams are involved.
Do not install one because it appeared in a diagram at a conference.
Remember this
- Two teams writing the same feature twice produces two different features.
- Every feature must use only what was known at decision time.
- A shared function is the cheap version. A feature store is the same idea with a database attached.
What to learn next
- ETL for machine learning — the pipeline that materialises these features.
- Monitoring and model drift — watching for skew after launch.
- Feature engineering — deciding what the features should be in the first place.
Developer — Code and libraries.
Two failures, measured. First the skew between two reasonable definitions. Then the fix.
Setup
pip install numpy pandas scikit-learnTwo teams, one feature, two answers
import numpy as np
import pandas as pd
rng = np.random.default_rng(5)
rows = []
for cid in ["c1", "c2", "c3"]:
for d in pd.date_range("2026-06-01", "2026-08-19", freq="D"):
if rng.random() < 0.35:
rows.append((cid, d, round(float(rng.gamma(2.0, 400.0)), 2)))
tx = pd.DataFrame(rows, columns=["customer", "ts", "amount"])
print("transactions:", len(tx), " range:", tx["ts"].min().date(), "to", tx["ts"].max().date())
# The training team's version: a batch job over the whole table.
train_feature = tx.groupby("customer")["amount"].mean().round(2)
# The serving team's version: last 30 days, computed live.
now = pd.Timestamp("2026-08-20")
serve_feature = (tx[tx["ts"] > now - pd.Timedelta(days=30)]
.groupby("customer")["amount"].mean().round(2))
skew = pd.concat([train_feature.rename("training"), serve_feature.rename("serving")], axis=1)
skew["difference"] = (skew["serving"] - skew["training"]).round(2)
skew["percent"] = (100 * skew["difference"] / skew["training"]).round(1)
print()
print(skew)
# One shared definition, used by both paths, with the cut-off passed in.
def avg_amount_30d(events, customer, as_of):
window = events[(events["customer"] == customer)
& (events["ts"] <= as_of)
& (events["ts"] > as_of - pd.Timedelta(days=30))]
return round(float(window["amount"].mean()), 2) if len(window) else 0.0
print()
print("shared definition, as_of 2026-08-20:", {c: avg_amount_30d(tx, c, now) for c in ["c1", "c2", "c3"]})
print("shared definition, as_of 2026-07-01:", {c: avg_amount_30d(tx, c, pd.Timestamp('2026-07-01')) for c in ["c1", "c2", "c3"]})transactions: 91 range: 2026-06-02 to 2026-08-18
training serving difference percent
customer
c1 732.32 666.85 -65.47 -8.9
c2 927.64 750.22 -177.42 -19.1
c3 827.92 1053.79 225.87 27.3
shared definition, as_of 2026-08-20: {'c1': 666.85, 'c2': 750.22, 'c3': 1053.79}
shared definition, as_of 2026-07-01: {'c1': 648.19, 'c2': 1097.97, 'c3': 788.79}Read the skew column
Three customers, three errors of -8.9%, -19.1% and +27.3%. Both definitions are defensible. Both were written by competent people. Neither is a bug you could find in a code review, because each file is correct on its own.
The fix is the avg_amount_30d function. One definition, with as_of as an argument. Training calls it with the historical decision date. Serving calls it with now. There is no second implementation to drift.
The second printed line is the point-in-time property: the same function, the same customer, a different as_of, and correspondingly different values. c2 was at 1097.97 on 1 July and 750.22 on 20 August. A training row for a July decision must use 1097.97, and a store that cannot reproduce that number will leak the future into your training set.
What the skew costs a model
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(5)
days = pd.date_range("2026-06-01", "2026-08-19", freq="D")
cust = [f"c{i:03d}" for i in range(400)]
level = dict(zip(cust, rng.gamma(2.0, 400.0, len(cust))))
rows = [(cid, d, float(rng.gamma(2.0, level[cid] / 2)))
for cid in cust for d in days if rng.random() < 0.35]
tx = pd.DataFrame(rows, columns=["customer", "ts", "amount"])
now = pd.Timestamp("2026-08-20")
whole_history = tx.groupby("customer")["amount"].mean()
last_30_days = tx[tx["ts"] > now - pd.Timedelta(days=30)].groupby("customer")["amount"].mean()
feat = pd.concat([whole_history.rename("training"), last_30_days.rename("serving")], axis=1).dropna()
churn = rng.binomial(1, 1 / (1 + np.exp(-(-0.004 * feat["training"] + 1.6))))
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(feat[["training"]], churn)
p_offline = model.predict_proba(feat[["training"]])[:, 1]
p_live = model.predict_proba(feat[["serving"]].rename(columns={"serving": "training"}))[:, 1]
print("customers scored :", len(feat))
print("median feature, training path :", round(feat["training"].median(), 1))
print("median feature, serving path :", round(feat["serving"].median(), 1))
print("mean absolute change in the predicted probability:", round(np.abs(p_offline - p_live).mean(), 3))
print("customers whose decision flips at a 0.5 cut-off :", int(((p_offline >= 0.5) != (p_live >= 0.5)).sum()))
print("offline AUC (training path) :", round(roc_auc_score(churn, p_offline), 3))
print("live AUC (serving path) :", round(roc_auc_score(churn, p_live), 3))customers scored : 400 median feature, training path : 685.3 median feature, serving path : 668.9 mean absolute change in the predicted probability: 0.043 customers whose decision flips at a 0.5 cut-off : 25 offline AUC (training path) : 0.813 live AUC (serving path) : 0.802
This is why skew survives so long
The model did not break. AUC fell from 0.813 to 0.802 — about one point. No alarm fires for that.
But 25 customers out of 400 get the opposite decision. That is 6% of your users receiving a different outcome because of a definition mismatch nobody documented.
Skew is a quiet tax, not a crash. It does not show up in error logs and it does not appear as a spike on a dashboard. It appears as a model that never quite performs the way the offline evaluation promised, and it is usually blamed on the model.
How to catch it without a feature store
Log the feature vector your production system computed, alongside every prediction. Then, once a day, recompute the same features offline for those same requests and compare.
import pandas as pd
# Both frames come from your own logs: what serving computed, and what you recomputed offline.
logged_at_serving = pd.DataFrame({
"request_id": ["r1", "r2", "r3", "r4"],
"avg_30d": [666.85, 750.22, 1053.79, 402.10],
"orders_7d": [3, 1, 5, 2],
})
recomputed_offline = pd.DataFrame({
"request_id": ["r1", "r2", "r3", "r4"],
"avg_30d": [666.85, 750.22, 1053.79, 388.44], # r4 disagrees
"orders_7d": [3, 1, 5, 2],
})
both = logged_at_serving.merge(recomputed_offline, on="request_id", suffixes=("_serve", "_offline"))
for col in ["avg_30d", "orders_7d"]:
gap = (both[f"{col}_serve"] - both[f"{col}_offline"]).abs()
print(f"{col:12s} max gap {gap.max():8.2f} rows disagreeing: {int((gap > 1e-9).sum())}")
if gap.max() > 1e-9:
print(" offenders:", list(both.loc[gap > 1e-9, "request_id"]))avg_30d max gap 13.66 rows disagreeing: 1
offenders: ['r4']
orders_7d max gap 0.00 rows disagreeing: 0orders_7d agrees exactly, so that feature has one definition. avg_30d disagrees on one request out of four, by 13.66. That is skew, and printing the offending request_id is what lets you go and read the two code paths side by side.
Run this on a daily sample of real traffic. If the maximum gap on any column is not near zero, you have skew. This one check catches more production ML problems than any monitoring dashboard, and it takes an afternoon to build. Google's Rules of Machine Learning makes the same recommendation, from the same experience.
Common mistakes
Two implementations of one feature. SQL for training, Python for serving. This is the default state of most teams and the root of the whole problem.
Computing a feature over the whole table for training. The leak in the churn example: an average over a customer's entire history includes days after the decision point. Always filter by as_of.
Fitting a scaler or an encoder on all rows. Same class of leak. The fitted parameters belong to the training window and must be saved, not recomputed at serving time.
Backfilling with today's code and yesterday's date. If the definition changed last month, recomputing history with the new code produces training data that never existed. Version the feature definitions the way you version data.
Installing a feature store to fix a discipline problem. If two teams still write two definitions, they will write them in the feature store. The tool enforces the pattern; it does not create the agreement.
Try it yourself
In skew.py, change the serving window from 30 days to 60. Watch the skew shrink but not vanish, because the training path still uses the entire history.
Then break point-in-time correctness deliberately: change events["ts"] <= as_of to events["ts"] <= as_of + pd.Timedelta(days=30). The function now reads a month into the future. Any model trained on it will look excellent and be worthless.
What to learn next
- ETL for machine learning — the pipeline that materialises these features.
- Monitoring and model drift — watching for skew after launch.
- Feature engineering — deciding what the features should be in the first place.
Researcher — Mathematics and papers.
The point-in-time join, stated precisely
Given an event stream $E = {(k, t, v)}$ of keys, timestamps and values, and a set of labelled decision points ${(k_i, t_i, y_i)}$, the correct training feature is:
$$ x_i \;=\; \phi\big({\,e \in E \;:\; e.k = k_i \;\wedge\; e.t \le t_i \,}\big) $$
Where $\phi$ is the aggregation. The naive alternative $\phi({e : e.k = k_i})$ over all time is a temporal leak, and it is undetectable by any check on the data alone — the values are individually correct, and only the join is wrong.
Implemented in SQL this is an AS OF join: for each entity and label timestamp, select the feature row with the greatest timestamp not exceeding it. Naively it is $O(|E| \cdot |L|)$. Sorting both sides by $(k, t)$ and merging reduces it to $O\big((|E| + |L|)\log(|E| + |L|)\big)$, which is why every feature store sorts before joining, and why pandas.merge_asof exists.
The bitemporal subtlety
A single timestamp is not enough when facts arrive late or are corrected. Two time axes are needed:
- Valid time $t_v$: when the fact was true in the world.
- Transaction time $t_x$: when your system learned it.
The correct training feature uses what was knowable, so the filter is $t_x \le t_i$, not $t_v \le t_i$. A payment that occurred at 10:00 but reached your database at 10:40 was not available for a 10:05 decision, and a training set that uses it overstates what production can do.
Most feature stores model valid time only. Systems where late arrival is common — payments, IoT, anything mobile — need both, and getting this wrong produces a model that is optimistic by exactly the amount of the arrival delay. See streaming data.
Architecture
The standard design is a dual store fed by one definition:
feature definition (one file)
/ \
batch materialisation streaming path
| |
offline store online store
(Parquet / warehouse) (Redis / DynamoDB)
| |
point-in-time join key lookup
| |
training set live inferenceRequirements differ sharply between the two sides:
| Offline store | Online store | |
|---|---|---|
| Access | Scan, joined by time | Point lookup by key |
| Latency | Minutes acceptable | Single-digit milliseconds |
| Volume | Full history | Latest value per key |
| Format | Columnar (Parquet) | Key-value |
Consistency between them is the hard engineering problem, and the reason feature stores exist as products rather than as libraries.
Lambda versus streaming-only
The lambda architecture runs a batch path and a streaming path over the same definition, reconciling them. It guarantees correctness through the batch layer while the streaming layer provides freshness. The cost is two implementations, which is the very problem being solved — the definition is shared, but the execution engines differ and diverge in edge cases.
The kappa architecture (Kreps, 2014) keeps only the streaming path and reprocesses from a retained log to correct history. Simpler in principle, and it requires a log with sufficient retention and deterministic, replayable transformations.
Modern systems increasingly compile one declarative feature definition into both engines, which is the honest solution to a genuinely hard problem rather than a way around it.
Detecting skew in production
Breck et al. (2019), Data Validation for Machine Learning, SysML, describe the TFX approach: log serving feature vectors, sample them, and compare distributions against the training data. They distinguish three failure classes worth naming separately:
- Feature skew — the computed value differs between paths, as demonstrated above.
- Distribution skew — the same code, different input distribution. This is genuine drift and needs a different response.
- Scoring/serving skew — the model sees a different subset than training assumed, typically because production filters candidates before scoring.
For continuous features, compare quantiles rather than means; the L-infinity distance between empirical CDFs is a robust default. For categorical features, watch for new categories, since an unseen category at serving time is a hard failure in most encoders. Chebyshev distance over category frequencies is what TFDV uses.
Cost model for the online store
Per prediction request with $f$ features across $g$ entity keys:
- Network round trips: $g$ with batched multi-get, $f$ without. Batching by entity is the single largest latency win.
- Storage: $O(\text{keys} \times f)$ for latest-value-only.
- Freshness lag: the streaming pipeline's end-to-end latency, which must be measured and alerted on, not assumed.
A p99 budget of 10 ms for feature retrieval is typical, and it constrains $f$ and $g$ more than most teams expect when designing the feature set.
When not to build one
The honest guidance, which vendor material omits: a feature store pays for itself when features are shared across models and teams, or when online and offline paths genuinely differ. With one model, one team and batch scoring, a versioned transformation module invoked by both paths delivers the same correctness at a fraction of the operational cost.
The pattern — one definition, as_of as a parameter, both paths calling it — is the part that matters. The database behind it is an implementation detail you should adopt only when scale demands it.
Reading
- Zinkevich, Rules of Machine Learning: Best Practices for ML Engineering, Google — rules 29 to 32 cover training/serving skew directly.
- Breck et al., Data Validation for Machine Learning, SysML 2019.
- Kleppmann, Designing Data-Intensive Applications, O'Reilly 2017 — chapter 11 on streams, batch and reprocessing.
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — on undeclared consumers and data dependencies.
- Feast documentation, Point-in-time correctness — the clearest worked explanation of the AS OF join in an ML context.
What to learn next
- ETL for machine learning — the pipeline that materialises these features.
- Monitoring and model drift — watching for skew after launch.
- Feature engineering — deciding what the features should be in the first place.