Feature engineering
Feature engineering is reshaping your columns into the form a model can learn from, and it routinely buys more accuracy than any change of algorithm.
- 23 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Feature engineering is reshaping the data you already have into a form the model can actually learn from.
The directions you could not use
Someone stops you on the road and asks how to reach your house. You could answer with four numbers: your latitude and longitude, and theirs. Every fact they need is in those four numbers.
They will still be lost.
Say instead "two kilometres straight down this road, third left after the petrol pump" and they arrive in ten minutes. Not one new fact was added. The same information was put into a form a person can use.
A model is in exactly the position of that stranger. It gets columns of numbers with no idea what they mean. The form you hand them over in decides whether it can learn anything at all.
Why this is the least glamorous skill and the most valuable one
Every tutorial spends its time on models. Pick this algorithm, tune that setting, try a bigger network.
In real work the ranking is close to the opposite. The columns you build usually matter more than the model you pick. In the Developer section you will watch a plain, boring linear regression beat a modern boosted-tree model. Its only advantage is one good extra column.
That result is not a trick. It is the normal state of affairs on tabular data. It is also why experienced people spend most of their time here, rather than on model selection.
The reason is straightforward. A model can only combine your columns in the ways its shape allows. A linear model can add columns together and scale them. It cannot square them, or divide one by another, or measure the distance between two points. If the answer depends on something the model cannot build, the model cannot find it. More data will not help.
You build that thing yourself, hand it over as a new column, and the problem becomes easy.
How it works
what you were given what you hand the model
pickup_lat 19.0760 trip_distance_km 4.2
pickup_lon 72.8777 ------> is_rush_hour yes
drop_lat 19.1136 you do day_of_week Tuesday
drop_lon 72.8697 this bit crosses_river no
timestamp 2026-08-20 18:42
four numbers no model can use four facts any model can useThe arrow in the middle is the job. There is no library that does it for you, because it needs something no library has: knowledge of what your data means.
Somebody who has driven in that city knows a river crossing adds twenty minutes. That is not in the coordinates in any form a model will discover on its own from three hundred rides.
Where you have already seen it
- Cab and food apps predicting arrival time. Coordinates go in, distance and traffic-hour come out.
- Spam filters counting how many links a message has, not reading it as one long string.
- Credit scoring using "how much of your limit you have used", built from balance divided by limit.
- Cricket score predictors using run rate and wickets in hand, not the raw ball-by-ball log.
- Shopping recommendations using "days since your last order", built from a date.
Every one of these is a number a person invented because they understood the problem.
What is honestly hard here
There is no checklist that works everywhere. Good features come from understanding what your data means. That understanding takes time with the people who generated the data, and there is no shortcut around it.
It is easy to cheat yourself by accident. Some columns look brilliant in testing and are useless in production. They secretly contain the answer. A column called refund_issued predicts order_cancelled beautifully, and is unavailable at the moment you actually need the prediction. This is called leakage, and train, test and validation splits covers how to catch it.
Read this twice. The work is slow, unglamorous and mostly conversation. It is also the part of the job that decides whether the project works. Nobody puts it in a demo video because there is nothing to show.
Remember this
- A model can only use the columns you give it, in the ways its own shape allows.
- Turning raw columns into meaningful ones often beats any change of algorithm.
- The good ideas come from understanding the data, not from a library.
What to learn next
- XGBoost — the model that gets the most out of well-built tabular columns.
- Train, test and validation splits — where feature code must sit so it does not leak.
- Model evaluation — the numbers that tell you a new column actually helped.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyEverything runs on CPU in a couple of seconds. All data is generated inline.
The experiment that makes the case
Three hundred auto rides on a city grid. We know the truth because we wrote it: the time a ride takes depends on how far it went. The model is given the four raw coordinates and must work that out.
Then we add one column — the distance — and run everything again.
import numpy as np
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
N = 300
# 300 auto rides on a 10 km by 10 km city grid.
pickup_x = rng.uniform(0, 10, N)
pickup_y = rng.uniform(0, 10, N)
drop_x = rng.uniform(0, 10, N)
drop_y = rng.uniform(0, 10, N)
# The truth we are trying to recover: time depends on how far the ride went.
km = np.sqrt((drop_x - pickup_x) ** 2 + (drop_y - pickup_y) ** 2)
minutes = 4.0 + 3.2 * km + rng.normal(0, 1.5, N)
raw = np.column_stack([pickup_x, pickup_y, drop_x, drop_y])
with_distance = np.column_stack([pickup_x, pickup_y, drop_x, drop_y, km])
for name, X in [("4 raw coordinates", raw), ("+ trip distance ", with_distance)]:
X_train, X_test, y_train, y_test = train_test_split(
X, minutes, test_size=0.3, random_state=0)
for model_name, model in [("linear regression", LinearRegression()),
("boosted trees ", HistGradientBoostingRegressor(random_state=0))]:
model.fit(X_train, y_train)
pred = model.predict(X_test)
print(f"{name} {model_name} R2 {r2_score(y_test, pred):6.3f}"
f" typical error {mean_absolute_error(y_test, pred):5.2f} min")
print()4 raw coordinates linear regression R2 -0.007 typical error 6.48 min 4 raw coordinates boosted trees R2 0.752 typical error 2.94 min + trip distance linear regression R2 0.962 typical error 1.29 min + trip distance boosted trees R2 0.949 typical error 1.48 min
Four numbers, and every one of them matters. R2 is the share of the variation the model explained; 0 means it did no better than always guessing the average.
Linear regression on raw coordinates: R2 = -0.007, off by 6.48 minutes. A negative score means it was slightly worse than guessing the average every time. The model was handed every fact it needed and learned nothing at all.
That is not a weak model. It is a model being asked to do something its shape forbids. Distance needs subtraction, then squaring, then a square root. A linear model can only scale each column and add. The answer was unreachable.
Boosted trees on raw coordinates: R2 = 0.752, off by 2.94 minutes. A powerful model partly rescued the situation. Trees carve the four-dimensional space into boxes and can crudely approximate distance with enough of them. It took a lot of machinery to half-solve a problem that one line of arithmetic solves.
Linear regression with the distance column: R2 = 0.962, off by 1.29 minutes. A five-fold reduction in error, from a model with five coefficients.
Now compare the last two lines against each other. The simple model with the good feature (1.29 min) beat the powerful model with the good feature (1.48 min), and beat the powerful model without it by more than a factor of two.
One column was worth more than the entire difference between the two algorithms. If your instinct on a stuck project is to try a bigger model, this table is the argument for spending that hour on your columns instead.
The mistake everyone makes with times and angles
Hours of the day, days of the week, compass bearings and months all wrap around. Written as plain numbers they lie about which values are close together.
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(1)
N = 500
hour = rng.integers(0, 24, N).astype(float)
# Orders peak around 8pm and bottom out around 8am. A smooth daily cycle.
orders = 60 + 40 * np.cos((hour - 20) / 24 * 2 * np.pi) + rng.normal(0, 4, N)
as_number = hour.reshape(-1, 1)
angle = hour / 24 * 2 * np.pi
as_cycle = np.column_stack([np.sin(angle), np.cos(angle)])
for name, X in [("hour as a plain number ", as_number),
("hour as sine and cosine", as_cycle)]:
X_train, X_test, y_train, y_test = train_test_split(
X, orders, test_size=0.3, random_state=0)
model = Ridge().fit(X_train, y_train)
err = mean_absolute_error(y_test, model.predict(X_test))
print(f"{name} average error {err:5.2f} orders")
print()
def gap(h1, h2):
a, b = as_cycle[hour == h1][0], as_cycle[hour == h2][0]
return float(np.hypot(*(a - b)))
print("how far apart are two clock times?")
print(f" 11pm and midnight, as plain numbers : {abs(23 - 0)}")
print(f" 11pm and midnight, as sine and cosine: {gap(23, 0):.3f}")
print(f" midnight and 1am, as sine and cosine: {gap(0, 1):.3f}")hour as a plain number average error 19.25 orders hour as sine and cosine average error 3.23 orders how far apart are two clock times? 11pm and midnight, as plain numbers : 23 11pm and midnight, as sine and cosine: 0.261 midnight and 1am, as sine and cosine: 0.261
The bottom three lines explain the top two.
Written as a plain number, 11pm and midnight are 23 apart — the largest possible gap in the column. In reality they are sixty minutes apart. Every model that treats the column as a quantity inherits that lie.
Place the hour on a circle instead, using its sine and cosine, and 11pm sits 0.261 from midnight. Midnight sits 0.261 from 1am. The two gaps are identical, which is exactly true of a clock.
The error dropped from 19.25 orders to 3.23. Two columns replaced one, and nothing else changed.
Use the same trick for day of week, day of year, month and compass bearing. The pattern is always sin(2π · value / period) and cos(2π · value / period), and you always need both — sine alone cannot tell 3am from 9pm.
Categories with too many values, and the trap inside them
A column such as pin code or product ID can have thousands of distinct values. One-hot encoding turns it into thousands of columns, most of them nearly empty.
Target encoding replaces each category with the average target for that category. Pin code 400058 becomes 0.31 because 31 percent of loans there defaulted. One column instead of two thousand.
It is a genuinely useful technique and it is also a trap. Here is a pin code column deliberately built to be pure noise, unrelated to the target in every way.
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.metrics import r2_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import TargetEncoder
rng = np.random.default_rng(2)
N = 800
# A pin code column with 200 levels that has NOTHING to do with the target.
pin = rng.integers(0, 200, N).reshape(-1, 1)
y = rng.normal(0, 1, N)
pin_tr, pin_te, y_tr, y_te = train_test_split(pin, y, test_size=0.3, random_state=0)
# WRONG: average the target per pin code over every training row, then reuse it.
means = {code: y_tr[pin_tr.ravel() == code].mean() for code in np.unique(pin_tr)}
overall = y_tr.mean()
leaky_tr = np.array([[means.get(c, overall)] for c in pin_tr.ravel()])
leaky_te = np.array([[means.get(c, overall)] for c in pin_te.ravel()])
model = Ridge().fit(leaky_tr, y_tr)
print("hand-rolled target encoding")
print(f" train R2 {r2_score(y_tr, model.predict(leaky_tr)):6.3f}")
print(f" test R2 {r2_score(y_te, model.predict(leaky_te)):6.3f}")
# RIGHT: TargetEncoder builds the training encoding with internal cross-fitting,
# so no row ever contributes to its own encoded value.
enc = TargetEncoder(random_state=0)
safe_tr = enc.fit_transform(pin_tr, y_tr)
safe_te = enc.transform(pin_te)
model2 = Ridge().fit(safe_tr, y_tr)
print("scikit-learn TargetEncoder")
print(f" train R2 {r2_score(y_tr, model2.predict(safe_tr)):6.3f}")
print(f" test R2 {r2_score(y_te, model2.predict(safe_te)):6.3f}")hand-rolled target encoding train R2 0.340 test R2 -0.465 scikit-learn TargetEncoder train R2 0.001 test R2 -0.005
The hand-rolled version reports R2 = 0.340 on training data, on a column containing no information whatsoever. It looks like a discovery. Someone will put it in a slide.
On unseen data it scores -0.465, which is substantially worse than predicting the average. The feature did not fail quietly. It actively damaged the model.
The cause is small and easy to miss. With 200 pin codes and 560 training rows, most codes appear two or three times. The average for a code is computed from rows that include the row being encoded. So each row's feature value partly is its own target. The model learns to read the answer off the feature, and at prediction time the answer is not there.
TargetEncoder reports 0.001 and -0.005. It builds the training encoding with internal cross-fitting, so a row is always encoded using other rows. It correctly says the column is worthless, which is the answer you needed.
Never hand-roll target encoding. Use sklearn.preprocessing.TargetEncoder, or CatBoost's ordered target statistics, both of which handle this properly.
Where feature code belongs
Feature steps must live inside the model object, not in a script that runs before it. Anything computed outside is computed once on all your data, which quietly mixes the test set into training.
Arithmetic done row by row is safe either way. Anything that looks at other rows — a mean, a scale factor, a category encoding — is not.
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import FunctionTransformer, StandardScaler
rng = np.random.default_rng(0)
N = 300
X = rng.uniform(0, 10, size=(N, 4)) # pickup x, y and drop x, y
km = np.hypot(X[:, 2] - X[:, 0], X[:, 3] - X[:, 1])
y = 4.0 + 3.2 * km + rng.normal(0, 1.5, N)
def add_distance(a):
"""Row-by-row arithmetic only. Nothing here looks at other rows."""
d = np.hypot(a[:, 2] - a[:, 0], a[:, 3] - a[:, 1])
return np.column_stack([a, d])
plain = make_pipeline(StandardScaler(), Ridge())
engineered = make_pipeline(FunctionTransformer(add_distance), StandardScaler(), Ridge())
for name, pipe in [("raw coordinates only", plain), ("with the distance step", engineered)]:
scores = cross_val_score(pipe, X, y, cv=5, scoring="neg_mean_absolute_error")
print(f"{name:24s} cross-validated error {-scores.mean():5.2f} min "
f"(spread {scores.std():.2f})")raw coordinates only cross-validated error 6.33 min (spread 0.51) with the distance step cross-validated error 1.26 min (spread 0.07)
The error result matches the earlier experiment, now measured across five folds rather than one split.
Look at the spread as well. It fell from 0.51 to 0.07. The engineered model is not only better on average, it is far more consistent from fold to fold. A model that has been handed the right shape stops depending on which rows it happened to see.
StandardScaler sits inside the pipeline, so its mean and standard deviation are computed on each fold's training rows alone. Fit it outside and every fold's "held-out" data has already influenced the scaling. Train, test and validation splits measures how large that error can get.
A working checklist
Not a recipe. A set of questions worth asking about every column you have.
| If you have | Consider building |
|---|---|
| Two positions | Distance, bearing, whether they are in the same zone |
| A timestamp | Hour and weekday as sine/cosine, is-weekend, is-holiday, days since a known event |
| Two related amounts | Their ratio — utilisation, conversion rate, price per square foot |
| A skewed amount | log1p of it, so the tail stops dominating |
| A category with few levels | One-hot encoding |
| A category with many levels | TargetEncoder, or frequency of occurrence |
| A history of events | Count, sum, mean and recency over the last 7, 30 and 90 days |
| Free text | Length, word count, digit count, link count, then embeddings |
| Anything with gaps | A was_missing flag alongside the filled value |
Two general notes. Ratios are consistently the highest-value construction on business data, because they remove size and leave behavior. And a was_missing flag costs one column and often carries real signal, as the XGBoost lesson demonstrates with loan data.
Common mistakes
Scaling or encoding before splitting. The most common leak in existence. Everything that learns from data goes inside the pipeline.
Building a feature from information that arrives later. delivery_rating predicts was_delivered perfectly and does not exist when you need to predict. Ask of every column: would this value be available, at this exact moment, in production?
Adding hundreds of features because you can. Each one is another chance to fit noise, and irrelevant columns dilute the useful ones. Add a feature because you can explain why it should matter.
Hand-rolling target encoding. Demonstrated above. Use the library version.
Believing importance scores about your new feature. Impurity-based importances are biased towards high-cardinality columns, as the random forest lesson measures. Judge a new feature by whether held-out error improved, and nothing else.
Different feature code in training and in production. Two implementations of days_since_last_order will drift apart, and the model degrades without any error appearing anywhere. One implementation, called from both places.
Try it yourself
Return to one_column.py and make the world less tidy. Replace the straight-line distance with a grid distance, np.abs(drop_x - pickup_x) + np.abs(drop_y - pickup_y), which is closer to how roads actually work. Keep giving the model the straight-line distance as its feature.
Now the feature is a good approximation rather than the exact truth. Measure how much of the improvement survives. Real features are never exactly right, and knowing what a nearly-right feature is worth is the useful skill.
Then delete the distance column and add drop_x - pickup_x and drop_y - pickup_y instead. Predict first whether linear regression can use these. Run it and find out. The answer explains precisely which shapes a linear model can and cannot reach.
What to learn next
- XGBoost — the model that gets the most out of well-built tabular columns.
- Train, test and validation splits — where feature code must sit so it does not leak.
- Model evaluation — the numbers that tell you a new column actually helped.
Researcher — Mathematics and papers.
What a feature map does to the hypothesis class
A learning algorithm searches a hypothesis class H. Feature engineering replaces the input x ∈ X with φ(x) ∈ R^m and searches H ∘ φ instead. The algorithm is unchanged; the reachable set of functions is not.
For a linear model, H = { x ↦ w^T x + b } is closed under scaling and addition of coordinates and under nothing else. The coordinate experiment in the Developer section makes the consequence concrete: sqrt((x3−x1)^2 + (x4−x2)^2) lies outside the span of {x1, x2, x3, x4, 1}, so the target is unreachable at any sample size. The measured R² = −0.007 is not an estimation failure. It is an approximation failure, and no amount of data repairs it.
This is the classical decomposition. Total error splits into approximation error — the gap between the best function in H ∘ φ and the truth — and estimation error, the gap between what you fit and that best function. Model selection and regularisation address estimation error. Feature engineering addresses approximation error, and it is the only tool that does.
The kernel view, and why explicit features persist
The kernel trick computes ⟨φ(x), φ(x')⟩ without materialising φ, permitting infinite-dimensional feature spaces at O(n^2) cost in the sample size. KernelRidge with an RBF kernel would solve the coordinate problem without a hand-built feature.
Explicit feature construction survives for reasons that are practical rather than theoretical:
- Kernel methods scale as
O(n^2)toO(n^3)and do not reach the sample sizes tabular problems now carry. Random Fourier features (Rahimi & Recht, 2007) approximate the map explicitly inO(n m)and are the standard compromise. - An explicit feature is auditable, monitorable and explainable to a regulator. A kernel expansion is not.
- Hand-built features encode domain constraints that no generic kernel contains. That a river crossing adds twenty minutes is not recoverable from three hundred coordinate pairs by any kernel.
Trees, rotation, and why they need features too
Axis-aligned recursive partitioning approximates any measurable function given unlimited depth and data, so trees are formally more expressive here than linear models. The Developer section's R² = 0.752 for boosted trees on raw coordinates reflects that.
The gap to 0.949 with the distance feature is a statement about sample efficiency. Representing a rotated or curved decision boundary with axis-aligned cuts requires a staircase whose step count grows with the required precision, and each step consumes data. Grinsztajn, Oyallon & Varoquaux (2022) identify non-rotation-invariance as one of the properties that makes tree ensembles strong on tabular data, and it is the same property that makes them inefficient on rotated targets. Oblique trees and models with learned rotations trade one bias for the other.
The practical rule that follows: trees need monotone transformations far less than linear models do — log, sqrt and any strictly increasing map leave the split structure invariant — and need constructed interactions and distances nearly as much.
Cyclical, periodic and spline encodings
For a value v with period p, the standard encoding is
φ(v) = ( sin(2π v / p) , cos(2π v / p) )This embeds the cyclic group into R^2 so that Euclidean distance in the embedding is a monotone function of angular separation — verified numerically in the Developer section, where d(23:00, 00:00) = d(00:00, 01:00) = 0.261. One term is insufficient: sin alone is not injective on [0, p).
The encoding imposes a single sinusoid. When the daily pattern has two peaks, add harmonics sin(2πkv/p), cos(2πkv/p) for k = 1, 2, …, which is a truncated Fourier basis and a design choice about smoothness.
The alternative is a periodic spline basis: sklearn.preprocessing.SplineTransformer(extrapolation="periodic"). Splines are local, so a lunchtime spike does not distort the evening, whereas a Fourier basis is global and will. For sharply peaked periodic effects, splines dominate.
Target encoding, formally
Replace category c with an estimate of E[Y | C = c]. The naive estimator is the within-category mean, and it is biased when the same rows contribute to the estimate and the fit.
With n_c rows in category c and a global mean μ, empirical-Bayes smoothing shrinks small categories towards the prior:
enc(c) = ( n_c · ȳ_c + m · μ ) / ( n_c + m )m is a smoothing weight, interpretable as a count of pseudo-observations from the prior. Micci-Barreca (2001) gives the original formulation with a sigmoid shrinkage schedule.
Smoothing alone does not remove leakage. The Developer section measures the residual: a pure-noise column with 200 levels over 560 rows yields train R² = 0.340 and test R² = −0.465. The row's own target is inside its own feature.
Two correct constructions exist.
Cross-fitted (K-fold) encoding. Partition the training rows; encode each fold using statistics from the other folds. This is what sklearn.preprocessing.TargetEncoder does internally, and it is why its fit_transform and transform return different values for the same input — an intentional asymmetry that surprises people the first time.
Ordered target statistics. Fix a random permutation and encode each row using only rows preceding it. Prokhorenkova et al. (2018) show that K-fold encoding still admits a residual prediction shift, and that ordering removes it. CatBoost extends the same permutation to the gradient estimates.
The general principle covers more than encoding: any statistic computed from y and attached to x must be computed out-of-fold. The same argument applies to feature selection, as demonstrated in train, test and validation splits.
Automated feature construction
Deep Feature Synthesis (Kanter & Veeramachaneni, 2015) generates features over relational schemas by stacking aggregation primitives (SUM, MEAN, COUNT across a one-to-many relation) and transform primitives (day-of-week, time-since) up to a depth limit. Featuretools implements it. The output space is combinatorial in depth, so selection afterwards is mandatory.
It works well for the aggregation-heavy features that relational business data invites, and it does not produce the coordinate-to-distance construction unless a distance primitive is supplied. That is the general boundary: automated systems compose primitives you gave them. Choosing the primitives is the domain work, relocated rather than removed.
Genetic programming approaches such as AutoFeat and symbolic regression search a wider space at substantially higher cost, with a strong tendency to produce expressions nobody can defend in review.
Training and serving skew
The dominant production failure mode is not a weak feature. It is a feature computed differently at training time and at serving time.
Two structural causes recur. The first is point-in-time correctness: a training feature computed from a table as it looks today includes information that did not exist at the timestamp of the row. Any aggregate over a customer's history must be truncated at the row's event time, which requires either an append-only event log or explicit as-of joins. The second is dual implementation: batch feature code in SQL or Python, serving feature code in application code. They diverge on edge cases and nothing raises an error.
Feature stores (Feast, Tecton, and the equivalents inside major platforms) exist to address exactly these two problems: one definition, materialised for both offline training and online serving, with as-of semantics enforced by the join. The engineering is unglamorous and the failure it prevents is silent, which is a fair summary of this whole subject.
Key references
- Micci-Barreca, D. (2001). A Preprocessing Scheme for High-Cardinality Categorical Attributes in Classification and Prediction Problems. SIGKDD Explorations 3(1).
- Rahimi, A. & Recht, B. (2007). Random Features for Large-Scale Kernel Machines. NeurIPS.
- Guyon, I. & Elisseeff, A. (2003). An Introduction to Variable and Feature Selection. JMLR 3.
- Kanter, J. & Veeramachaneni, K. (2015). Deep Feature Synthesis: Towards Automating Data Science Endeavors. IEEE DSAA.
- Kaufman, S., Rosset, S. & Perlich, C. (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM TKDD 6(4).
- Prokhorenkova, L. et al. (2018). CatBoost: Unbiased Boosting with Categorical Features. NeurIPS.
- Sculley, D. et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS.
- Grinsztajn, L., Oyallon, E. & Varoquaux, G. (2022). Why Do Tree-Based Models Still Outperform Deep Learning on Tabular Data? NeurIPS.
- Kapoor, S. & Narayanan, A. (2023). Leakage and the Reproducibility Crisis in ML-Based Science. Patterns 4(9).
What to learn next
- XGBoost — the model that gets the most out of well-built tabular columns.
- Train, test and validation splits — where feature code must sit so it does not leak.
- Model evaluation — the numbers that tell you a new column actually helped.