Machine Learning

Feature engineering

Feature engineering is reshaping your columns into the form a model can learn from, and it routinely buys more accuracy than any change of algorithm.

On this page 8
  1. The short answer
  2. The directions you could not use
  3. Why this is the least glamorous skill and the most valuable one
  4. How it works
  5. Where you have already seen it
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Feature engineering is reshaping the data you already have into a form the model can actually learn from.

The directions you could not use

Someone stops you on the road and asks how to reach your house. You could answer with four numbers: your latitude and longitude, and theirs. Every fact they need is in those four numbers.

They will still be lost.

Say instead "two kilometres straight down this road, third left after the petrol pump" and they arrive in ten minutes. Not one new fact was added. The same information was put into a form a person can use.

A model is in exactly the position of that stranger. It gets columns of numbers with no idea what they mean. The form you hand them over in decides whether it can learn anything at all.

Why this is the least glamorous skill and the most valuable one

Every tutorial spends its time on models. Pick this algorithm, tune that setting, try a bigger network.

In real work the ranking is close to the opposite. The columns you build usually matter more than the model you pick. In the Developer section you will watch a plain, boring linear regression beat a modern boosted-tree model. Its only advantage is one good extra column.

That result is not a trick. It is the normal state of affairs on tabular data. It is also why experienced people spend most of their time here, rather than on model selection.

The reason is straightforward. A model can only combine your columns in the ways its shape allows. A linear model can add columns together and scale them. It cannot square them, or divide one by another, or measure the distance between two points. If the answer depends on something the model cannot build, the model cannot find it. More data will not help.

You build that thing yourself, hand it over as a new column, and the problem becomes easy.

How it works

   what you were given                    what you hand the model

   pickup_lat   19.0760                   trip_distance_km   4.2
   pickup_lon   72.8777        ------>    is_rush_hour       yes
   drop_lat     19.1136        you do     day_of_week        Tuesday
   drop_lon     72.8697        this bit   crosses_river      no
   timestamp    2026-08-20 18:42

   four numbers no model can use          four facts any model can use

The arrow in the middle is the job. There is no library that does it for you, because it needs something no library has: knowledge of what your data means.

Somebody who has driven in that city knows a river crossing adds twenty minutes. That is not in the coordinates in any form a model will discover on its own from three hundred rides.

Where you have already seen it

  • Cab and food apps predicting arrival time. Coordinates go in, distance and traffic-hour come out.
  • Spam filters counting how many links a message has, not reading it as one long string.
  • Credit scoring using "how much of your limit you have used", built from balance divided by limit.
  • Cricket score predictors using run rate and wickets in hand, not the raw ball-by-ball log.
  • Shopping recommendations using "days since your last order", built from a date.

Every one of these is a number a person invented because they understood the problem.

What is honestly hard here

There is no checklist that works everywhere. Good features come from understanding what your data means. That understanding takes time with the people who generated the data, and there is no shortcut around it.

It is easy to cheat yourself by accident. Some columns look brilliant in testing and are useless in production. They secretly contain the answer. A column called refund_issued predicts order_cancelled beautifully, and is unavailable at the moment you actually need the prediction. This is called leakage, and train, test and validation splits covers how to catch it.

Read this twice. The work is slow, unglamorous and mostly conversation. It is also the part of the job that decides whether the project works. Nobody puts it in a demo video because there is nothing to show.

Remember this

  • A model can only use the columns you give it, in the ways its own shape allows.
  • Turning raw columns into meaningful ones often beats any change of algorithm.
  • The good ideas come from understanding the data, not from a library.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Everything runs on CPU in a couple of seconds. All data is generated inline.

The experiment that makes the case

Three hundred auto rides on a city grid. We know the truth because we wrote it: the time a ride takes depends on how far it went. The model is given the four raw coordinates and must work that out.

Then we add one column — the distance — and run everything again.

one_column.py
import numpy as np
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
N = 300

# 300 auto rides on a 10 km by 10 km city grid.
pickup_x = rng.uniform(0, 10, N)
pickup_y = rng.uniform(0, 10, N)
drop_x = rng.uniform(0, 10, N)
drop_y = rng.uniform(0, 10, N)

# The truth we are trying to recover: time depends on how far the ride went.
km = np.sqrt((drop_x - pickup_x) ** 2 + (drop_y - pickup_y) ** 2)
minutes = 4.0 + 3.2 * km + rng.normal(0, 1.5, N)

raw = np.column_stack([pickup_x, pickup_y, drop_x, drop_y])
with_distance = np.column_stack([pickup_x, pickup_y, drop_x, drop_y, km])

for name, X in [("4 raw coordinates", raw), ("+ trip distance  ", with_distance)]:
    X_train, X_test, y_train, y_test = train_test_split(
        X, minutes, test_size=0.3, random_state=0)
    for model_name, model in [("linear regression", LinearRegression()),
                              ("boosted trees    ", HistGradientBoostingRegressor(random_state=0))]:
        model.fit(X_train, y_train)
        pred = model.predict(X_test)
        print(f"{name}   {model_name}   R2 {r2_score(y_test, pred):6.3f}"
              f"   typical error {mean_absolute_error(y_test, pred):5.2f} min")
    print()
Output
4 raw coordinates   linear regression   R2 -0.007   typical error  6.48 min
4 raw coordinates   boosted trees       R2  0.752   typical error  2.94 min

+ trip distance     linear regression   R2  0.962   typical error  1.29 min
+ trip distance     boosted trees       R2  0.949   typical error  1.48 min

Four numbers, and every one of them matters. R2 is the share of the variation the model explained; 0 means it did no better than always guessing the average.

Linear regression on raw coordinates: R2 = -0.007, off by 6.48 minutes. A negative score means it was slightly worse than guessing the average every time. The model was handed every fact it needed and learned nothing at all.

That is not a weak model. It is a model being asked to do something its shape forbids. Distance needs subtraction, then squaring, then a square root. A linear model can only scale each column and add. The answer was unreachable.

Boosted trees on raw coordinates: R2 = 0.752, off by 2.94 minutes. A powerful model partly rescued the situation. Trees carve the four-dimensional space into boxes and can crudely approximate distance with enough of them. It took a lot of machinery to half-solve a problem that one line of arithmetic solves.

Linear regression with the distance column: R2 = 0.962, off by 1.29 minutes. A five-fold reduction in error, from a model with five coefficients.

Now compare the last two lines against each other. The simple model with the good feature (1.29 min) beat the powerful model with the good feature (1.48 min), and beat the powerful model without it by more than a factor of two.

One column was worth more than the entire difference between the two algorithms. If your instinct on a stuck project is to try a bigger model, this table is the argument for spending that hour on your columns instead.

The mistake everyone makes with times and angles

Hours of the day, days of the week, compass bearings and months all wrap around. Written as plain numbers they lie about which values are close together.

clock_time.py
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(1)
N = 500

hour = rng.integers(0, 24, N).astype(float)
# Orders peak around 8pm and bottom out around 8am. A smooth daily cycle.
orders = 60 + 40 * np.cos((hour - 20) / 24 * 2 * np.pi) + rng.normal(0, 4, N)

as_number = hour.reshape(-1, 1)
angle = hour / 24 * 2 * np.pi
as_cycle = np.column_stack([np.sin(angle), np.cos(angle)])

for name, X in [("hour as a plain number ", as_number),
                ("hour as sine and cosine", as_cycle)]:
    X_train, X_test, y_train, y_test = train_test_split(
        X, orders, test_size=0.3, random_state=0)
    model = Ridge().fit(X_train, y_train)
    err = mean_absolute_error(y_test, model.predict(X_test))
    print(f"{name}   average error {err:5.2f} orders")

print()


def gap(h1, h2):
    a, b = as_cycle[hour == h1][0], as_cycle[hour == h2][0]
    return float(np.hypot(*(a - b)))


print("how far apart are two clock times?")
print(f"  11pm and midnight, as plain numbers  : {abs(23 - 0)}")
print(f"  11pm and midnight, as sine and cosine: {gap(23, 0):.3f}")
print(f"  midnight and 1am,  as sine and cosine: {gap(0, 1):.3f}")
Output
hour as a plain number    average error 19.25 orders
hour as sine and cosine   average error  3.23 orders

how far apart are two clock times?
  11pm and midnight, as plain numbers  : 23
  11pm and midnight, as sine and cosine: 0.261
  midnight and 1am,  as sine and cosine: 0.261

The bottom three lines explain the top two.

Written as a plain number, 11pm and midnight are 23 apart — the largest possible gap in the column. In reality they are sixty minutes apart. Every model that treats the column as a quantity inherits that lie.

Place the hour on a circle instead, using its sine and cosine, and 11pm sits 0.261 from midnight. Midnight sits 0.261 from 1am. The two gaps are identical, which is exactly true of a clock.

The error dropped from 19.25 orders to 3.23. Two columns replaced one, and nothing else changed.

Use the same trick for day of week, day of year, month and compass bearing. The pattern is always sin(2π · value / period) and cos(2π · value / period), and you always need both — sine alone cannot tell 3am from 9pm.

Categories with too many values, and the trap inside them

A column such as pin code or product ID can have thousands of distinct values. One-hot encoding turns it into thousands of columns, most of them nearly empty.

Target encoding replaces each category with the average target for that category. Pin code 400058 becomes 0.31 because 31 percent of loans there defaulted. One column instead of two thousand.

It is a genuinely useful technique and it is also a trap. Here is a pin code column deliberately built to be pure noise, unrelated to the target in every way.

target_encoding.py
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.metrics import r2_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import TargetEncoder

rng = np.random.default_rng(2)
N = 800

# A pin code column with 200 levels that has NOTHING to do with the target.
pin = rng.integers(0, 200, N).reshape(-1, 1)
y = rng.normal(0, 1, N)
pin_tr, pin_te, y_tr, y_te = train_test_split(pin, y, test_size=0.3, random_state=0)

# WRONG: average the target per pin code over every training row, then reuse it.
means = {code: y_tr[pin_tr.ravel() == code].mean() for code in np.unique(pin_tr)}
overall = y_tr.mean()
leaky_tr = np.array([[means.get(c, overall)] for c in pin_tr.ravel()])
leaky_te = np.array([[means.get(c, overall)] for c in pin_te.ravel()])

model = Ridge().fit(leaky_tr, y_tr)
print("hand-rolled target encoding")
print(f"   train R2 {r2_score(y_tr, model.predict(leaky_tr)):6.3f}")
print(f"   test  R2 {r2_score(y_te, model.predict(leaky_te)):6.3f}")

# RIGHT: TargetEncoder builds the training encoding with internal cross-fitting,
# so no row ever contributes to its own encoded value.
enc = TargetEncoder(random_state=0)
safe_tr = enc.fit_transform(pin_tr, y_tr)
safe_te = enc.transform(pin_te)
model2 = Ridge().fit(safe_tr, y_tr)
print("scikit-learn TargetEncoder")
print(f"   train R2 {r2_score(y_tr, model2.predict(safe_tr)):6.3f}")
print(f"   test  R2 {r2_score(y_te, model2.predict(safe_te)):6.3f}")
Output
hand-rolled target encoding
   train R2  0.340
   test  R2 -0.465
scikit-learn TargetEncoder
   train R2  0.001
   test  R2 -0.005

The hand-rolled version reports R2 = 0.340 on training data, on a column containing no information whatsoever. It looks like a discovery. Someone will put it in a slide.

On unseen data it scores -0.465, which is substantially worse than predicting the average. The feature did not fail quietly. It actively damaged the model.

The cause is small and easy to miss. With 200 pin codes and 560 training rows, most codes appear two or three times. The average for a code is computed from rows that include the row being encoded. So each row's feature value partly is its own target. The model learns to read the answer off the feature, and at prediction time the answer is not there.

TargetEncoder reports 0.001 and -0.005. It builds the training encoding with internal cross-fitting, so a row is always encoded using other rows. It correctly says the column is worthless, which is the answer you needed.

Never hand-roll target encoding. Use sklearn.preprocessing.TargetEncoder, or CatBoost's ordered target statistics, both of which handle this properly.

Where feature code belongs

Feature steps must live inside the model object, not in a script that runs before it. Anything computed outside is computed once on all your data, which quietly mixes the test set into training.

Arithmetic done row by row is safe either way. Anything that looks at other rows — a mean, a scale factor, a category encoding — is not.

inside_the_pipeline.py
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import FunctionTransformer, StandardScaler

rng = np.random.default_rng(0)
N = 300
X = rng.uniform(0, 10, size=(N, 4))                       # pickup x, y and drop x, y
km = np.hypot(X[:, 2] - X[:, 0], X[:, 3] - X[:, 1])
y = 4.0 + 3.2 * km + rng.normal(0, 1.5, N)


def add_distance(a):
    """Row-by-row arithmetic only. Nothing here looks at other rows."""
    d = np.hypot(a[:, 2] - a[:, 0], a[:, 3] - a[:, 1])
    return np.column_stack([a, d])


plain = make_pipeline(StandardScaler(), Ridge())
engineered = make_pipeline(FunctionTransformer(add_distance), StandardScaler(), Ridge())

for name, pipe in [("raw coordinates only", plain), ("with the distance step", engineered)]:
    scores = cross_val_score(pipe, X, y, cv=5, scoring="neg_mean_absolute_error")
    print(f"{name:24s} cross-validated error {-scores.mean():5.2f} min  "
          f"(spread {scores.std():.2f})")
Output
raw coordinates only     cross-validated error  6.33 min  (spread 0.51)
with the distance step   cross-validated error  1.26 min  (spread 0.07)

The error result matches the earlier experiment, now measured across five folds rather than one split.

Look at the spread as well. It fell from 0.51 to 0.07. The engineered model is not only better on average, it is far more consistent from fold to fold. A model that has been handed the right shape stops depending on which rows it happened to see.

StandardScaler sits inside the pipeline, so its mean and standard deviation are computed on each fold's training rows alone. Fit it outside and every fold's "held-out" data has already influenced the scaling. Train, test and validation splits measures how large that error can get.

A working checklist

Not a recipe. A set of questions worth asking about every column you have.

If you haveConsider building
Two positionsDistance, bearing, whether they are in the same zone
A timestampHour and weekday as sine/cosine, is-weekend, is-holiday, days since a known event
Two related amountsTheir ratio — utilisation, conversion rate, price per square foot
A skewed amountlog1p of it, so the tail stops dominating
A category with few levelsOne-hot encoding
A category with many levelsTargetEncoder, or frequency of occurrence
A history of eventsCount, sum, mean and recency over the last 7, 30 and 90 days
Free textLength, word count, digit count, link count, then embeddings
Anything with gapsA was_missing flag alongside the filled value

Two general notes. Ratios are consistently the highest-value construction on business data, because they remove size and leave behavior. And a was_missing flag costs one column and often carries real signal, as the XGBoost lesson demonstrates with loan data.

Common mistakes

Scaling or encoding before splitting. The most common leak in existence. Everything that learns from data goes inside the pipeline.

Building a feature from information that arrives later. delivery_rating predicts was_delivered perfectly and does not exist when you need to predict. Ask of every column: would this value be available, at this exact moment, in production?

Adding hundreds of features because you can. Each one is another chance to fit noise, and irrelevant columns dilute the useful ones. Add a feature because you can explain why it should matter.

Hand-rolling target encoding. Demonstrated above. Use the library version.

Believing importance scores about your new feature. Impurity-based importances are biased towards high-cardinality columns, as the random forest lesson measures. Judge a new feature by whether held-out error improved, and nothing else.

Different feature code in training and in production. Two implementations of days_since_last_order will drift apart, and the model degrades without any error appearing anywhere. One implementation, called from both places.

Try it yourself

Return to one_column.py and make the world less tidy. Replace the straight-line distance with a grid distance, np.abs(drop_x - pickup_x) + np.abs(drop_y - pickup_y), which is closer to how roads actually work. Keep giving the model the straight-line distance as its feature.

Now the feature is a good approximation rather than the exact truth. Measure how much of the improvement survives. Real features are never exactly right, and knowing what a nearly-right feature is worth is the useful skill.

Then delete the distance column and add drop_x - pickup_x and drop_y - pickup_y instead. Predict first whether linear regression can use these. Run it and find out. The answer explains precisely which shapes a linear model can and cannot reach.

What to learn next

Researcher — Mathematics and papers.

What a feature map does to the hypothesis class

A learning algorithm searches a hypothesis class H. Feature engineering replaces the input x ∈ X with φ(x) ∈ R^m and searches H ∘ φ instead. The algorithm is unchanged; the reachable set of functions is not.

For a linear model, H = { x ↦ w^T x + b } is closed under scaling and addition of coordinates and under nothing else. The coordinate experiment in the Developer section makes the consequence concrete: sqrt((x3−x1)^2 + (x4−x2)^2) lies outside the span of {x1, x2, x3, x4, 1}, so the target is unreachable at any sample size. The measured R² = −0.007 is not an estimation failure. It is an approximation failure, and no amount of data repairs it.

This is the classical decomposition. Total error splits into approximation error — the gap between the best function in H ∘ φ and the truth — and estimation error, the gap between what you fit and that best function. Model selection and regularisation address estimation error. Feature engineering addresses approximation error, and it is the only tool that does.

The kernel view, and why explicit features persist

The kernel trick computes ⟨φ(x), φ(x')⟩ without materialising φ, permitting infinite-dimensional feature spaces at O(n^2) cost in the sample size. KernelRidge with an RBF kernel would solve the coordinate problem without a hand-built feature.

Explicit feature construction survives for reasons that are practical rather than theoretical:

  • Kernel methods scale as O(n^2) to O(n^3) and do not reach the sample sizes tabular problems now carry. Random Fourier features (Rahimi & Recht, 2007) approximate the map explicitly in O(n m) and are the standard compromise.
  • An explicit feature is auditable, monitorable and explainable to a regulator. A kernel expansion is not.
  • Hand-built features encode domain constraints that no generic kernel contains. That a river crossing adds twenty minutes is not recoverable from three hundred coordinate pairs by any kernel.

Trees, rotation, and why they need features too

Axis-aligned recursive partitioning approximates any measurable function given unlimited depth and data, so trees are formally more expressive here than linear models. The Developer section's R² = 0.752 for boosted trees on raw coordinates reflects that.

The gap to 0.949 with the distance feature is a statement about sample efficiency. Representing a rotated or curved decision boundary with axis-aligned cuts requires a staircase whose step count grows with the required precision, and each step consumes data. Grinsztajn, Oyallon & Varoquaux (2022) identify non-rotation-invariance as one of the properties that makes tree ensembles strong on tabular data, and it is the same property that makes them inefficient on rotated targets. Oblique trees and models with learned rotations trade one bias for the other.

The practical rule that follows: trees need monotone transformations far less than linear models do — log, sqrt and any strictly increasing map leave the split structure invariant — and need constructed interactions and distances nearly as much.

Cyclical, periodic and spline encodings

For a value v with period p, the standard encoding is

φ(v) = ( sin(2π v / p) ,  cos(2π v / p) )

This embeds the cyclic group into R^2 so that Euclidean distance in the embedding is a monotone function of angular separation — verified numerically in the Developer section, where d(23:00, 00:00) = d(00:00, 01:00) = 0.261. One term is insufficient: sin alone is not injective on [0, p).

The encoding imposes a single sinusoid. When the daily pattern has two peaks, add harmonics sin(2πkv/p), cos(2πkv/p) for k = 1, 2, …, which is a truncated Fourier basis and a design choice about smoothness.

The alternative is a periodic spline basis: sklearn.preprocessing.SplineTransformer(extrapolation="periodic"). Splines are local, so a lunchtime spike does not distort the evening, whereas a Fourier basis is global and will. For sharply peaked periodic effects, splines dominate.

Target encoding, formally

Replace category c with an estimate of E[Y | C = c]. The naive estimator is the within-category mean, and it is biased when the same rows contribute to the estimate and the fit.

With n_c rows in category c and a global mean μ, empirical-Bayes smoothing shrinks small categories towards the prior:

enc(c) = ( n_c · ȳ_c  +  m · μ ) / ( n_c + m )

m is a smoothing weight, interpretable as a count of pseudo-observations from the prior. Micci-Barreca (2001) gives the original formulation with a sigmoid shrinkage schedule.

Smoothing alone does not remove leakage. The Developer section measures the residual: a pure-noise column with 200 levels over 560 rows yields train R² = 0.340 and test R² = −0.465. The row's own target is inside its own feature.

Two correct constructions exist.

Cross-fitted (K-fold) encoding. Partition the training rows; encode each fold using statistics from the other folds. This is what sklearn.preprocessing.TargetEncoder does internally, and it is why its fit_transform and transform return different values for the same input — an intentional asymmetry that surprises people the first time.

Ordered target statistics. Fix a random permutation and encode each row using only rows preceding it. Prokhorenkova et al. (2018) show that K-fold encoding still admits a residual prediction shift, and that ordering removes it. CatBoost extends the same permutation to the gradient estimates.

The general principle covers more than encoding: any statistic computed from y and attached to x must be computed out-of-fold. The same argument applies to feature selection, as demonstrated in train, test and validation splits.

Automated feature construction

Deep Feature Synthesis (Kanter & Veeramachaneni, 2015) generates features over relational schemas by stacking aggregation primitives (SUM, MEAN, COUNT across a one-to-many relation) and transform primitives (day-of-week, time-since) up to a depth limit. Featuretools implements it. The output space is combinatorial in depth, so selection afterwards is mandatory.

It works well for the aggregation-heavy features that relational business data invites, and it does not produce the coordinate-to-distance construction unless a distance primitive is supplied. That is the general boundary: automated systems compose primitives you gave them. Choosing the primitives is the domain work, relocated rather than removed.

Genetic programming approaches such as AutoFeat and symbolic regression search a wider space at substantially higher cost, with a strong tendency to produce expressions nobody can defend in review.

Training and serving skew

The dominant production failure mode is not a weak feature. It is a feature computed differently at training time and at serving time.

Two structural causes recur. The first is point-in-time correctness: a training feature computed from a table as it looks today includes information that did not exist at the timestamp of the row. Any aggregate over a customer's history must be truncated at the row's event time, which requires either an append-only event log or explicit as-of joins. The second is dual implementation: batch feature code in SQL or Python, serving feature code in application code. They diverge on edge cases and nothing raises an error.

Feature stores (Feast, Tecton, and the equivalents inside major platforms) exist to address exactly these two problems: one definition, materialised for both offline training and online serving, with as-of semantics enforced by the join. The engineering is unglamorous and the failure it prevents is silent, which is a fair summary of this whole subject.

Key references

  • Micci-Barreca, D. (2001). A Preprocessing Scheme for High-Cardinality Categorical Attributes in Classification and Prediction Problems. SIGKDD Explorations 3(1).
  • Rahimi, A. & Recht, B. (2007). Random Features for Large-Scale Kernel Machines. NeurIPS.
  • Guyon, I. & Elisseeff, A. (2003). An Introduction to Variable and Feature Selection. JMLR 3.
  • Kanter, J. & Veeramachaneni, K. (2015). Deep Feature Synthesis: Towards Automating Data Science Endeavors. IEEE DSAA.
  • Kaufman, S., Rosset, S. & Perlich, C. (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM TKDD 6(4).
  • Prokhorenkova, L. et al. (2018). CatBoost: Unbiased Boosting with Categorical Features. NeurIPS.
  • Sculley, D. et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS.
  • Grinsztajn, L., Oyallon, E. & Varoquaux, G. (2022). Why Do Tree-Based Models Still Outperform Deep Learning on Tabular Data? NeurIPS.
  • Kapoor, S. & Narayanan, A. (2023). Leakage and the Reproducibility Crisis in ML-Based Science. Patterns 4(9).

What to learn next