Feature engineering
In one sentence Feature engineering is creating better input columns from raw data, using what you know about the problem.
Updated
Feature engineering is transforming raw data into features that make the pattern easier for a model to see.
A raw timestamp like 2026-08-31 02:14:07 is nearly useless to a fraud model. But derived from it: hour of day, weekday or weekend, seconds since this card's previous transaction. Those columns carry the signal — 2 a.m. purchases minutes apart are suspicious in a way no model can read directly off a timestamp string. Nothing new was collected; knowledge of the problem was folded into the columns.
The craft has recurring moves. Ratios and differences (price ÷ area beats price and area separately, for judging flats). Aggregations (a customer's average order value over 90 days). Extractions (domain name out of an email address). Encodings that turn categories into numbers, such as one-hot-encoding. Each move injects human understanding the raw data only implies.
For gradient-boosted trees and other classical models on tabular data, this is routinely where competitions and production systems are won — more than by the choice of algorithm. Deep learning automates much of it for images, audio and text, which is precisely why those fields moved to neural networks.
The standing danger is data-leakage: building a feature from information that will not exist at prediction time, like "total purchases this month" computed with the month's end included.
Where to go next
- Full lesson: What is machine learning?
- Related terms: feature, label, feature-scaling, data-leakage