ColumnTransformer for mixed data
Real tables mix numbers and categories, and ColumnTransformer sends each kind of column through its own preparation before the model sees one combined table.
- 6 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ColumnTransformer lets each column of your table get its own preparation — numbers scaled, categories encoded — before everything merges back into one table.
Think of doing the household laundry. Whites go in one pile with hot water, colours in another with cold, and delicates get hand-washed. Different treatment per pile, yet everything ends up folded in the same wardrobe.
Real data tables need the same sorting. A loan application has income (a number), city (a choice from a list), and age (a number). One treatment cannot fit all of them.
Why it exists
Models only eat numbers. A column like city = "pune" means nothing to a model until it becomes numbers. That job is done by an encoder, a tool that turns categories into columns of 0s and 1s.
Number columns have the opposite need. They are already numbers, but on wildly different scales — income in lakhs, age in years. A scaler squashes them to a comparable range.
Before ColumnTransformer, people wrote loops that sliced the table apart, treated the pieces, and glued them back. That glue code was where bugs and leakage lived. ColumnTransformer replaced the glue with one declared routing.
How it works
┌─ city ──→ encoder (categories → 0/1 columns) ─┐
raw table ───┼─ rooms ──→ scaler ├─→ one wide
└─ sqft ──→ scaler ┘ number table
↓
modelYou declare the routing once: these columns to this tool, those columns to that tool. Fitting the whole thing fits every branch on the same training rows. The branches' outputs sit side by side in the final table.
A real example you have seen
Any online form you have filled mixes types: dropdowns for state, a number box for pincode, free text for a name. The software behind it treats each field by its type. ColumnTransformer is that same idea for model input.
Remember this
- Real tables mix numbers and categories, and each kind needs different preparation.
- ColumnTransformer routes named columns to named tools, then merges the outputs.
- It slots into a Pipeline like any other step, so the leakage protection still holds.
What to learn next
- Encoders and scalers, and the unknown-category trap — what each branch tool actually does.
- Writing your own transformer — when no built-in branch fits your data.
- Feature engineering — deciding what columns deserve to exist.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasTested against scikit-learn 1.7 and pandas 2.2.
Route columns, then train as one unit
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
flats = pd.DataFrame({
"city": ["mumbai", "delhi", "mumbai", "pune", "delhi", "pune"],
"rooms": [2, 3, 1, 2, 4, 3],
"sqft": [650, 900, 400, 750, 1200, 850],
})
pricey = pd.Series([1, 1, 0, 0, 1, 0]) # 1 = expensive flat
prep = ColumnTransformer([
("cat", OneHotEncoder(handle_unknown="ignore"), ["city"]),
("num", StandardScaler(), ["rooms", "sqft"]),
])
model = Pipeline([("prep", prep), ("clf", LogisticRegression())])
model.fit(flats, pricey)
print(model.named_steps["prep"].get_feature_names_out())
new_flat = pd.DataFrame({"city": ["mumbai"], "rooms": [3], "sqft": [1100]})
print("pricey?", model.predict(new_flat))['cat__city_delhi' 'cat__city_mumbai' 'cat__city_pune' 'num__rooms' 'num__sqft'] pricey? [1]
The walkthrough
Each entry is (name, transformer, columns). The name is yours, and it prefixes the output feature names — cat__city_delhi came from the cat branch. Column lists use DataFrame column names, which is the main reason to feed pipelines DataFrames rather than bare arrays.
The model never sees "mumbai". By the time LogisticRegression gets the data, the city column has become three 0/1 columns and the numbers are scaled. Five columns in total, matching the printed names.
Prediction takes a raw row. new_flat looks exactly like the training table. All the routing replays inside predict, using what each branch learned during fit.
Selecting columns by type instead of by name scales better than maintaining lists:
import numpy as np
from sklearn.compose import make_column_selector
prep = ColumnTransformer([
("cat", OneHotEncoder(handle_unknown="ignore"),
make_column_selector(dtype_include=object)),
("num", StandardScaler(),
make_column_selector(dtype_include=np.number)),
])This picks every text column for encoding and every numeric column for scaling, whatever the table grows into.
Common mistakes
Unlisted columns vanish. The default is remainder="drop": any column not routed anywhere is silently removed. If the model's accuracy is mysteriously poor, count your output columns. Pass remainder="passthrough" to keep the rest untouched — those come through named remainder__sqft and so on.
Scaling one-hot columns. Routing the encoder's 0/1 output through a scaler turns clean indicators into odd values and helps nothing. Branches run in parallel on the raw columns, so this mistake usually appears when people chain two ColumnTransformers. One layer of routing is almost always enough.
Feeding numpy arrays but selecting by name. Name-based selection needs a DataFrame. With a bare array you must select by column index, and index lists rot the moment someone reorders columns. Keep data in pandas until it enters the pipeline.
Assuming output column order matches the input table. Output order follows the transformer list order: all cat columns, then all num columns. Never index the output positionally; use get_feature_names_out() when you need to find a column.
Try it yourself
Add a furnished column with values "yes" and "no", and route it through the encoder branch. Before running, predict the new output of get_feature_names_out(). Then set remainder="passthrough" with furnished unrouted, and see how the names differ.
What to learn next
- Encoders and scalers, and the unknown-category trap — what each branch tool actually does.
- Writing your own transformer — when no built-in branch fits your data.
- Feature engineering — deciding what columns deserve to exist.
Researcher — Mathematics and papers.
Structure
ColumnTransformer applies transformers T_1..T_m to column subsets C_1..C_m of input X and concatenates outputs horizontally: the composed map is X ↦ [T_1(X[C_1]) | ... | T_m(X[C_m])]. It is the column-restricted analogue of FeatureUnion, which applies each transformer to all columns. Fit cost is the sum of branch costs; branches are independent, so n_jobs parallelises them.
Design notes that matter in practice:
- Sparse promotion. If the density of the stacked output falls below
sparse_threshold(default 0.3), the result is a scipy sparse matrix. A one-hot branch with high-cardinality categories routinely triggers this; downstream estimators must accept sparse input or you densify explicitly. set_output(transform="pandas")propagates names through the whole composition, making post-hoc coefficient inspection reliable: coefficient i belongs toget_feature_names_out()[i].- Column specification is resolved at fit time.
make_column_selectorstores a predicate, not a column list, so the fitted object records which columns it actually consumed (transformers_).
Statistical considerations for the branches
The encoder-versus-scaler split is a special case of a more general point: preprocessing choice is per-feature-type, and some choices are estimators with real variance. TargetEncoder (added in scikit-learn 1.3) replaces a category with a shrunken estimate of E[y | category]; naively fitted, it is a textbook leaker, so the implementation cross-fits internally during fit_transform. Micci-Barreca (2001), A preprocessing scheme for high-cardinality categorical attributes, is the standard reference for the shrinkage form.
For linear downstream models, one-hot encoding with an intercept induces perfect collinearity across each category block; regularisation absorbs it, which is why scikit-learn's LogisticRegression (L2 by default) is untroubled while unpenalised OLS needs drop="first".
Provenance
There is no standalone paper; the component follows the composition principles of Buitinck et al. (2013), API design for machine learning software. The historical antecedent is the DataFrameMapper from the sklearn-pandas package, which ColumnTransformer absorbed into core in version 0.20 (2018).
What to learn next
- Encoders and scalers, and the unknown-category trap — what each branch tool actually does.
- Writing your own transformer — when no built-in branch fits your data.
- Feature engineering — deciding what columns deserve to exist.