ValueError: X has N features, but the model is expecting M features
The data you passed at prediction time has a different column count than the data used for fitting. Train-time and inference-time preprocessing drifted apart — a pipeline fixes it for good.
Updated
The error
ValueError: X has 10 features, but RandomForestClassifier is expecting 12 features as input.
With DataFrames, a related complaint names the columns:
ValueError: The feature names should match those that were passed during fit. Feature names unseen at fit time: - city_Indore Feature names seen at fit time, yet now missing: - city_Jaipur - city_Nagpur
What it means
A fitted model remembers how many features it was trained on, and (for DataFrames) their names. The rows you now want predictions for have a different set. The model cannot guess which columns moved, appeared or vanished, so it refuses.
Why it happens
The features themselves rarely change — the preprocessing does. The classic culprit is pd.get_dummies: it creates one column per category present in the data it sees. Training data had customers from 12 cities; today's batch has 10 of them plus one new city. Different columns, same code.
Other routes: a column dropped in the training script but not the serving script, features built in a different order by hand, or a raw DataFrame passed where the scaled version was expected.
How to fix it
1. Compare the two column sets — the diff names the drift.
expected = list(model.feature_names_in_) # what fit saw (DataFrame input)
got = list(X_new.columns)
print(set(expected) - set(got)) # missing now
print(set(got) - set(expected)) # unexpected now2. The durable fix: put encoding inside a Pipeline, and save the whole thing.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.pipeline import make_pipeline
from sklearn.ensemble import RandomForestClassifier
import joblib
pre = ColumnTransformer(
[("cat", OneHotEncoder(handle_unknown="ignore"), ["city", "plan"])],
remainder="passthrough",
)
model = make_pipeline(pre, RandomForestClassifier())
model.fit(X_train, y_train)
joblib.dump(model, "model.joblib")OneHotEncoder learns the category list at fit time and reproduces the exact same columns forever. handle_unknown="ignore" makes new categories encode as all zeros instead of crashing.
3. Stuck with get_dummies? Reindex to the training columns.
X_new = pd.get_dummies(X_new)
X_new = X_new.reindex(columns=train_columns, fill_value=0)This requires you to have saved train_columns at training time — which is exactly the bookkeeping the pipeline does for you.
4. Feed columns in the same structure fit received. If you fitted on a DataFrame, predict on a DataFrame with the same column names, not a bare NumPy array — the names warning exists to protect you from silent column-order bugs.
How to prevent it
One rule prevents this whole error class: every transformation between raw data and model lives inside the saved Pipeline. Serving code then does joblib.load(...).predict(raw_df) and cannot drift.
Related errors
- NotFittedError: instance is not fitted yet
- Found input variables with inconsistent numbers of samples — the row-count version of this column-count error
- Expected 2D array, got 1D array