Error database

ValueError: X has N features, but the model is expecting M features

The data you passed at prediction time has a different column count than the data used for fitting. Train-time and inference-time preprocessing drifted apart — a pipeline fixes it for good.

The message you saw
ValueError: X has N features, but the model is expecting M features

By Updated

The error

Output
ValueError: X has 10 features, but RandomForestClassifier is expecting 12 features as input.

With DataFrames, a related complaint names the columns:

Output
ValueError: The feature names should match those that were passed during fit.
Feature names unseen at fit time:
- city_Indore
Feature names seen at fit time, yet now missing:
- city_Jaipur
- city_Nagpur

What it means

A fitted model remembers how many features it was trained on, and (for DataFrames) their names. The rows you now want predictions for have a different set. The model cannot guess which columns moved, appeared or vanished, so it refuses.

Why it happens

The features themselves rarely change — the preprocessing does. The classic culprit is pd.get_dummies: it creates one column per category present in the data it sees. Training data had customers from 12 cities; today's batch has 10 of them plus one new city. Different columns, same code.

Other routes: a column dropped in the training script but not the serving script, features built in a different order by hand, or a raw DataFrame passed where the scaled version was expected.

How to fix it

1. Compare the two column sets — the diff names the drift.

python
expected = list(model.feature_names_in_)     # what fit saw (DataFrame input)
got = list(X_new.columns)
print(set(expected) - set(got))              # missing now
print(set(got) - set(expected))              # unexpected now

2. The durable fix: put encoding inside a Pipeline, and save the whole thing.

python
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.pipeline import make_pipeline
from sklearn.ensemble import RandomForestClassifier
import joblib

pre = ColumnTransformer(
    [("cat", OneHotEncoder(handle_unknown="ignore"), ["city", "plan"])],
    remainder="passthrough",
)
model = make_pipeline(pre, RandomForestClassifier())
model.fit(X_train, y_train)
joblib.dump(model, "model.joblib")

OneHotEncoder learns the category list at fit time and reproduces the exact same columns forever. handle_unknown="ignore" makes new categories encode as all zeros instead of crashing.

3. Stuck with get_dummies? Reindex to the training columns.

python
X_new = pd.get_dummies(X_new)
X_new = X_new.reindex(columns=train_columns, fill_value=0)

This requires you to have saved train_columns at training time — which is exactly the bookkeeping the pipeline does for you.

4. Feed columns in the same structure fit received. If you fitted on a DataFrame, predict on a DataFrame with the same column names, not a bare NumPy array — the names warning exists to protect you from silent column-order bugs.

How to prevent it

One rule prevents this whole error class: every transformation between raw data and model lives inside the saved Pipeline. Serving code then does joblib.load(...).predict(raw_df) and cannot drift.