Error database

ValueError: Found input variables with inconsistent numbers of samples

Your X and y have different numbers of rows. Usually one of them was filtered, shifted or reshaped while the other was not.

The message you saw
ValueError: Found input variables with inconsistent numbers of samples

By Updated

The error

Output
Traceback (most recent call last):
  File "train.py", line 24, in <module>
    model.fit(X_train, y_train)
ValueError: Found input variables with inconsistent numbers of samples: [1000, 800]

The same message appears from train_test_split, cross_val_score, and every metric function:

Output
ValueError: Found input variables with inconsistent numbers of samples: [250, 1000]

What it means

scikit-learn compared the number of rows in each array you passed and they disagree. The list at the end is the row count of each argument, in the order you gave them. [1000, 800] means the features have 1,000 rows and the labels have 800.

Every estimator needs one label per row of features. There is no rule it could apply to decide which 200 rows to drop, so it stops instead of guessing.

Why it happens

The rows were separated somewhere, and only one side was changed.

The most frequent version is cleaning X but not y. X = X.dropna() removes rows from the features and leaves the labels untouched, so the counts diverge silently and fail later at fit. Filtering with a boolean mask applied to one array only does the same thing.

Time-series work has its own version. Creating a target with df["target"] = df["price"].shift(-1) leaves a missing value in the last row; dropping it from y without dropping it from X leaves you one row apart.

Then there is the transposed array — a features matrix stored as (features × samples) instead of (samples × features), which produces wildly mismatched numbers like [13, 506]. And the metric-function version, where you compare y_test against predictions made from a different X:

python
y_pred = model.predict(X_train)         # 800 predictions
accuracy_score(y_test, y_pred)          # against 200 true labels → this error

A subtler one: train_test_split returns four arrays in a fixed order, and unpacking them in the wrong order gives you X_train, X_test, y_train, y_test where you expected something else, which shows up as this error one line later.

How to fix it

1. Print the shapes. The answer is usually visible immediately.

python
print("X:", X.shape, "y:", len(y))
print("X_train:", X_train.shape, "y_train:", y_train.shape)
print("X_test:", X_test.shape, "y_test:", y_test.shape)

Whichever pair disagrees tells you where in the pipeline the rows were lost.

2. Do all filtering on the DataFrame, before splitting X and y. This is the fix that prevents the whole family of bugs, because the rows cannot separate if they never leave the same table.

python
import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv("data.csv")
features = ["age", "income", "region"]
target = "churned"

df = df.dropna(subset=features + [target])       # both sides drop together
X = df[features]
y = df[target]

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print(X_train.shape, y_train.shape)
Output
(3184, 3) (3184,)

3. If you must filter after splitting, apply the same mask to both.

python
mask = X["age"] > 18
X, y = X[mask], y[mask]                   # never one without the other

For NumPy arrays the pattern is identical: X, y = X[mask], y[mask].

4. For a shifted target, drop the incomplete rows from both sides.

python
df["target"] = df["price"].shift(-1)      # tomorrow's price
df = df.dropna(subset=["target"])         # removes the final row from X and y at once
X = df[["price", "volume"]]
y = df["target"]

5. If the counts look transposed, transpose. Only when the printed shapes confirm it — X.shape == (13, 506) for 506 houses and 13 features is unmistakable.

python
X = X.T

6. Check the metric call is comparing matching sets. Predictions must come from the X that matches the y.

python
y_pred = model.predict(X_test)            # not X_train
print(len(y_test), len(y_pred))
accuracy_score(y_test, y_pred)

7. When the index is involved, reset it. Concatenating or filtering pandas objects preserves the original index, and combining two objects with different indexes can align in unexpected ways.

python
X = X.reset_index(drop=True)
y = y.reset_index(drop=True)

How to prevent it

Keep features and labels in one DataFrame for as long as possible, and separate them in the last step before splitting. Every row-dropping operation should happen while they are still together.

Add one assertion after each step that could change row counts. It turns a confusing failure at fit into an obvious failure at the exact line responsible:

python
assert len(X) == len(y), f"X has {len(X)} rows, y has {len(y)}"

Use scikit-learn Pipeline objects for preprocessing rather than transforming arrays by hand. A pipeline applies steps consistently to train and test data and cannot drop rows from one and not the other. And when a transformation genuinely has to remove rows, do it in pandas before the pipeline, not inside it — scikit-learn transformers are built to change columns, not to change the number of rows.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors