ValueError: Found input variables with inconsistent numbers of samples
Your X and y have different numbers of rows. Usually one of them was filtered, shifted or reshaped while the other was not.
Updated
The error
Traceback (most recent call last):
File "train.py", line 24, in <module>
model.fit(X_train, y_train)
ValueError: Found input variables with inconsistent numbers of samples: [1000, 800]The same message appears from train_test_split, cross_val_score, and every metric function:
ValueError: Found input variables with inconsistent numbers of samples: [250, 1000]
What it means
scikit-learn compared the number of rows in each array you passed and they disagree. The list at the end is the row count of each argument, in the order you gave them. [1000, 800] means the features have 1,000 rows and the labels have 800.
Every estimator needs one label per row of features. There is no rule it could apply to decide which 200 rows to drop, so it stops instead of guessing.
Why it happens
The rows were separated somewhere, and only one side was changed.
The most frequent version is cleaning X but not y. X = X.dropna() removes rows from the features and leaves the labels untouched, so the counts diverge silently and fail later at fit. Filtering with a boolean mask applied to one array only does the same thing.
Time-series work has its own version. Creating a target with df["target"] = df["price"].shift(-1) leaves a missing value in the last row; dropping it from y without dropping it from X leaves you one row apart.
Then there is the transposed array — a features matrix stored as (features × samples) instead of (samples × features), which produces wildly mismatched numbers like [13, 506]. And the metric-function version, where you compare y_test against predictions made from a different X:
y_pred = model.predict(X_train) # 800 predictions
accuracy_score(y_test, y_pred) # against 200 true labels → this errorA subtler one: train_test_split returns four arrays in a fixed order, and unpacking them in the wrong order gives you X_train, X_test, y_train, y_test where you expected something else, which shows up as this error one line later.
How to fix it
1. Print the shapes. The answer is usually visible immediately.
print("X:", X.shape, "y:", len(y))
print("X_train:", X_train.shape, "y_train:", y_train.shape)
print("X_test:", X_test.shape, "y_test:", y_test.shape)Whichever pair disagrees tells you where in the pipeline the rows were lost.
2. Do all filtering on the DataFrame, before splitting X and y. This is the fix that prevents the whole family of bugs, because the rows cannot separate if they never leave the same table.
import pandas as pd
from sklearn.model_selection import train_test_split
df = pd.read_csv("data.csv")
features = ["age", "income", "region"]
target = "churned"
df = df.dropna(subset=features + [target]) # both sides drop together
X = df[features]
y = df[target]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print(X_train.shape, y_train.shape)(3184, 3) (3184,)
3. If you must filter after splitting, apply the same mask to both.
mask = X["age"] > 18
X, y = X[mask], y[mask] # never one without the otherFor NumPy arrays the pattern is identical: X, y = X[mask], y[mask].
4. For a shifted target, drop the incomplete rows from both sides.
df["target"] = df["price"].shift(-1) # tomorrow's price
df = df.dropna(subset=["target"]) # removes the final row from X and y at once
X = df[["price", "volume"]]
y = df["target"]5. If the counts look transposed, transpose. Only when the printed shapes confirm it — X.shape == (13, 506) for 506 houses and 13 features is unmistakable.
X = X.T6. Check the metric call is comparing matching sets. Predictions must come from the X that matches the y.
y_pred = model.predict(X_test) # not X_train
print(len(y_test), len(y_pred))
accuracy_score(y_test, y_pred)7. When the index is involved, reset it. Concatenating or filtering pandas objects preserves the original index, and combining two objects with different indexes can align in unexpected ways.
X = X.reset_index(drop=True)
y = y.reset_index(drop=True)How to prevent it
Keep features and labels in one DataFrame for as long as possible, and separate them in the last step before splitting. Every row-dropping operation should happen while they are still together.
Add one assertion after each step that could change row counts. It turns a confusing failure at fit into an obvious failure at the exact line responsible:
assert len(X) == len(y), f"X has {len(X)} rows, y has {len(y)}"Use scikit-learn Pipeline objects for preprocessing rather than transforming arrays by hand. A pipeline applies steps consistently to train and test data and cannot drop rows from one and not the other. And when a transformation genuinely has to remove rows, do it in pandas before the pipeline, not inside it — scikit-learn transformers are built to change columns, not to change the number of rows.
Related errors
- ValueError: Input contains NaN, infinity or a value too large for dtype('float64') — the error you often hit next, after deciding not to drop the missing rows
- mat1 and mat2 shapes cannot be multiplied — the PyTorch equivalent of a shape disagreement