ValueError: Length of values does not match length of index (pandas)
You assigned a list or array with the wrong number of rows as a DataFrame column. The values were computed from a different frame — usually the unfiltered one.
Updated
The error
ValueError: Length of values (900) does not match length of index (1000)
What it means
A DataFrame column must have exactly one value per row. You assigned a sequence of 900 values to a frame with 1000 rows. Pandas cannot decide which rows should get values and which should not, so it refuses.
Why it happens
Almost always, the values were computed from a different version of the frame. A typical sequence:
df = pd.read_csv("orders.csv") # 1000 rows
clean = df.dropna(subset=["price"]) # 900 rows
scores = model.predict(clean[features]) # 900 predictions
df["score"] = scores # boom: 900 into 1000Other routes: predictions from a train/test split assigned to the full frame, a list built in a loop that skipped some rows, or .unique() output (deduplicated, so shorter) assigned back as a column.
How to fix it
1. Assign to the frame the values actually came from.
clean["score"] = scores # 900 into 900If clean was a slice, create it with .copy() first to avoid the SettingWithCopyWarning.
2. To place partial results into the full frame, assign a Series with the right index. Pandas aligns Series assignments by index label; rows without a match get NaN.
df["score"] = pd.Series(scores, index=clean.index)This works because clean kept the original row labels. It fails silently after a reset_index, so check which index the slice carries.
3. For lookups, use map or merge instead of positional assignment.
df["city_tier"] = df["city"].map(tier_lookup)map matches by value and cannot go out of step with row counts.
4. When lengths should match but do not, find the dropped rows.
print(len(df), len(clean))
print(df.index.difference(clean.index))Then decide: filter the frame first, or stop filtering.
How to prevent it
Keep one rule: compute new columns from the same frame you assign them to, in the line directly above. When you filter, give the filtered frame a clear name and keep using it. Note that assigning a plain NumPy array skips index alignment entirely — lengths must match exactly — while assigning a Series aligns on labels. Choose deliberately.
Related errors
- Found input variables with inconsistent numbers of samples — the scikit-learn version of the same drift
- SettingWithCopyWarning
- KeyError: column not in index
- You are trying to merge on object and int64 columns