ValueError: cannot convert float NaN to integer
You cast a column with missing values to int, and NaN has no integer form. Fill or drop the missing values first — or use the nullable Int64 dtype.
Updated
The error
ValueError: cannot convert float NaN to integer
Pandas raises its own wording for the same situation:
ValueError: Cannot convert non-finite values (NA or inf) to integer
What it means
NaN ("not a number") is how pandas and NumPy mark a missing value, and it only exists as a float. There is no integer that means "missing". So astype(int) on a column containing NaN has no valid answer, and the conversion refuses.
This is also why a column of whole numbers mysteriously shows as float64: one missing value forces the entire column to float, and 1042 displays as 1042.0.
Why it happens
The NaN often arrives earlier than you think. A merge that found no match fills the gap with NaN. pd.to_numeric(..., errors="coerce") turns bad strings into NaN. A CSV with empty cells loads them as NaN. Then, sometime later, an innocent astype(int) hits the mine.
How to fix it
1. Count and inspect the missing values first.
print(df["units"].isna().sum())
print(df[df["units"].isna()].head())Look at the rows. Are they genuinely missing data, a failed merge, or bad strings that got coerced? The right fix depends on the answer.
2. Fill with a meaningful value, then convert.
df["units"] = df["units"].fillna(0).astype(int)Fill with 0 only when 0 is truthful. For a count of items ordered, 0 may be right. For a rating, filling with 0 invents a terrible review — use the median, or keep it missing.
3. Drop the rows when they cannot be repaired.
df = df.dropna(subset=["units"])
df["units"] = df["units"].astype(int)4. Keep missing values AND integers with the nullable dtype.
df["units"] = df["units"].astype("Int64") # capital IPandas' nullable integer type stores whole numbers plus a proper missing marker (<NA>). This is the honest option when missing is a real state you want to preserve.
How to prevent it
Treat any float64 column that should be whole numbers as a signal: missing values are hiding in it. Audit NaN counts right after loading and right after every merge — df.isna().sum() is one line. Decide the missing-value policy per column, on purpose, instead of letting a late astype make the decision for you.
Related errors
- ValueError: Input contains NaN, infinity or a value too large — the same NaNs reaching scikit-learn
- ValueError: could not convert string to float
- You are trying to merge on object and int64 columns — merges are a common NaN source