ValueError: Input contains NaN, infinity or a value too large for dtype('float64')
Your array has missing values or infinities. Find which columns and rows carry them, then fix them inside a Pipeline so training and prediction are treated identically.
Updated
The error
Traceback (most recent call last):
File "train.py", line 31, in <module>
model.fit(X_train, y_train)
ValueError: Input contains NaN, infinity or a value too large for dtype('float64').Newer scikit-learn versions often name the culprit directly:
ValueError: Input X contains NaN. LinearRegression does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingRegressor which accept missing values encoded as NaNs natively.
What it means
scikit-learn checks every array before using it, and yours contains at least one value that is not a finite number: NaN (a missing value), or inf / -inf.
Most estimators cannot work with these. There is no meaningful distance to a missing value, no sensible average that includes infinity, and no correct split point for a value that is not a number. Rather than producing a silently wrong model, scikit-learn stops.
The "value too large" part of the message covers a rarer case: numbers beyond what the dtype can represent, which shows up when a float64 array holding very large values is cast down to float32.
Why it happens
Missing cells in the source data. Empty fields in a CSV become NaN when pandas reads them. So do values pandas could not parse as numbers, and placeholders like "N/A" or "-" once you tell pandas to treat them as missing.
A merge that did not match. pd.merge(..., how="left") fills every unmatched row with NaN across all the columns from the right-hand table. This one is easy to miss because the merge itself succeeds.
Feature engineering that divides. df["ratio"] = df["a"] / df["b"] produces inf wherever b is zero, and NaN where both are zero. np.log(x) gives -inf at zero and NaN for negatives. pct_change() and diff() leave NaN in the first row by design, and rolling windows leave NaN at the start of the series.
Manual scaling. Writing (x - x.mean()) / x.std() yourself produces NaN for any column where every value is the same, because the standard deviation is zero. scikit-learn's StandardScaler handles that case by using a scale of 1. That alone is a good reason to prefer it.
A gap between training and prediction. You cleaned the training data in the notebook and then called predict on fresh data that was never cleaned. The error arrives in production rather than in development, which is the worst time to meet it.
How to fix it
1. Find exactly where the bad values are. Do not guess which column is responsible.
import numpy as np
import pandas as pd
X = pd.DataFrame(X) # if it is a numpy array
nums = X.select_dtypes(include=[np.number])
print("NaN per column:")
print(X.isna().sum()[lambda s: s > 0].sort_values(ascending=False))
print("\ninf per column:")
print(np.isinf(nums).sum()[lambda s: s > 0].sort_values(ascending=False))
print("\nrows affected:", (X.isna().any(axis=1) | np.isinf(nums).any(axis=1)).sum(), "of", len(X))NaN per column: income 412 region_code 38 dtype: int64 inf per column: debt_ratio 17 dtype: int64 rows affected: 449 of 12000
Now you know both the size of the problem and where it came from. 412 missing incomes is a data collection issue; 17 infinite ratios is a division by zero in your own feature code.
2. Fix inf at its source, not with a blanket replace. An infinity means a calculation went wrong, and knowing which one matters.
df["debt_ratio"] = df["debt"] / df["income"].replace(0, np.nan) # 0 income → missing, not infinite
df["log_amount"] = np.log1p(df["amount"]) # log1p is defined at 0Turning inf into NaN first is often right, because it says "this value is unknown" rather than "this value is enormous":
X = X.replace([np.inf, -np.inf], np.nan)3. Impute inside a Pipeline. This is the fix that also prevents the error coming back, because the same filling rule is learned on the training data and reapplied automatically at prediction time.
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
("impute", SimpleImputer(strategy="median")), # median resists outliers better than mean
("scale", StandardScaler()),
("model", LogisticRegression(max_iter=1000)),
])
pipe.fit(X_train, y_train)
print(pipe.score(X_test, y_test))Filling with X.fillna(X.mean()) by hand before splitting is a genuine mistake, not a shortcut. The mean is computed over the test rows too, which leaks information. You get an optimistic score that will not survive contact with real data.
For a mix of numeric and categorical columns, ColumnTransformer applies a different imputer to each:
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
prep = ColumnTransformer([
("num", Pipeline([("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler())]), numeric_cols),
("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
("encode", OneHotEncoder(handle_unknown="ignore"))]), categorical_cols),
])4. Drop the rows only when there are few of them, and drop from X and y together.
mask = X.notna().all(axis=1)
X, y = X[mask], y[mask]Dropping from X alone gives you Found input variables with inconsistent numbers of samples on the next line. As a rough guide, dropping under 1-2% of rows is usually fine; dropping 30% throws away real signal, and the fact that a value is missing may itself be informative.
5. Add a "was missing" flag when missingness carries meaning. A blank income field on a loan application is not random.
SimpleImputer(strategy="median", add_indicator=True)6. Use a model that accepts missing values. scikit-learn's histogram gradient boosting handles NaN natively, routing missing values down whichever branch reduces the loss:
from sklearn.ensemble import HistGradientBoostingClassifier
model = HistGradientBoostingClassifier().fit(X_train, y_train) # NaN is fine, inf is notSome other tree-based estimators gained native missing-value support in recent scikit-learn versions — check the documentation for the version you have. Infinities are never accepted, by any of them.
7. For the "value too large" variant, look for numbers outside the dtype's range, which usually appear after a downcast to float32.
print(nums.abs().max().sort_values(ascending=False).head())Values above roughly 3.4 × 10^38 overflow float32. A log transform, or keeping the column in float64, resolves it.
How to prevent it
Validate at the boundary. Run one check as soon as data is loaded, so a bad file fails immediately with a clear message instead of failing twenty minutes later inside fit:
def check_frame(df: pd.DataFrame) -> None:
nums = df.select_dtypes(include=[np.number])
bad_inf = np.isinf(nums).sum()
bad_nan = df.isna().sum()
if bad_inf.any() or bad_nan.any():
raise ValueError(
f"infinite: {bad_inf[bad_inf > 0].to_dict()} | missing: {bad_nan[bad_nan > 0].to_dict()}"
)Do every transformation inside a Pipeline so training and serving cannot diverge. Save the fitted pipeline, not the bare model, and the imputation travels with it.
Guard divisions and logarithms as you write them rather than cleaning up afterwards — np.log1p instead of np.log, and a replaced zero denominator instead of an infinity. And keep an eye on df.describe() after a merge: a count that dropped for some columns and not others is the signature of unmatched rows.
Related errors
- ValueError: Found input variables with inconsistent numbers of samples — what happens when you drop rows from X without dropping them from y
- Loss is NaN during training — the deep learning version, where bad values pass through silently instead of being rejected