Error database

ValueError: Input contains NaN, infinity or a value too large for dtype('float64')

Your array has missing values or infinities. Find which columns and rows carry them, then fix them inside a Pipeline so training and prediction are treated identically.

The message you saw
ValueError: Input contains NaN, infinity or a value too large for dtype('float64')

By Updated

The error

Output
Traceback (most recent call last):
  File "train.py", line 31, in <module>
    model.fit(X_train, y_train)
ValueError: Input contains NaN, infinity or a value too large for dtype('float64').

Newer scikit-learn versions often name the culprit directly:

Output
ValueError: Input X contains NaN.
LinearRegression does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingRegressor which accept missing values encoded as NaNs natively.

What it means

scikit-learn checks every array before using it, and yours contains at least one value that is not a finite number: NaN (a missing value), or inf / -inf.

Most estimators cannot work with these. There is no meaningful distance to a missing value, no sensible average that includes infinity, and no correct split point for a value that is not a number. Rather than producing a silently wrong model, scikit-learn stops.

The "value too large" part of the message covers a rarer case: numbers beyond what the dtype can represent, which shows up when a float64 array holding very large values is cast down to float32.

Why it happens

Missing cells in the source data. Empty fields in a CSV become NaN when pandas reads them. So do values pandas could not parse as numbers, and placeholders like "N/A" or "-" once you tell pandas to treat them as missing.

A merge that did not match. pd.merge(..., how="left") fills every unmatched row with NaN across all the columns from the right-hand table. This one is easy to miss because the merge itself succeeds.

Feature engineering that divides. df["ratio"] = df["a"] / df["b"] produces inf wherever b is zero, and NaN where both are zero. np.log(x) gives -inf at zero and NaN for negatives. pct_change() and diff() leave NaN in the first row by design, and rolling windows leave NaN at the start of the series.

Manual scaling. Writing (x - x.mean()) / x.std() yourself produces NaN for any column where every value is the same, because the standard deviation is zero. scikit-learn's StandardScaler handles that case by using a scale of 1. That alone is a good reason to prefer it.

A gap between training and prediction. You cleaned the training data in the notebook and then called predict on fresh data that was never cleaned. The error arrives in production rather than in development, which is the worst time to meet it.

How to fix it

1. Find exactly where the bad values are. Do not guess which column is responsible.

python
import numpy as np
import pandas as pd

X = pd.DataFrame(X)                      # if it is a numpy array
nums = X.select_dtypes(include=[np.number])

print("NaN per column:")
print(X.isna().sum()[lambda s: s > 0].sort_values(ascending=False))

print("\ninf per column:")
print(np.isinf(nums).sum()[lambda s: s > 0].sort_values(ascending=False))

print("\nrows affected:", (X.isna().any(axis=1) | np.isinf(nums).any(axis=1)).sum(), "of", len(X))
Output
NaN per column:
income        412
region_code    38
dtype: int64

inf per column:
debt_ratio    17
dtype: int64

rows affected: 449 of 12000

Now you know both the size of the problem and where it came from. 412 missing incomes is a data collection issue; 17 infinite ratios is a division by zero in your own feature code.

2. Fix inf at its source, not with a blanket replace. An infinity means a calculation went wrong, and knowing which one matters.

python
df["debt_ratio"] = df["debt"] / df["income"].replace(0, np.nan)   # 0 income → missing, not infinite
df["log_amount"] = np.log1p(df["amount"])                          # log1p is defined at 0

Turning inf into NaN first is often right, because it says "this value is unknown" rather than "this value is enormous":

python
X = X.replace([np.inf, -np.inf], np.nan)

3. Impute inside a Pipeline. This is the fix that also prevents the error coming back, because the same filling rule is learned on the training data and reapplied automatically at prediction time.

python
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([
    ("impute", SimpleImputer(strategy="median")),   # median resists outliers better than mean
    ("scale",  StandardScaler()),
    ("model",  LogisticRegression(max_iter=1000)),
])

pipe.fit(X_train, y_train)
print(pipe.score(X_test, y_test))

Filling with X.fillna(X.mean()) by hand before splitting is a genuine mistake, not a shortcut. The mean is computed over the test rows too, which leaks information. You get an optimistic score that will not survive contact with real data.

For a mix of numeric and categorical columns, ColumnTransformer applies a different imputer to each:

python
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder

prep = ColumnTransformer([
    ("num", Pipeline([("impute", SimpleImputer(strategy="median")),
                      ("scale", StandardScaler())]), numeric_cols),
    ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                      ("encode", OneHotEncoder(handle_unknown="ignore"))]), categorical_cols),
])

4. Drop the rows only when there are few of them, and drop from X and y together.

python
mask = X.notna().all(axis=1)
X, y = X[mask], y[mask]

Dropping from X alone gives you Found input variables with inconsistent numbers of samples on the next line. As a rough guide, dropping under 1-2% of rows is usually fine; dropping 30% throws away real signal, and the fact that a value is missing may itself be informative.

5. Add a "was missing" flag when missingness carries meaning. A blank income field on a loan application is not random.

python
SimpleImputer(strategy="median", add_indicator=True)

6. Use a model that accepts missing values. scikit-learn's histogram gradient boosting handles NaN natively, routing missing values down whichever branch reduces the loss:

python
from sklearn.ensemble import HistGradientBoostingClassifier

model = HistGradientBoostingClassifier().fit(X_train, y_train)   # NaN is fine, inf is not

Some other tree-based estimators gained native missing-value support in recent scikit-learn versions — check the documentation for the version you have. Infinities are never accepted, by any of them.

7. For the "value too large" variant, look for numbers outside the dtype's range, which usually appear after a downcast to float32.

python
print(nums.abs().max().sort_values(ascending=False).head())

Values above roughly 3.4 × 10^38 overflow float32. A log transform, or keeping the column in float64, resolves it.

How to prevent it

Validate at the boundary. Run one check as soon as data is loaded, so a bad file fails immediately with a clear message instead of failing twenty minutes later inside fit:

python
def check_frame(df: pd.DataFrame) -> None:
    nums = df.select_dtypes(include=[np.number])
    bad_inf = np.isinf(nums).sum()
    bad_nan = df.isna().sum()
    if bad_inf.any() or bad_nan.any():
        raise ValueError(
            f"infinite: {bad_inf[bad_inf > 0].to_dict()} | missing: {bad_nan[bad_nan > 0].to_dict()}"
        )

Do every transformation inside a Pipeline so training and serving cannot diverge. Save the fitted pipeline, not the bare model, and the imputation travels with it.

Guard divisions and logarithms as you write them rather than cleaning up afterwards — np.log1p instead of np.log, and a replaced zero denominator instead of an infinity. And keep an eye on df.describe() after a merge: a count that dropped for some columns and not others is the signature of unmatched rows.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors