ValueError: could not convert string to float
A text value sat where a number was expected — often a category column fed to a model, or numbers stored with commas and currency signs. Find the text columns and encode or clean them.
Updated
The error
ValueError: could not convert string to float: 'Mumbai'
The value after the colon changes: 'yes', '12,500', '₹499', '?', 'NA'. The cause is the same.
What it means
Something tried to turn text into a number and failed. Models in scikit-learn, and NumPy's float conversions, work on numbers only. The message helpfully shows the exact value that refused to convert — read it. It tells you which of the two situations below you are in.
Why it happens
Situation one: a genuine category column. 'Mumbai' or 'yes' is not a number and never will be. You passed a raw DataFrame with text columns straight into model.fit(X, y).
Situation two: numbers stored as text. '12,500' and '₹499' are numbers wearing costume — commas, currency symbols, percent signs, or placeholder strings like '?' for missing data. One such value makes pandas store the whole column as text.
How to fix it
1. Find which columns are text.
print(X.select_dtypes(include="object").columns.tolist())2. For category columns, encode them instead of dropping them. The clean way is a ColumnTransformer so the same encoding applies at training and prediction time:
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
cat_cols = ["city", "product"]
pre = ColumnTransformer(
[("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)],
remainder="passthrough",
)
model = make_pipeline(pre, LogisticRegression(max_iter=1000))
model.fit(X, y)One-hot encoding turns each category into its own 0/1 column, which is what linear models and most others need.
3. For numbers-as-text, strip the costume and convert.
X["price"] = (
X["price"].astype(str)
.str.replace(",", "", regex=False)
.str.replace("₹", "", regex=False)
)
X["price"] = pd.to_numeric(X["price"], errors="coerce")errors="coerce" turns anything unconvertible into NaN. Count the NaNs afterwards — a large count means the column held something you did not expect.
4. Tell read_csv about placeholder strings up front.
df = pd.read_csv("data.csv", na_values=["?", "NA", "n/a", "-"])Now placeholders arrive as proper missing values, and numeric columns stay numeric.
How to prevent it
Run df.dtypes right after loading any file. A column you expected to be numeric showing object is this error waiting to happen. Handle categories with encoders inside a pipeline, not by hand, so training and inference always agree.
Related errors
- ValueError: Input contains NaN, infinity or a value too large — the usual next error, after coercing to NaN
- ValueError: Unknown label type: 'continuous'
- ValueError: cannot convert float NaN to integer