Error database

ValueError: could not convert string to float

A text value sat where a number was expected — often a category column fed to a model, or numbers stored with commas and currency signs. Find the text columns and encode or clean them.

The message you saw
ValueError: could not convert string to float

By Updated

The error

Output
ValueError: could not convert string to float: 'Mumbai'

The value after the colon changes: 'yes', '12,500', '₹499', '?', 'NA'. The cause is the same.

What it means

Something tried to turn text into a number and failed. Models in scikit-learn, and NumPy's float conversions, work on numbers only. The message helpfully shows the exact value that refused to convert — read it. It tells you which of the two situations below you are in.

Why it happens

Situation one: a genuine category column. 'Mumbai' or 'yes' is not a number and never will be. You passed a raw DataFrame with text columns straight into model.fit(X, y).

Situation two: numbers stored as text. '12,500' and '₹499' are numbers wearing costume — commas, currency symbols, percent signs, or placeholder strings like '?' for missing data. One such value makes pandas store the whole column as text.

How to fix it

1. Find which columns are text.

python
print(X.select_dtypes(include="object").columns.tolist())

2. For category columns, encode them instead of dropping them. The clean way is a ColumnTransformer so the same encoding applies at training and prediction time:

python
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression

cat_cols = ["city", "product"]
pre = ColumnTransformer(
    [("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols)],
    remainder="passthrough",
)
model = make_pipeline(pre, LogisticRegression(max_iter=1000))
model.fit(X, y)

One-hot encoding turns each category into its own 0/1 column, which is what linear models and most others need.

3. For numbers-as-text, strip the costume and convert.

python
X["price"] = (
    X["price"].astype(str)
    .str.replace(",", "", regex=False)
    .str.replace("₹", "", regex=False)
)
X["price"] = pd.to_numeric(X["price"], errors="coerce")

errors="coerce" turns anything unconvertible into NaN. Count the NaNs afterwards — a large count means the column held something you did not expect.

4. Tell read_csv about placeholder strings up front.

python
df = pd.read_csv("data.csv", na_values=["?", "NA", "n/a", "-"])

Now placeholders arrive as proper missing values, and numeric columns stay numeric.

How to prevent it

Run df.dtypes right after loading any file. A column you expected to be numeric showing object is this error waiting to happen. Handle categories with encoders inside a pipeline, not by hand, so training and inference always agree.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

  • Python for AI

    NumPy

    NumPy lets you do one operation to millions of numbers at once instead of one at a time. It is the foundation every AI library in Python is built on.

Back to all errors