Error database

ValueError: Unknown label type: 'continuous' (scikit-learn)

You gave a classifier a target of continuous numbers. If you are predicting a quantity, use a regressor; if your labels are real classes stored as floats, convert them to integers.

The message you saw
ValueError: Unknown label type: 'continuous' (scikit-learn)

By Updated

The error

Output
ValueError: Unknown label type: 'continuous'

Newer scikit-learn versions spell out the diagnosis:

Output
ValueError: Unknown label type: continuous. Maybe you are trying to fit a classifier, which expects discrete classes on a regression target with continuous values.

What it means

Classifiers predict categories: spam or not, cat or dog or horse. Their target must be a set of discrete labels. You passed a target full of continuous numbers — 4.7, 12.31, 899.5 — which reads as a quantity, not categories. The model cannot treat every distinct float as its own class, so it stops.

Why it happens

The most common cause is a naming trap: LogisticRegression is a classifier, despite the word "regression" in its name. Students predicting house prices reach for it and hit this error immediately.

The second cause is real class labels stored as floats. Labels exported as 0.0 and 1.0, or labels that picked up NaN (which forces a column to float), look continuous to the type checker.

How to fix it

1. Predicting a quantity? Use a regressor.

python
from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor

model = LinearRegression()          # not LogisticRegression
model.fit(X_train, y_train)

Every popular classifier has a regressor twin: RandomForestRegressor, GradientBoostingRegressor, SVR. Score with regression metrics — MAE, RMSE, R² — not accuracy.

2. Labels are real classes stored as floats? Convert them.

python
print(y.unique())                   # [0.0, 1.0] — classes in float costume
y = y.astype(int)

If NaN blocks the conversion, missing labels exist — handle those rows first.

3. Text or mixed labels? Encode them.

python
from sklearn.preprocessing import LabelEncoder
y = LabelEncoder().fit_transform(y)

4. Want to classify a continuous outcome? Make the threshold an explicit decision.

python
y_class = (df["income"] > 50_000).astype(int)

Binning a quantity throws away information. Do it because the question is genuinely categorical ("will churn: yes/no"), not to dodge this error — otherwise prefer regression.

How to prevent it

Before fitting, print y.dtype and y.nunique(). Thousands of unique float values means regression; a handful of values means classification. Choose the estimator to match the question, and remember the one misleading name: logistic regression classifies.