ValueError: Unknown label type: 'continuous' (scikit-learn)
You gave a classifier a target of continuous numbers. If you are predicting a quantity, use a regressor; if your labels are real classes stored as floats, convert them to integers.
Updated
The error
ValueError: Unknown label type: 'continuous'
Newer scikit-learn versions spell out the diagnosis:
ValueError: Unknown label type: continuous. Maybe you are trying to fit a classifier, which expects discrete classes on a regression target with continuous values.
What it means
Classifiers predict categories: spam or not, cat or dog or horse. Their target must be a set of discrete labels. You passed a target full of continuous numbers — 4.7, 12.31, 899.5 — which reads as a quantity, not categories. The model cannot treat every distinct float as its own class, so it stops.
Why it happens
The most common cause is a naming trap: LogisticRegression is a classifier, despite the word "regression" in its name. Students predicting house prices reach for it and hit this error immediately.
The second cause is real class labels stored as floats. Labels exported as 0.0 and 1.0, or labels that picked up NaN (which forces a column to float), look continuous to the type checker.
How to fix it
1. Predicting a quantity? Use a regressor.
from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor
model = LinearRegression() # not LogisticRegression
model.fit(X_train, y_train)Every popular classifier has a regressor twin: RandomForestRegressor, GradientBoostingRegressor, SVR. Score with regression metrics — MAE, RMSE, R² — not accuracy.
2. Labels are real classes stored as floats? Convert them.
print(y.unique()) # [0.0, 1.0] — classes in float costume
y = y.astype(int)If NaN blocks the conversion, missing labels exist — handle those rows first.
3. Text or mixed labels? Encode them.
from sklearn.preprocessing import LabelEncoder
y = LabelEncoder().fit_transform(y)4. Want to classify a continuous outcome? Make the threshold an explicit decision.
y_class = (df["income"] > 50_000).astype(int)Binning a quantity throws away information. Do it because the question is genuinely categorical ("will churn: yes/no"), not to dodge this error — otherwise prefer regression.
How to prevent it
Before fitting, print y.dtype and y.nunique(). Thousands of unique float values means regression; a handful of values means classification. Choose the estimator to match the question, and remember the one misleading name: logistic regression classifies.
Related errors
- Classification metrics can't handle a mix — the same confusion at evaluation time
- ValueError: could not convert string to float
- The least populated class in y has only 1 member