Error database

NotFittedError: This instance is not fitted yet (scikit-learn)

You called predict or transform on an estimator that never had fit called on it — or on a different object than the one you fitted. Check which instance you are actually using.

The message you saw
NotFittedError: This instance is not fitted yet (scikit-learn)

By Updated

The error

Output
sklearn.exceptions.NotFittedError: This LogisticRegression instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator.

The class name varies — StandardScaler, RandomForestClassifier, TfidfVectorizer — but the sentence is the same.

What it means

Every scikit-learn estimator has two lives. Before fit, it holds only settings. After fit, it holds learned state — coefficients, trees, vocabulary, means. You asked an unfitted instance to do a fitted instance's job.

The subtle part: the object you called is often not the object you fitted. Same class, different instance.

Why it happens

The genuinely-forgot case is rare. The common cases are identity mix-ups:

  • You fitted a pipeline step through the pipeline, then called the bare step — or fitted the bare step and called it through a new pipeline.
  • cross_val_score(model, X, y) fits clones of your model. The original model stays unfitted afterwards, by design.
  • At inference time you constructed a fresh StandardScaler() instead of loading the one fitted on training data.
  • You saved the model before fitting it, then loaded the unfitted version.

How to fix it

1. Fit before you predict — on the same object.

python
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
preds = model.predict(X_test)

2. After cross-validation, fit once more on the full training data.

python
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X_train, y_train, cv=5)
model.fit(X_train, y_train)          # cross_val_score fitted clones, not this one

3. For transformers, fit on train, transform everywhere — with one instance.

python
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)   # same scaler, no second fit

Fitting a second scaler on test data is both this error's cousin and a data leak.

4. Persist the fitted object, and load that at inference time.

python
import joblib
joblib.dump(model, "model.joblib")     # after fit
model = joblib.load("model.joblib")    # elsewhere: already fitted

5. When unsure, ask.

python
from sklearn.utils.validation import check_is_fitted
check_is_fitted(model)                 # raises NotFittedError if not

How to prevent it

Bundle preprocessing and model into one Pipeline, fit the pipeline, save the pipeline. One object, one fit, nothing to desynchronise. Name fitted things for what they are (scaler_fitted, or the pipeline itself) so an unfitted twin cannot masquerade.