ValueError: Expected 2D array, got 1D array instead (scikit-learn)
scikit-learn wants X as a table — rows by columns — even when there is one feature or one sample. Reshape with (-1, 1) for one feature, or pass a double-bracketed row.
Updated
The error
ValueError: Expected 2D array, got 1D array instead: array=[4.5 6.1 7.2]. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
A related complaint appears with scalars: Expected 2D array, got scalar array instead.
What it means
Every scikit-learn X is a 2D table: one row per sample, one column per feature. That stays true when there is only one feature, or only one sample. You passed a flat 1D array, and the library cannot tell which of the two you meant: three samples of one feature, or one sample of three features? Rather than guess, it asks you to say.
Why it happens
Selecting a single DataFrame column with single brackets gives a 1D Series:
X = df["experience"] # 1D — this will failPredicting for one new sample with a flat list does the same:
model.predict([5.0]) # ambiguous 1DHow to fix it
1. One feature, many samples: reshape to a column.
X = df["experience"].to_numpy().reshape(-1, 1) # shape (n, 1)
model.fit(X, y)The -1 means "however many rows there are".
2. Cleaner with pandas: double brackets keep 2D.
X = df[["experience"]] # DataFrame, shape (n, 1)Single brackets give a Series (1D); double brackets give a DataFrame (2D). This one-character habit prevents the error entirely.
3. One sample at prediction time: wrap it in a list of lists.
model.predict([[5.0, 60000, 2]]) # one row, three featuresThe outer list is the batch, the inner list is the sample.
4. Note the direction of each reshape.
a.reshape(-1, 1) # column: many samples, one feature
a.reshape(1, -1) # row: one sample, many featuresPicking the wrong one runs without error and gives nonsense results, which is worse. Check X.shape — the first number must equal your sample count.
5. Targets are the exception: y should stay 1D. If you see a DataConversionWarning about a column-vector y, flatten it with y.ravel().
How to prevent it
Print X.shape and y.shape before every fit while learning — five seconds, no surprises. Internalise the pair of rules: X is always 2D, y is usually 1D. Prefer df[["col"]] over reshape gymnastics when working from DataFrames.
Related errors
- X has N features, but the model is expecting M
- operands could not be broadcast together — the (N,) versus (N,1) confusion in NumPy form
- NotFittedError: instance is not fitted yet