Error database

KeyError: 'column' / "['column'] not in index" (pandas)

The column name you asked for does not exist in the DataFrame. Print df.columns.tolist() and look for whitespace, capitalisation, or a column that became the index.

The message you saw
KeyError: 'column' / "['column'] not in index" (pandas)

By Updated

The error

Output
KeyError: 'price'

Selecting several columns at once produces a wordier variant:

Output
KeyError: "['price', 'tax'] not in index"

And KeyError: 0 appears when you write df[0] hoping for the first column.

What it means

You asked the DataFrame for a column by name, and no column has that exact name. Pandas matches names character for character. "Price", "price " and "price" are three different columns.

Why it happens

Real CSV files have messy headers. A trailing space ("price ") is invisible when you print the frame. Capitalisation differs between files. Sometimes the real header sits on row 3 and pandas took row 1, so your "columns" are metadata text.

Two pandas-specific routes also lead here. A column stops being a column after set_index, groupby or read_csv(..., index_col=...) — it moved into the index. And df[0] fails because plain brackets select by label, not position.

How to fix it

1. Print the real column names. Not the frame — the list, where whitespace shows.

python
print(df.columns.tolist())
Output
['Date', 'price ', 'City name', 'tax%']

Now the mismatch is visible.

2. Clean the headers once, right after loading.

python
df.columns = df.columns.str.strip().str.lower().str.replace(" ", "_")

3. If the column moved into the index, bring it back.

python
df = df.reset_index()

4. If the header row is wrong, tell read_csv where it is.

python
df = pd.read_csv("sales.csv", header=2)     # real header on the third line

5. For positional access, use .iloc.

python
first_col = df.iloc[:, 0]

6. For optional columns, select what is present instead of a fixed list.

python
wanted = ["price", "tax"]
present = [c for c in wanted if c in df.columns]
subset = df[present]

How to prevent it

Normalise column names in the first line after every read_csv — the strip/lower/replace one-liner above. Treat df.columns.tolist() as your first debugging move for any pandas KeyError. When a pipeline consumes files from other people, validate the expected columns up front and fail with a message that names what is missing.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors