Error database

ParserError: Error tokenizing data. C error: Expected N fields (pandas)

A row in your CSV has more columns than the header suggested — wrong delimiter, unquoted commas, or junk lines before the real header. Look at the raw line the message names.

The message you saw
ParserError: Error tokenizing data. C error: Expected N fields (pandas)

By Updated

The error

Output
pandas.errors.ParserError: Error tokenizing data. C error: Expected 5 fields in line 12, saw 7

What it means

Pandas read the header, counted 5 columns, and then met a line that splits into 7 pieces. A CSV is a fragile contract: every row must produce the same number of fields under the same splitting rules. Line 12 broke the contract — or, more often, pandas is splitting with the wrong rules.

Why it happens

Four causes account for nearly every case:

  • Wrong delimiter. The file uses ; or tab, and stray commas inside values create phantom fields.
  • Unquoted text containing commas. Sharma, Priya in an unquoted field becomes two fields.
  • Junk before the real header. Export tools love writing a title line and a date line first. Pandas reads the short title line as the header and every real row looks too wide.
  • Genuinely broken rows. Log files with occasional garbage, or files truncated mid-write.

How to fix it

1. Look at the raw file, especially the named line. This tells you which cause you have.

python
with open("export.csv", encoding="utf-8") as f:
    for i, line in enumerate(f, 1):
        if i <= 3 or i == 12:
            print(i, repr(line))

repr makes tabs and stray quotes visible.

2. Set the real delimiter. Semicolon files are standard in much of Europe, where the comma is the decimal separator.

python
df = pd.read_csv("export.csv", sep=";")
df = pd.read_csv("export.tsv", sep="\t")

3. Skip the junk lines above the real header.

python
df = pd.read_csv("export.csv", skiprows=2)

4. For occasional broken rows, skip them — and count what you lost.

python
df = pd.read_csv("logs.csv", on_bad_lines="skip")
print(len(df), "rows loaded")

Compare the count against the file's line count before trusting the result. Silently dropping 40% of a file is worse than the error.

5. For messy quoting or multi-character separators, use the Python engine.

python
df = pd.read_csv("export.csv", engine="python", sep=None)   # sniff the delimiter

It is slower but far more forgiving than the default C parser.

How to prevent it

Open the first lines of any new file in a text editor before writing loader code — a five-second look answers delimiter, header and quoting at once. When you export CSVs yourself, always quote text fields. For data that round-trips between systems, prefer Parquet, which carries its schema and has no parsing ambiguity.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors