ParserError: Error tokenizing data. C error: Expected N fields (pandas)
A row in your CSV has more columns than the header suggested — wrong delimiter, unquoted commas, or junk lines before the real header. Look at the raw line the message names.
Updated
The error
pandas.errors.ParserError: Error tokenizing data. C error: Expected 5 fields in line 12, saw 7
What it means
Pandas read the header, counted 5 columns, and then met a line that splits into 7 pieces. A CSV is a fragile contract: every row must produce the same number of fields under the same splitting rules. Line 12 broke the contract — or, more often, pandas is splitting with the wrong rules.
Why it happens
Four causes account for nearly every case:
- Wrong delimiter. The file uses
;or tab, and stray commas inside values create phantom fields. - Unquoted text containing commas.
Sharma, Priyain an unquoted field becomes two fields. - Junk before the real header. Export tools love writing a title line and a date line first. Pandas reads the short title line as the header and every real row looks too wide.
- Genuinely broken rows. Log files with occasional garbage, or files truncated mid-write.
How to fix it
1. Look at the raw file, especially the named line. This tells you which cause you have.
with open("export.csv", encoding="utf-8") as f:
for i, line in enumerate(f, 1):
if i <= 3 or i == 12:
print(i, repr(line))repr makes tabs and stray quotes visible.
2. Set the real delimiter. Semicolon files are standard in much of Europe, where the comma is the decimal separator.
df = pd.read_csv("export.csv", sep=";")
df = pd.read_csv("export.tsv", sep="\t")3. Skip the junk lines above the real header.
df = pd.read_csv("export.csv", skiprows=2)4. For occasional broken rows, skip them — and count what you lost.
df = pd.read_csv("logs.csv", on_bad_lines="skip")
print(len(df), "rows loaded")Compare the count against the file's line count before trusting the result. Silently dropping 40% of a file is worse than the error.
5. For messy quoting or multi-character separators, use the Python engine.
df = pd.read_csv("export.csv", engine="python", sep=None) # sniff the delimiterIt is slower but far more forgiving than the default C parser.
How to prevent it
Open the first lines of any new file in a text editor before writing loader code — a five-second look answers delimiter, header and quoting at once. When you export CSVs yourself, always quote text fields. For data that round-trips between systems, prefer Parquet, which carries its schema and has no parsing ambiguity.
Related errors
- UnicodeDecodeError: 'utf-8' codec can't decode byte
- KeyError: column not in index — often the sequel, when the wrong line became the header
- ValueError: could not convert string to float