Error database

UnicodeDecodeError: 'utf-8' codec can't decode byte (read_csv)

The file is not UTF-8 — usually a cp1252 export from Excel or an old Windows system. Pass the right encoding to read_csv, or detect it first.

The message you saw
UnicodeDecodeError: 'utf-8' codec can't decode byte (read_csv)

By Updated

The error

Output
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 2892: invalid continuation byte

The byte value varies: 0xe9, 0x92, 0xff. The meaning does not.

What it means

Text files are bytes plus an encoding — the rulebook mapping bytes to characters. Pandas assumed the modern default, UTF-8, and hit a byte sequence that is illegal under UTF-8 rules. The file was written with a different rulebook. Nothing is corrupted; you are reading French, German or symbol characters with the wrong decoder.

Byte 0xe9 is the letter é in cp1252, the classic Windows/Excel encoding. That single clue solves most cases.

Why it happens

Excel's "CSV" export on Windows traditionally writes cp1252, not UTF-8. Files from older databases, government portals and legacy systems come in cp1252, latin-1, or regional encodings. Any accented name or currency symbol in the data becomes an illegal UTF-8 byte.

A special case: 0xff in position 0 usually means a UTF-16 file — or not a text file at all (opening an Excel .xlsx with read_csv gives exactly this).

How to fix it

1. Try cp1252 first for anything that touched Excel or Windows.

python
df = pd.read_csv("export.csv", encoding="cp1252")

2. If that fails, detect the encoding instead of guessing.

bash
pip install charset-normalizer
python
from charset_normalizer import from_path
print(from_path("export.csv").best().encoding)
Output
cp1252

Then pass that value to read_csv.

3. latin-1 always "works" — use it knowingly. Latin-1 maps every possible byte to some character, so it never raises. Wrong guesses show up later as mangled text (é where é should be), not as errors. It is a fallback, not a diagnosis.

4. For a damaged file, replace the bad bytes and move on.

python
df = pd.read_csv("export.csv", encoding_errors="replace")

Each unreadable byte becomes the � character. Acceptable for exploration; not for data you ship.

5. If the first column name starts with , use utf-8-sig. That is a UTF-8 BOM (byte-order mark), a marker Excel adds at the start of the file.

python
df = pd.read_csv("export.csv", encoding="utf-8-sig")

6. If the file is actually .xlsx, use the right reader.

python
df = pd.read_excel("export.xlsx")

How to prevent it

Ask data providers for UTF-8 exports — modern Excel offers "CSV UTF-8" explicitly. Record the encoding of every external source next to its loader code. Convert files to UTF-8 once at the boundary of your project, so everything downstream can assume one encoding.

The lessons behind this error.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors