UnicodeDecodeError: 'utf-8' codec can't decode byte (read_csv)
The file is not UTF-8 — usually a cp1252 export from Excel or an old Windows system. Pass the right encoding to read_csv, or detect it first.
Updated
The error
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 2892: invalid continuation byte
The byte value varies: 0xe9, 0x92, 0xff. The meaning does not.
What it means
Text files are bytes plus an encoding — the rulebook mapping bytes to characters. Pandas assumed the modern default, UTF-8, and hit a byte sequence that is illegal under UTF-8 rules. The file was written with a different rulebook. Nothing is corrupted; you are reading French, German or symbol characters with the wrong decoder.
Byte 0xe9 is the letter é in cp1252, the classic Windows/Excel encoding. That single clue solves most cases.
Why it happens
Excel's "CSV" export on Windows traditionally writes cp1252, not UTF-8. Files from older databases, government portals and legacy systems come in cp1252, latin-1, or regional encodings. Any accented name or currency symbol in the data becomes an illegal UTF-8 byte.
A special case: 0xff in position 0 usually means a UTF-16 file — or not a text file at all (opening an Excel .xlsx with read_csv gives exactly this).
How to fix it
1. Try cp1252 first for anything that touched Excel or Windows.
df = pd.read_csv("export.csv", encoding="cp1252")2. If that fails, detect the encoding instead of guessing.
pip install charset-normalizerfrom charset_normalizer import from_path
print(from_path("export.csv").best().encoding)cp1252
Then pass that value to read_csv.
3. latin-1 always "works" — use it knowingly. Latin-1 maps every possible byte to some character, so it never raises. Wrong guesses show up later as mangled text (é where é should be), not as errors. It is a fallback, not a diagnosis.
4. For a damaged file, replace the bad bytes and move on.
df = pd.read_csv("export.csv", encoding_errors="replace")Each unreadable byte becomes the � character. Acceptable for exploration; not for data you ship.
5. If the first column name starts with , use utf-8-sig. That is a UTF-8 BOM (byte-order mark), a marker Excel adds at the start of the file.
df = pd.read_csv("export.csv", encoding="utf-8-sig")6. If the file is actually .xlsx, use the right reader.
df = pd.read_excel("export.xlsx")How to prevent it
Ask data providers for UTF-8 exports — modern Excel offers "CSV UTF-8" explicitly. Record the encoding of every external source next to its loader code. Convert files to UTF-8 once at the boundary of your project, so everything downstream can assume one encoding.
Related errors
- ParserError: Error tokenizing data — the other classic read_csv failure
- KeyError: column not in index — what a BOM does to your first column name
- ValueError: could not convert string to float