Text Preprocessing

Unicode, UTF-8 and mojibake

Text is numbers wearing a costume — Unicode assigns the numbers, UTF-8 packs them into bytes, and mojibake is what unpacking with the wrong rule looks like.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Unicode gives every character in every language its own number.

UTF-8 is the packing rule that turns those numbers into bytes. Mojibake is the garbage you see when text is unpacked with the wrong rule.

Think of a parcel service. Unicode is the address book: every character has one house number, worldwide. UTF-8 is the packing style for shipping. Mojibake — café where café should be — is a parcel unpacked by someone using the wrong manual.

Why it exists

Computers store only numbers. So the world needed an agreement: which number means which character. Early on, every region invented its own small table. The same number meant one letter in Greece and a different one in Russia. Files crossing borders turned to soup.

Unicode ended the chaos with one giant shared table — over 150,000 characters, from ₹ to emoji, each with one number, called a code point.

Numbers still need to become bytes on disk. UTF-8 is the dominant packing rule: common English characters take one byte, and ₹, devanagari and emoji take two to four. Nearly the whole web ships as UTF-8 today.

How it works

The trouble starts when packing and unpacking disagree:

"café"  → packed with UTF-8 →  4 characters become 5 bytes
                                (é takes two bytes)

unpacked with UTF-8      → "café"     correct
unpacked with an old
one-byte-per-character
table (Latin-1)          → "café"    é's two bytes read as two characters

That two-characters-where-one-belongs pattern — à followed by something — is the fingerprint of mojibake, the Japanese word for "transformed characters". You have seen it in subtitles, emails and CSV files: ’ where an apostrophe belongs, ₹ where ₹ belongs.

A real example you have seen

Downloaded subtitles where every apostrophe shows as ’. A bank statement CSV that opens in a spreadsheet with names full of à characters. An old SMS where an emoji became empty boxes. All the same disease: bytes packed one way, unpacked another.

Remember this

  • Unicode numbers the characters; UTF-8 packs the numbers into bytes.
  • Mojibake = unpacked with the wrong rule. à and ’ are its fingerprints.
  • The fix is knowing (or detecting) the true packing rule — not deleting the weird characters.

What to learn next

  • Stopwords — the next cleanup decision after the bytes are sane.
  • Text normalisation — the pass that unifies look-alike characters.
  • Tokenization — where normalisation contracts become part of the model.

Developer — Code and libraries.

Setup

bash
pip install ftfy

Verified with Python 3.10 and ftfy 6.3.1. Everything else is standard library.

Break text on purpose, then fix it

mojibake.py
import ftfy

word = "café"
print("code points:", [hex(ord(c)) for c in word])
print("as UTF-8 bytes:", word.encode("utf-8"))

# the classic accident: UTF-8 bytes read back with the wrong decoder
broken = word.encode("utf-8").decode("latin-1")
print("mojibake:", broken)

# doing it twice is depressingly common in real pipelines
double = broken.encode("utf-8").decode("latin-1")
print("double mojibake:", ascii(double))   # ascii() exposes invisible characters

print("ftfy fixes:", ftfy.fix_text(broken), "/", ftfy.fix_text(double))

# same-looking string, different bytes: é vs e + combining accent
a, b = "caf\u00e9", "cafe\u0301"   # same look, different code points
print("look equal:", a, b, "| == is", a == b)
Output
code points: ['0x63', '0x61', '0x66', '0xe9']
as UTF-8 bytes: b'caf\xc3\xa9'
mojibake: café
double mojibake: 'caf\xc3\x83\xc2\xa9'
ftfy fixes: café / café
look equal: café café | == is False

The walkthrough

é is one code point, two bytes. The list shows four code points; the byte string shows five bytes, with \xc3\xa9 for é. That gap between "how many characters" and "how many bytes" is the whole subject.

.decode("latin-1") cannot fail — that is the trap. Latin-1 maps every possible byte to some character, so wrong decodes never raise errors. They produce confident garbage. An error would have been kinder.

ftfy ("fixes text for you") recognises mojibake fingerprints and reverses the bad decode chain — even the double-mangled case. Notice the double-mangled string contains ƒ, an invisible control character — printed plainly it would hide; ascii() drags it into view. ftfy is the standard rescue tool when the damage already happened upstream.

The last line is a different disease. Two strings display identically but compare unequal: one uses the single character é, the other e plus a combining accent mark. unicodedata.normalize("NFC", s) unifies them — this is the normalisation pass from the previous lesson working at the byte level of meaning.

Common mistakes

Opening files without naming the encoding. open(path) uses a platform-dependent default — UTF-8 on modern Linux and macOS, but often legacy code pages on Windows. Write open(path, encoding="utf-8") every time. Python's UnicodeDecodeError at position 51,842 usually means exactly this.

Deleting characters that error. errors="ignore" makes the crash disappear along with real data — names, currency signs, whole words in non-English text. Prefer errors="replace" while debugging (damage stays visible as �) and a proper encoding fix for production.

Trusting chardet-style detection blindly. Detection is statistical guessing. It works well on long documents and fails on short strings. Check the source's declared encoding first: HTTP headers, HTML meta tags, database column settings.

Confusing str and bytes in Python. str is code points (decoded); bytes is packed data. Mixing them raises TypeError: can't concat str to bytes. Decode at the border when data enters, encode when it leaves, keep everything str in between.

Try it yourself

Reproduce a real-world fingerprint: encode "₹500 off" as UTF-8 and decode with "latin-1". Count how many characters the ₹ became. Then check whether ftfy.fix_text recovers it.

What to learn next

  • Stopwords — the next cleanup decision after the bytes are sane.
  • Text normalisation — the pass that unifies look-alike characters.
  • Tokenization — where normalisation contracts become part of the model.

Researcher — Mathematics and papers.

UTF-8's design

UTF-8 (Thompson and Pike, 1992) encodes code point U in 1–4 bytes: ASCII (U ≤ 0x7F) as itself; larger values with a length-announcing lead byte followed by continuation bytes of the form 10xxxxxx. Three properties made it win: ASCII transparency (every ASCII file is already valid UTF-8), self-synchronisation (a decoder can find the next character boundary from any byte, since lead and continuation bytes are disjoint ranges), and sorted-order preservation of code points under byte-wise comparison. Overlong encodings (encoding a value in more bytes than needed) are invalid by specification — historically a security vector for filter bypass (CVE-class: overlong ../), which is why strict decoders reject them.

Surrogates (U+D800–DFFF) exist for UTF-16's pairing mechanism and are illegal in UTF-8 interchange; "WTF-8" (Sapin, 2014) is the deliberate superset used internally by systems bridging ill-formed UTF-16 (Windows filenames, JavaScript strings).

Mojibake, formally

Mojibake is composition of encode/decode functions that are not mutual inverses: d_B ∘ e_A with B ≠ A. Because single-byte codecs (Latin-1, cp1252) are total functions on bytes, the composition is injective and therefore reversible — the basis of ftfy's approach (Speer, 2019): search a small space of plausible codec chains, score candidates by character-class plausibility heuristics, and invert. cp1252 versus Latin-1 differ in the 0x80–0x9F range (cp1252 assigns printable characters — €, ™, curly quotes — where Latin-1 has control codes), which is why ’ (bytes E2 80 99 of U+2019 read as cp1252) is the single most common fingerprint in English web text.

Encoding detection (chardet, charset-normalizer, ICU's detector) is Bayesian classification over byte n-gram statistics per codec-language pair — reliable above a few kilobytes, weak on tweets; the W3C encoding standard therefore mandates header/meta declarations take precedence.

Normalisation forms and security

UAX #15 defines NFC/NFD (canonical) and NFKC/NFKD (compatibility). Canonical equivalence covers composed é (U+00E9) versus decomposed e+U+0301; compatibility additionally folds fi, fullwidth forms, superscripts. NFC is the web's de-facto storage form; NFD appears in macOS filenames — a classic cross-platform bug source. Security work cares because visually-identical strings with different code points enable homograph attacks (Cyrillic а vs Latin a in domains — IDN, UTS #39 confusables) and normalisation-bypass injection; the mitigation is normalising then validating, in that order, at trust boundaries.

For ML pipelines: tokenizers pin a normalisation form as part of the model contract (BERT's uncased preprocessing applies NFD + accent stripping; SentencePiece defaults to NMT-flavoured NFKC). Changing normalisation between training and inference is silent distribution shift — the tokenization lesson picks this up.

What to learn next

  • Stopwords — the next cleanup decision after the bytes are sane.
  • Text normalisation — the pass that unifies look-alike characters.
  • Tokenization — where normalisation contracts become part of the model.