Text Preprocessing

Detecting the language of a text

Character patterns give every language a fingerprint a small model can read — reliably on paragraphs, shakily on tweets, and wrongly on romanised Hindi.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Language detection is a small model that reads a text's letter patterns — its fingerprint — and names the language, before any other processing runs.

Think of hearing two neighbours talk through a wall. You cannot make out one word, yet you know within seconds whether it is Telugu, Hindi or English. The rhythm and sounds give it away. Written language leaks the same way: which letters appear, and which letter pairs follow each other, form a fingerprint.

Why it exists

Every tool downstream of this decision is language-specific. Stopword lists are per-language. Sentence rules are per-language. spaCy models, translation systems, even keyboards — all need to know what they are reading first.

Get it wrong and everything downstream fails at once: English stopwords sifting a German email, an English parser shredding a French sentence. Detection is the routing desk of a multilingual pipeline — cheap, early, and load-bearing.

How it works

The classic method counts short letter sequences. "th", "he", "ing" scream English. "sch" and "der" lean German. "ção" is Portuguese's autograph. A text's counts get compared against each language's known profile, and the closest profile wins.

"Der Zug nach Berlin"
   │  count letter pairs: de, er, zu, ug, na, ch ...
   │  compare with stored fingerprints
   ├── German profile   → very close  ✓
   ├── English profile  → far
   └── French profile   → far

Longer text means a steadier fingerprint. One or two words — "ok", "amit" — barely leave a mark, and detectors guess. Honest systems return a confidence along with the answer, and honest engineers check it.

A real example you have seen

The "Translate this post?" button appears because a detector already identified the language of what you are reading. Chrome's translation bar, WhatsApp's per-message translate option, and spam filters that route mail by language all run detection silently, millions of times a day.

Remember this

  • Detection reads letter-pattern fingerprints, not meaning.
  • Long text → reliable. Short text → guessy. Check the confidence.
  • It is the first step of a multilingual pipeline, because everything after it is per-language.

What to learn next

  • Stopwords — the first per-language decision this routing enables.
  • What is NLP — the field this routing desk serves.
  • Tokenization — how multilingual models sidestep some of this entirely.

Developer — Code and libraries.

Setup

bash
pip install langdetect

Verified with langdetect 1.0.9 on Python 3.10, CPU only.

Detect, and meet the two classic failures

detect_lang.py
from langdetect import DetectorFactory, detect, detect_langs

DetectorFactory.seed = 0   # langdetect is randomised; pin it for stable answers

samples = [
    "The train to Hyderabad leaves at nine.",
    "Der Zug nach Berlin ist schon weg.",
    "Le train pour Paris est parti.",
    "aap kaise hain bhai",          # Hindi written in Latin letters
    "ok",                           # too short to judge
]
for s in samples:
    print(f"{detect(s):5s} <- {s!r}")

print()
print(detect_langs("The train to Hyderabad leaves at nine."))
Output
en    <- 'The train to Hyderabad leaves at nine.'
de    <- 'Der Zug nach Berlin ist schon weg.'
fr    <- 'Le train pour Paris est parti.'
fi    <- 'aap kaise hain bhai'
sk    <- 'ok'

[en:0.9999950127756445]

The walkthrough

Three clean wins, two instructive faceplants. Full sentences in English, German and French were called correctly. Then: romanised Hindi came back as Finnish, and "ok" as Slovak. Neither is a bug — they are the tool's honest limits, printed.

Why "aap kaise hain bhai" broke it. The model knows Hindi by its Devanagari fingerprint. Hindi written in Latin letters matches no stored profile, so the detector picks whatever European language's letter pairs land closest. Romanised and code-mixed text (Hinglish) is the number-one real-world failure mode for Indian products — test for it explicitly.

Why the seed matters. langdetect adds random noise internally, so the same borderline string can return different languages across runs. DetectorFactory.seed = 0 makes results reproducible — set it once at import time, every time.

detect_langs returns probabilities. A production system should threshold: accept the label above 0.9, else fall back — route to a default, ask the user, or try a stronger model. detect() alone hides exactly the uncertainty you need.

Stronger options when it matters. fastText's lid.176.bin (Facebook, ~130 MB download, 176 languages) is markedly more robust on short text. Google's compact CLD3 ships in Chrome. Both handle more scripts; none solves romanised Hindi well — for that, current practice is a fine-tuned transformer or a dedicated codemix identifier.

Common mistakes

Detecting per short string instead of per document. Running detection on every chat message wastes accuracy where it is weakest. Detect at the largest sensible unit — the document, the conversation — and inherit the label downward.

Trusting the top label without the score. "sk" for "ok" arrived with high internal uncertainty. Thresholding on detect_langs output is two lines of code and prevents entire categories of silent misrouting.

Forgetting mixed-language documents. An English email with a German quotation gets one label, and half the document gets the wrong pipeline. If mixing is common in your data, detect per paragraph and reconcile.

Cleaning too aggressively first. Stripping accents ("café" → "cafe") erases exactly the characters the fingerprint relies on. Detection runs on raw-ish text, before the heavy normalisation passes.

Try it yourself

Feed detect_langs a sentence that is half English and half German, and watch the probability split across both. Then find the shortest English sentence that gets a confidence above 0.99 — you will be surprised how long it needs to be.

What to learn next

  • Stopwords — the first per-language decision this routing enables.
  • What is NLP — the field this routing desk serves.
  • Tokenization — how multilingual models sidestep some of this entirely.

Researcher — Mathematics and papers.

Models

Classical LID is Naive Bayes or nearest-profile over character n-gram statistics: Cavnar and Trenkle (1994), N-gram-based text categorization, established the rank-profile method; Dunning (1994) the likelihood framing. langdetect is a port of Nakatani's (2010) Java library — character n-gram Naive Bayes over Wikipedia priors for 55 languages, with random feature sampling explaining its nondeterminism. langid.py (Lui and Baldwin, 2012) selects features cross-domain-stably via information gain and is the classical academic baseline. fastText LID (Joulin et al., 2016/2017) trains a linear classifier over character n-gram embeddings for 176 languages; CLD3 uses a small feed-forward net over n-gram embeddings. Current research-grade systems (OpenLID — Burchell et al., 2023; GlotLID — Kargaran et al., 2023, 1600+ languages; the NLLB LID covering 200+) push coverage to low-resource languages, where training data quality, not architecture, is the binding constraint.

Error structure

Accuracy is a function of input length: near-ceiling above ~50 characters for high-resource languages in native scripts, collapsing below ~20 (Lui and Baldwin, 2014 quantify the short-text degradation; tweet-length LID was a shared-task genre of its own). Systematic confusions cluster in: (1) close language pairs — Bosnian/Croatian/Serbian, Indonesian/Malay, Czech/Slovak — where lexical overlap defeats character statistics (the VarDial workshop series benchmarks exactly this); (2) romanisation, which moves a language into another script's feature space, as demonstrated above; and (3) code-switching, where the single-label assumption is itself wrong — token-level LID for code-mixed text (Solorio et al., 2014, the first shared task; much subsequent Hinglish work) reframes the problem as sequence labelling. Priors matter operationally: P(language | script) is extremely peaked, so script detection (a cheap Unicode-range check) should gate the classifier — running a 176-way model on pure Devanagari input wastes both compute and error budget.

Cost and deployment

All classical models are O(n) in characters with tiny constants — microseconds per document — so LID runs comfortably at web-corpus scale; the large multilingual crawls (CommonCrawl-derived: CCNet, OSCAR, C4's multilingual variants) apply fastText LID plus a confidence threshold as a primary filter, and their documented misrouting rates for low-resource languages (Kreutzer et al., 2022, Quality at a glance) show detector error compounding into corpus pollution — worth remembering when a pretrained model behaves oddly in a rare language. In-product, thresholded prediction with a human-fallback path outperforms raising model size in most cost-benefit analyses; per-paragraph voting with document-level reconciliation is the standard mixed-document strategy.

What to learn next

  • Stopwords — the first per-language decision this routing enables.
  • What is NLP — the field this routing desk serves.
  • Tokenization — how multilingual models sidestep some of this entirely.

What to learn next

These follow on from what you just read.

  • Classical NLP That Still Works

    Bag of words

    Bag of words turns a sentence into a list of word counts, throwing away word order but keeping enough signal to search and classify text cheaply.

  • Classical NLP That Still Works

    TF-IDF

    TF-IDF scores a word higher when it is frequent in one document but rare across all others, so common words stop drowning out the words that actually matter.

  • Classical NLP That Still Works

    N-gram language models and smoothing

    An n-gram language model predicts the next word from the last few words it saw, and smoothing stops it from calling any unseen sentence completely impossible.