Multilingual and Indic NLP

Hinglish and code-mixed text

Code-mixed text switches between two languages inside a single sentence, and most language tools were never built to expect that.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Code-mixed text is one sentence built from two languages at once, the way most bilingual people actually talk.

Think about a conversation with a friend in Bengaluru: "yaar aaj traffic bahut zyada tha." Hindi and English arrive in the same breath, in the same sentence. Nobody planned that. It is how bilingual speech comes out.

Hinglish mixes Hindi and English, usually written in the Latin alphabet. It is the most common form of this on Indian social media, chat and reviews. Most NLP tools assume one sentence uses one language. Code-mixed text breaks that assumption constantly.

Why it exists

Every language-processing tool needs to know which language it is looking at. It has to, before it can pick a tokenizer, a spellchecker, a translation direction.

Code-mixing is not rare or accidental. Large-scale studies of Indian social media agree: most posts mix two or more languages. Sometimes the mixing happens inside a single word. A tool built to "detect the language, then apply single-language processing" fails at the first step. Often, there is no single language to detect.

Tools built specifically for code-mixed text handle this differently. Some tag each word with its own language. Others skip language detection and work directly on the mixed sequence.

How it works

  "yaar aaj traffic bahut zyada tha, I was late"
     |     |    |       |      |    |    \    \
    HI    HI   EN       HI    HI   HI    EN    EN

  Whole-sentence language detection sees:  mostly Latin script  -->  guesses "English"
                                             (wrong — it is mostly Hindi, romanised)

The trap: Hinglish is written in Latin letters. A detector that only looks at script — which alphabet is used — cannot tell Hindi-in-Latin-letters from English at all.

Where you have already seen it

  • YouTube and Instagram comments on Indian content, switching languages mid-sentence without a second thought.
  • Customer support chats, where "mera order abhi tak deliver nahi hua" mixes Hindi grammar with an English noun.
  • Voice assistants stumbling on "Alexa, kal ka weather batao" — grammatically Hindi, with one English word inside it.

Remember this

  • Code-mixed text uses two languages inside one sentence, often inside one clause.
  • Hinglish is usually written in Latin letters, which defeats language detectors that only look at script.
  • Real code-mixed processing tags language per word, rather than assuming one language per sentence.

What to learn next

Developer — Code and libraries.

Two demos: whole-sentence language detection failing on Hinglish, then a simple per-word script tagger that catches genuine script mixing.

Setup

bash
pip install langdetect

Whole-sentence detection fails on Hinglish

langdetect_fails.py
from langdetect import detect, detect_langs, DetectorFactory

DetectorFactory.seed = 0   # makes results repeatable across runs

samples = [
    "Yaar this movie was mast, I loved it!",
    "Kal mausam bahut accha tha, we went for a walk.",
    "This is a fully English sentence with no mixing at all.",
]
for s in samples:
    print(f"{detect(s)}  {detect_langs(s)}  {s}")
Output
en  [en:0.9999979578727662]  Yaar this movie was mast, I loved it!
en  [en:0.8571404109979762, so:0.14285785185815808]  Kal mausam bahut accha tha, we went for a walk.
en  [en:0.9999969960500934]  This is a fully English sentence with no mixing at all.

The first sentence is mostly Hindi, written in Latin letters. langdetect reports "en" with near-total confidence. It has no way to know otherwise — every character it sees is an ordinary Latin letter.

Per-word script tagging, for genuinely mixed scripts

script_tag.py
import re

def tag_scripts(text):
    tags = []
    for word in text.split():
        letters = re.sub(r"[^\w]", "", word)
        if not letters:
            continue
        devanagari = sum(1 for c in letters if "ऀ" <= c <= "ॿ")
        latin = sum(1 for c in letters if "a" <= c.lower() <= "z")
        if devanagari and latin:
            tag = "MIXED"
        elif devanagari:
            tag = "HI"
        elif latin:
            tag = "EN/ROMAN"
        else:
            tag = "OTHER"
        tags.append((word, tag))
    return tags

mixed_script = "aaj मीटिंग है, please कल तक भेज दो the report"
for word, tag in tag_scripts(mixed_script):
    print(f"{word:12} {tag}")
Output
aaj          EN/ROMAN
मीटिंग       HI
है,          HI
please       EN/ROMAN
कल           HI
तक           HI
भेज          HI
दो           HI
the          EN/ROMAN
report       EN/ROMAN

Line by line

detect_langs returns a probability distribution, not one flat guess. For the second sentence it briefly considers Somali ("so") at 14%, which shows the model grasping at straws — it has no real signal for anything but English here.

The Unicode range ऀ to ॿ is the Devanagari block. Any character in that range is Hindi, Marathi or another Devanagari-script language. This check is cheap, exact and has nothing to do with a language model — it is pure script detection.

Script tagging catches real script-mixing perfectly, because "मीटिंग" cannot be mistaken for English. It is powerless on the first sentence's Hinglish, because "yaar", "mast" and "accha" are ordinary Latin letters with no script signal at all.

Common mistakes

Treating "language detection" and "script detection" as the same problem. They are not. Script detection is close to perfect and needs no model. Language detection on romanised text needs actual word-level modelling, trained on romanised Hindi specifically.

Running a single-language spellchecker or POS tagger on Hinglish text directly. It will confidently mislabel Hindi-origin words dressed in Latin letters as misspelled English, and correct them into nonsense.

Assuming Hinglish patterns generalise to other code-mixed pairs. Tamil-English and Bengali-English code-mixing follow different switching patterns. A rule tuned for Hinglish will not transfer cleanly.

Try it yourself

Write three sentences that mix Hindi and English at different points — one that starts Hindi and ends English, one the reverse, one that alternates every few words. Run tag_scripts on a Devanagari version of each and see whether the switch points match where you intended them.

What to learn next

Researcher — Mathematics and papers.

Formalising code-switching

Given a sentence s = (w_1, ..., w_n), code-mixed processing assigns each word a language label l_i in L, where L is the set of languages in play — commonly {HI, EN, MIXED, NE} for Hindi-English data, with NE marking named entities that are language-ambiguous by nature (a brand name, a person's name).

This is framed as sequence labelling, structurally identical to named entity recognition: one label per token, predicted jointly rather than independently, since adjacent tokens' languages are correlated (switches cluster at clause boundaries more often than mid-word).

Why script does not solve it

For script-mixed text (Devanagari plus Latin in one sentence), a per-character Unicode range lookup is a near-perfect language-ID signal, as the developer demo shows directly.

Romanised code-mixing removes this signal entirely. The word "gaya" is Hindi ("went") transliterated to Latin script, indistinguishable at the character level from an English proper noun. Language identification here requires a model trained on labelled romanised code-mixed data, typically a character or subword-level sequence tagger (Bhat et al., 2018), since word-level dictionaries fail on transliteration spelling variation — "gaya", "gya" and "gaia" might all appear for the same word.

The Matrix Language Frame

Linguistically, code-switching is not random. The Matrix Language Frame model (Myers-Scotton, 1993) describes one language as the grammatical "matrix" supplying sentence structure, with the other contributing inserted content words — usually nouns — into that frame. "mera order abhi tak deliver nahi hua" has Hindi as the matrix (verb morphology, word order, negation) with "order" and "deliver" inserted from English. This asymmetry is why naive word-count-based language detection ("mostly Hindi words, so it's Hindi") often gets the functional language wrong even when it counts correctly.

Datasets and shared tasks

The ICON and CALCS shared tasks (Molina et al., 2016; Solorio et al., 2014) established standard benchmarks for Hindi-English, Spanish-English and Modern Standard Arabic-dialectal code-switching, covering language identification, POS tagging and named entity recognition on code-mixed data. Performance on these tasks consistently trails equivalent monolingual benchmarks by a wide margin, even for well-resourced language pairs like Hindi-English.

Key references

  • Myers-Scotton, C. (1993). Duelling Languages: Grammatical Structure in Codeswitching. Oxford University Press. — the Matrix Language Frame model.
  • Solorio, T. et al. (2014). Overview for the First Shared Task on Language Identification in Code-Switched Data. Proceedings of CodeSwitch.
  • Molina, G. et al. (2016). Overview for the Second Shared Task on Language Identification in Code-Switched Data. Proceedings of CodeSwitch.
  • Bhat, I. et al. (2018). Universal Dependency Parsing for Hindi-English Code-Switching. NAACL. — sequence-model approaches to romanised Hindi-English tagging.
  • Khanuja, S. et al. (2020). GLUECoS: An Evaluation Benchmark for Code-Switched NLP. arXiv:2004.12376

Current state and open problems

Large multilingual LLMs handle code-mixed input noticeably better than earlier pipeline-based systems, since they were pretrained on web text that already contains real code-mixed sentences, rather than being trained only on clean monolingual corpora. They are not solved on it — informal romanised code-mixed text remains one of the more reliable ways to degrade an LLM's output quality, particularly for anything requiring precise grammatical structure.

Transliteration normalisation before processing — converting romanised Hindi to Devanagari first, covered in transliteration between scripts — is a common practical mitigation, though it introduces its own error rate and does not fully resolve the underlying ambiguity.

What to learn next