Hinglish and code-mixed text
Code-mixed text switches between two languages inside a single sentence, and most language tools were never built to expect that.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Code-mixed text is one sentence built from two languages at once, the way most bilingual people actually talk.
Think about a conversation with a friend in Bengaluru: "yaar aaj traffic bahut zyada tha." Hindi and English arrive in the same breath, in the same sentence. Nobody planned that. It is how bilingual speech comes out.
Hinglish mixes Hindi and English, usually written in the Latin alphabet. It is the most common form of this on Indian social media, chat and reviews. Most NLP tools assume one sentence uses one language. Code-mixed text breaks that assumption constantly.
Why it exists
Every language-processing tool needs to know which language it is looking at. It has to, before it can pick a tokenizer, a spellchecker, a translation direction.
Code-mixing is not rare or accidental. Large-scale studies of Indian social media agree: most posts mix two or more languages. Sometimes the mixing happens inside a single word. A tool built to "detect the language, then apply single-language processing" fails at the first step. Often, there is no single language to detect.
Tools built specifically for code-mixed text handle this differently. Some tag each word with its own language. Others skip language detection and work directly on the mixed sequence.
How it works
"yaar aaj traffic bahut zyada tha, I was late"
| | | | | | \ \
HI HI EN HI HI HI EN EN
Whole-sentence language detection sees: mostly Latin script --> guesses "English"
(wrong — it is mostly Hindi, romanised)The trap: Hinglish is written in Latin letters. A detector that only looks at script — which alphabet is used — cannot tell Hindi-in-Latin-letters from English at all.
Where you have already seen it
- YouTube and Instagram comments on Indian content, switching languages mid-sentence without a second thought.
- Customer support chats, where "mera order abhi tak deliver nahi hua" mixes Hindi grammar with an English noun.
- Voice assistants stumbling on "Alexa, kal ka weather batao" — grammatically Hindi, with one English word inside it.
Remember this
- Code-mixed text uses two languages inside one sentence, often inside one clause.
- Hinglish is usually written in Latin letters, which defeats language detectors that only look at script.
- Real code-mixed processing tags language per word, rather than assuming one language per sentence.
What to learn next
- Detecting the language of a text — the tool that struggles hardest with code-mixed input.
- Transliteration between scripts — converting the Latin-script Hindi seen here back into Devanagari.
- Tokenization — the step that runs into trouble first when the language assumption breaks.
Developer — Code and libraries.
Two demos: whole-sentence language detection failing on Hinglish, then a simple per-word script tagger that catches genuine script mixing.
Setup
pip install langdetectWhole-sentence detection fails on Hinglish
from langdetect import detect, detect_langs, DetectorFactory
DetectorFactory.seed = 0 # makes results repeatable across runs
samples = [
"Yaar this movie was mast, I loved it!",
"Kal mausam bahut accha tha, we went for a walk.",
"This is a fully English sentence with no mixing at all.",
]
for s in samples:
print(f"{detect(s)} {detect_langs(s)} {s}")en [en:0.9999979578727662] Yaar this movie was mast, I loved it! en [en:0.8571404109979762, so:0.14285785185815808] Kal mausam bahut accha tha, we went for a walk. en [en:0.9999969960500934] This is a fully English sentence with no mixing at all.
The first sentence is mostly Hindi, written in Latin letters. langdetect reports "en" with near-total confidence. It has no way to know otherwise — every character it sees is an ordinary Latin letter.
Per-word script tagging, for genuinely mixed scripts
import re
def tag_scripts(text):
tags = []
for word in text.split():
letters = re.sub(r"[^\w]", "", word)
if not letters:
continue
devanagari = sum(1 for c in letters if "ऀ" <= c <= "ॿ")
latin = sum(1 for c in letters if "a" <= c.lower() <= "z")
if devanagari and latin:
tag = "MIXED"
elif devanagari:
tag = "HI"
elif latin:
tag = "EN/ROMAN"
else:
tag = "OTHER"
tags.append((word, tag))
return tags
mixed_script = "aaj मीटिंग है, please कल तक भेज दो the report"
for word, tag in tag_scripts(mixed_script):
print(f"{word:12} {tag}")aaj EN/ROMAN मीटिंग HI है, HI please EN/ROMAN कल HI तक HI भेज HI दो HI the EN/ROMAN report EN/ROMAN
Line by line
detect_langs returns a probability distribution, not one flat guess. For the second sentence it briefly considers Somali ("so") at 14%, which shows the model grasping at straws — it has no real signal for anything but English here.
The Unicode range ऀ to ॿ is the Devanagari block. Any character in that range is Hindi, Marathi or another Devanagari-script language. This check is cheap, exact and has nothing to do with a language model — it is pure script detection.
Script tagging catches real script-mixing perfectly, because "मीटिंग" cannot be mistaken for English. It is powerless on the first sentence's Hinglish, because "yaar", "mast" and "accha" are ordinary Latin letters with no script signal at all.
Common mistakes
Treating "language detection" and "script detection" as the same problem. They are not. Script detection is close to perfect and needs no model. Language detection on romanised text needs actual word-level modelling, trained on romanised Hindi specifically.
Running a single-language spellchecker or POS tagger on Hinglish text directly. It will confidently mislabel Hindi-origin words dressed in Latin letters as misspelled English, and correct them into nonsense.
Assuming Hinglish patterns generalise to other code-mixed pairs. Tamil-English and Bengali-English code-mixing follow different switching patterns. A rule tuned for Hinglish will not transfer cleanly.
Try it yourself
Write three sentences that mix Hindi and English at different points — one that starts Hindi and ends English, one the reverse, one that alternates every few words. Run tag_scripts on a Devanagari version of each and see whether the switch points match where you intended them.
What to learn next
- Detecting the language of a text — the tool this lesson shows breaking, and why.
- Named entity recognition — another task that gets harder once a sentence stops being one language.
- Unicode, UTF-8 and mojibake — the character-encoding foundation the script-range check above relies on.
Researcher — Mathematics and papers.
Formalising code-switching
Given a sentence s = (w_1, ..., w_n), code-mixed processing assigns each word a language label l_i in L, where L is the set of languages in play — commonly {HI, EN, MIXED, NE} for Hindi-English data, with NE marking named entities that are language-ambiguous by nature (a brand name, a person's name).
This is framed as sequence labelling, structurally identical to named entity recognition: one label per token, predicted jointly rather than independently, since adjacent tokens' languages are correlated (switches cluster at clause boundaries more often than mid-word).
Why script does not solve it
For script-mixed text (Devanagari plus Latin in one sentence), a per-character Unicode range lookup is a near-perfect language-ID signal, as the developer demo shows directly.
Romanised code-mixing removes this signal entirely. The word "gaya" is Hindi ("went") transliterated to Latin script, indistinguishable at the character level from an English proper noun. Language identification here requires a model trained on labelled romanised code-mixed data, typically a character or subword-level sequence tagger (Bhat et al., 2018), since word-level dictionaries fail on transliteration spelling variation — "gaya", "gya" and "gaia" might all appear for the same word.
The Matrix Language Frame
Linguistically, code-switching is not random. The Matrix Language Frame model (Myers-Scotton, 1993) describes one language as the grammatical "matrix" supplying sentence structure, with the other contributing inserted content words — usually nouns — into that frame. "mera order abhi tak deliver nahi hua" has Hindi as the matrix (verb morphology, word order, negation) with "order" and "deliver" inserted from English. This asymmetry is why naive word-count-based language detection ("mostly Hindi words, so it's Hindi") often gets the functional language wrong even when it counts correctly.
Datasets and shared tasks
The ICON and CALCS shared tasks (Molina et al., 2016; Solorio et al., 2014) established standard benchmarks for Hindi-English, Spanish-English and Modern Standard Arabic-dialectal code-switching, covering language identification, POS tagging and named entity recognition on code-mixed data. Performance on these tasks consistently trails equivalent monolingual benchmarks by a wide margin, even for well-resourced language pairs like Hindi-English.
Key references
- Myers-Scotton, C. (1993). Duelling Languages: Grammatical Structure in Codeswitching. Oxford University Press. — the Matrix Language Frame model.
- Solorio, T. et al. (2014). Overview for the First Shared Task on Language Identification in Code-Switched Data. Proceedings of CodeSwitch.
- Molina, G. et al. (2016). Overview for the Second Shared Task on Language Identification in Code-Switched Data. Proceedings of CodeSwitch.
- Bhat, I. et al. (2018). Universal Dependency Parsing for Hindi-English Code-Switching. NAACL. — sequence-model approaches to romanised Hindi-English tagging.
- Khanuja, S. et al. (2020). GLUECoS: An Evaluation Benchmark for Code-Switched NLP. arXiv:2004.12376
Current state and open problems
Large multilingual LLMs handle code-mixed input noticeably better than earlier pipeline-based systems, since they were pretrained on web text that already contains real code-mixed sentences, rather than being trained only on clean monolingual corpora. They are not solved on it — informal romanised code-mixed text remains one of the more reliable ways to degrade an LLM's output quality, particularly for anything requiring precise grammatical structure.
Transliteration normalisation before processing — converting romanised Hindi to Devanagari first, covered in transliteration between scripts — is a common practical mitigation, though it introduces its own error rate and does not fully resolve the underlying ambiguity.
What to learn next
- Transliteration between scripts — the normalisation step often applied before further processing.
- Named entity recognition — the closest published task to the per-word tagging problem described here.
- Detecting the language of a text — the single-language tool this section extends.