Transliteration between scripts
Transliteration rewrites text in a different script while keeping the sound, unlike translation which changes the meaning-carrying words.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Transliteration rewrites a word in a different alphabet while keeping how it sounds, without changing what it means.
Think about writing your own name on a foreign visa form that only accepts the Latin alphabet. "प्रणय" becomes "Pranay." The sound survives. The Devanagari letters do not.
That is transliteration: script changes, sound stays, meaning is untouched. It is not translation. "Pranay" was never translated into an English word, because it is not an English word. It is the same word, in different letters.
Why it exists
Many people type Indian languages using an English keyboard. They spell words out phonetically in Latin letters — "namaste," "kaise ho," "dhanyawad." This is not Hinglish code-mixing. The words are still Hindi, written in the wrong alphabet for Hindi.
Search engines, voice assistants and chat apps all need this. They connect romanised typing back to the real Devanagari word. Sometimes they work the other way. A Devanagari word becomes Latin letters, for someone who cannot read the script.
Transliteration systems handle that conversion, in both directions, between scripts that were never designed to line up letter-for-letter.
How it works
Devanagari: प्र ण य (script changes)
| | |
Roman: pra ṇ a (sound is kept as closely as the target alphabet allows)Some sounds in Devanagari have no single obvious Latin letter — retroflex consonants, aspirated sounds, long versus short vowels. Different transliteration schemes (ITRANS, IAST, Harvard-Kyoto) make different choices about how to spell those sounds in Latin letters.
Where you have already seen it
- Typing Hindi with an English keyboard in Google Input Tools or Gboard, where "namaste" turns into "नमस्ते" as you type.
- Passports and visas, where an Indian name gets one official Latin-letter spelling.
- Song lyrics and prayer texts printed in Latin letters for readers who cannot read Devanagari.
Remember this
- Transliteration changes the script, keeps the sound, and never touches the meaning — translation is a different operation entirely.
- Romanised Hindi typing ("kaise ho") is transliteration going one direction; converting it back to Devanagari is the other.
- No single Latin spelling is "the" correct one — several competing schemes exist, and they disagree on tricky sounds.
What to learn next
- Normalising Indic scripts — cleaning up the Devanagari text transliteration produces.
- Hinglish and code-mixed text — the related, harder problem of language mixed with script.
- Tokenization — how a model would chop the transliterated text that comes out of this step.
Developer — Code and libraries.
Setup
pip install indic-transliterationThis is a pure-Python package — no model download, no network call after installation.
Converting between Devanagari and Roman letters
from indic_transliteration import sanscript
from indic_transliteration.sanscript import transliterate
hindi = "नमस्ते दुनिया" # "namaste duniya" -- hello, world
roman = transliterate(hindi, sanscript.DEVANAGARI, sanscript.ITRANS)
print("Devanagari -> ITRANS:", roman)
back = transliterate("namaste duniya", sanscript.ITRANS, sanscript.DEVANAGARI)
print("ITRANS -> Devanagari:", back)
mixed_name = transliterate("mera naam Pranay hai", sanscript.ITRANS, sanscript.DEVANAGARI)
print("A name inside the sentence:", mixed_name)Devanagari -> ITRANS: namaste duniyA ITRANS -> Devanagari: नमस्ते दुनिय A name inside the sentence: मेर नाम् Pरनय् है
Line by line
"duniyA" has a capital A, not a typo. ITRANS uses capital letters for long vowels — lowercase a is a short "uh" sound, capital A is a long "aa" sound. This is a real convention, not decoration.
The round trip lost something. "नमस्ते दुनिया" became "namaste duniyA," but converting "namaste duniya" (lowercase a, as most people would actually type it) back only produced "नमस्ते दुनिय" — missing the final vowel. Casual typists rarely bother with ITRANS's capitalisation rules, and the conversion degrades when they don't.
The proper noun broke badly. "Pranay," an English-spelled Indian name, does not follow ITRANS conventions at all. The converter partially transliterated it into "Pरनय्" — a genuine mix of Latin and Devanagari inside one broken word. Proper nouns are one of the hardest cases in transliteration, for exactly this reason.
Common mistakes
Assuming one romanisation scheme covers casual typing. ITRANS, IAST and Harvard-Kyoto are formal academic schemes with strict capitalisation and diacritic rules. Nobody types "duniyA" on their phone — they type "duniya" — and formal schemes were not built to tolerate that.
Running transliteration on proper nouns without a name list. As shown above, real names frequently break formal transliteration rules. Production systems usually keep a dictionary of known names to bypass automatic conversion entirely.
Treating transliteration output as guaranteed correct Devanagari. It is a best-effort mapping, not a spelling authority. Always sanity-check output against a dictionary before using it downstream.
Try it yourself
Take a sentence you would actually type in romanised Hindi on your phone — informal spelling, no capital letters for long vowels — and run it through transliterate. Compare the output to what you would have written by hand in Devanagari, and note every place they disagree.
What to learn next
- Normalising Indic scripts — fixing the kind of Unicode inconsistencies transliteration output can introduce.
- The AI4Bharat and IndicNLP stack — a toolkit with purpose-built, model-based transliteration for casual typing.
- Unicode, UTF-8 and mojibake — the character encoding rules underneath every script conversion.
Researcher — Mathematics and papers.
Rule-based versus learned transliteration
The demo above is rule-based: a fixed lookup table maps each Devanagari grapheme to a Latin sequence, applied deterministically. This is exact and fast, but brittle wherever a language's spelling conventions diverge from the scheme's assumptions — informal romanisation, proper nouns, and any Devanagari sequence the table's author did not anticipate.
Learned transliteration reframes the task as a sequence-to-sequence problem: a character-level encoder-decoder, trained on pairs of native-script and romanised words, predicting the target script character by character. Given enough paired data, this handles informal spelling variation directly, because the model has seen many casual, inconsistent spellings during training rather than expecting one canonical form.
P(y_1, ..., y_m | x_1, ..., x_n) = product over j = 1..m of P(y_j | y_{<j}, x_1, ..., x_n)x_1, ..., x_nare source-script characters,y_1, ..., y_mtarget-script characters.- The product is the standard autoregressive character-generation factorisation, identical in form to any sequence-to-sequence model — see neural machine translation for the same factorisation at the word or subword level.
Why proper nouns are the hard case
Formal transliteration schemes (IAST, ITRANS, Harvard-Kyoto) were designed for Sanskrit and classical text, where spelling is comparatively regular. Modern proper nouns frequently violate the phonetic assumptions those schemes encode — English loanwords transcribed into Devanagari, family names with non-standard spellings, brand names invented outside any linguistic convention.
Google's transliteration research (Jia et al., 2013, on Indic language input methods) and subsequent production systems handle this with a hybrid approach: a learned model for general text, backed by a large dictionary of known named entities that bypass the model and get looked up directly.
Key references
- Kunchukuttan, A. & Bhattacharyya, P. (2020). Utilizing Language Relatedness to Improve Machine Translation: A Case Study on Languages of the Indian Subcontinent. arXiv:2003.08925 — covers script relatedness across the Brahmic family, foundational to Indic transliteration tooling.
- Kunchukuttan, A., Puduppully, R. & Bhattacharyya, P. (2015). Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinent. NAACL Demo. — the system behind much of the open-source Indic transliteration tooling referenced in the AI4Bharat and IndicNLP stack.
- Jia, Y. et al. (2013). Devanagari to Roman Transliteration. Proceedings of the workshop on South and Southeast Asian NLP.
Current state and open problems
Deep-learning transliteration models now outperform rule-based schemes on informal input by a wide margin, precisely because they learn from real, messy typed data instead of formal linguistic rules. AI4Bharat's IndicXlit is the current widely-used open model for this in the Indian-language setting.
The unresolved case remains proper nouns and code-switched text together — a sentence that mixes a formally-spelled brand name with informally-romanised Hindi grammar, which is exactly the shape of real chat and search-query text. No current system handles that combination reliably without a maintained name dictionary sitting alongside the learned model.
What to learn next
- How neural machine translation works — the same autoregressive generation pattern, at word or subword granularity instead of characters.
- The AI4Bharat and IndicNLP stack — where IndicXlit and related tools actually live.
- Normalising Indic scripts — the Unicode-level cleanup that usually follows transliteration.