Normalising Indic scripts
Two pieces of Devanagari text can look identical and still be stored as different bytes, and normalisation is the step that makes them match again.
- 9 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Script normalisation makes two identical-looking pieces of text actually match, even when typed or saved in different ways.
Think about writing the number "100" versus "one hundred." Same value, different representation. A cash register that only recognises one form will refuse the other. A human sees no difference at all.
Devanagari and other Indic scripts have this problem inside single letters, not only whole words. The same visible character can be stored as different underlying data. It depends on which keyboard or software typed it.
Why it exists
Unicode, the standard that assigns every character a number, allows more than one way to represent certain Indic letters. Take a consonant with a nukta mark — the small dot under a letter, changing its sound, as in क़. It can be stored as one single combined character. Or it can be stored as the plain consonant plus a separate nukta mark, typed right after it.
Both render identically on screen. To a computer comparing raw text, they are different strings. A search, a spellcheck, or a machine-learning model may quietly treat them as two unrelated words.
Normalisation picks one consistent representation and converts everything to it, so identical-looking text is actually identical underneath.
How it works
Typed one way: क + ़ (nukta typed separately) --> two Unicode characters
Typed another way: क़ (nukta built into one character) --> one Unicode character
Both look the same on screen.
Normalisation converts both to one agreed form, so string comparison works correctly.Where you have already seen it
- Search boxes on Indian-language websites failing to find a page. The query and the page text encoded the same word differently.
- Spellcheckers flagging correctly-spelled Devanagari words as errors, because the dictionary stored a different Unicode form.
- Duplicate detection on multilingual forums missing genuine duplicates, because visually identical posts were typed through different input methods.
Remember this
- The same visible Indic character can exist as more than one sequence of Unicode code points.
- This is invisible on screen — it only matters when a computer compares, searches or counts text.
- Normalisation converts every representation of the same character to one standard form before any further processing.
What to learn next
- The AI4Bharat and IndicNLP stack — a toolkit with normalisation built in for several Indian languages.
- Unicode, UTF-8 and mojibake — the character-encoding foundation this problem sits on top of.
- Text normalisation: case, punctuation and whitespace — the general-purpose version of this cleanup step.
Developer — Code and libraries.
Two demos: Python's built-in Unicode normalisation on a Devanagari nukta case, and the AI4Bharat toolkit's dedicated Indic normaliser confirming the same result.
Setup
pip install indic-nlp-libraryunicodedata needs no install — it ships with Python.
The same visible letter, two different byte sequences
import unicodedata
# Two ways to write the same visible letter "क़" (qa, ka with a nukta):
decomposed = chr(0x0915) + chr(0x093C) # KA, then a separate combining nukta mark
precomposed = chr(0x0958) # QA, one single pre-combined character
print("Do they compare equal as raw strings?", decomposed == precomposed)
print("Decomposed code points: ", [hex(ord(c)) for c in decomposed])
print("Precomposed code points:", [hex(ord(c)) for c in precomposed])
for form in ["NFC", "NFD", "NFKC", "NFKD"]:
a = unicodedata.normalize(form, decomposed)
b = unicodedata.normalize(form, precomposed)
print(f"{form}: equal after normalising? {a == b}")Do they compare equal as raw strings? False Decomposed code points: ['0x915', '0x93c'] Precomposed code points: ['0x958'] NFC: equal after normalising? True NFD: equal after normalising? True NFKC: equal after normalising? True NFKD: equal after normalising? True
Line by line
Every normalisation form agrees here, which is not guaranteed for every Unicode character. Devanagari nukta letters like this one are marked in Unicode as "composition exclusions" — even NFC, which normally merges separate marks back into one combined character, leaves this one split apart. All four forms converge on the decomposed, two-character version.
This is a genuinely surprising detail even for people who know Unicode well. Most developers expect NFC to always produce the shortest, most "combined" form. For this class of Devanagari letters, it deliberately does not.
The practical fix is simple, once you know about it: run unicodedata.normalize("NFC", text) — or any of the four forms, since they agree here — on all incoming text before comparing, searching or tokenizing it.
The AI4Bharat normaliser, for comparison
from indicnlp.normalize.indic_normalize import IndicNormalizerFactory
normalizer = IndicNormalizerFactory().get_normalizer("hi")
decomposed_word = chr(0x0915) + chr(0x093C) + chr(0x093E) # क + nukta + aa-matra
precomposed_word = chr(0x0958) + chr(0x093E) # क़ + aa-matra
norm_a = normalizer.normalize(decomposed_word)
norm_b = normalizer.normalize(precomposed_word)
print("Equal after AI4Bharat normalisation?", norm_a == norm_b)Equal after AI4Bharat normalisation? True
Common mistakes
Assuming Unicode normalisation is one universal fix for all Indic-text quirks. It fixes representation differences for the same character. It does not fix whitespace inconsistencies, mixed danda (।) versus period (.) usage, or genuine spelling variation — those need the separate cleanup covered in text normalisation.
Comparing or hashing raw user-typed Indic text without normalising first. Two visibly identical strings from two different keyboards can hash differently, silently breaking deduplication and caching.
Expecting NFC to always shorten text. As shown above, some Devanagari sequences stay multi-character under every normalisation form. Do not assume "normalised" means "one character per visible letter."
Try it yourself
Pick a Devanagari word from any Hindi website and type it two different ways on your keyboard — once using an on-screen Devanagari keyboard, once using a transliteration input method like the one in transliteration between scripts. Compare the raw code points with [hex(ord(c)) for c in word] and check whether they match before and after unicodedata.normalize("NFC", word).
What to learn next
- The AI4Bharat and IndicNLP stack — where the normaliser used above lives, alongside other Indic-specific tools.
- Unicode, UTF-8 and mojibake — the encoding layer underneath every example in this lesson.
- Transliteration between scripts — a common source of the representation inconsistency shown here.
Researcher — Mathematics and papers.
Unicode normalisation forms
Unicode defines four normalisation forms, built from two operations: canonical decomposition (splitting a character into its defined base-plus-marks sequence) and canonical composition (recombining a decomposed sequence back into a single character, where one exists).
NFD = decompose
NFC = decompose, then recompose
NFKD = decompose using compatibility mappings, not only canonical ones
NFKC = decompose using compatibility mappings, then recomposeRecomposition in NFC and NFKC is governed by the Composition Exclusion Table (Unicode Standard Annex #15), a fixed list of characters that are canonically decomposable but explicitly excluded from automatic recomposition. Most Devanagari nukta letters — U+0958 through U+095F — sit on this list, which is why the developer demo above shows NFC leaving them decomposed rather than recombining them.
Why nukta letters are excluded
Nukta letters represent sounds borrowed into Hindi and related languages from Persian, Arabic and English — क़ (qa), ख़ (kha with aspiration shift), ग़ (ghayn), ज़ (za), फ़ (fa). Regional and individual spelling practice varies on whether writers use the nukta-marked letter or its plain base consonant, since many speakers do not distinguish the sounds in casual speech.
Excluding these from automatic recomposition preserves the author's original choice through round-trip normalisation, rather than silently merging what a writer may have deliberately kept distinct — a linguistically-motivated design decision, not an oversight.
Beyond code-point normalisation
Unicode normalisation addresses representation equivalence at the character level. It does not address a separate, higher-level class of problems specific to Indic scripts:
- Matra reordering. In several Indic scripts, a vowel sign that appears visually before its consonant is stored in memory after it, per the script's defined logical order — a frequent source of rendering and processing bugs in naively-written text tools.
- Zero-width joiner and non-joiner usage, controlling conjunct consonant formation, which affects rendering but is easy to apply inconsistently across different input methods.
- Danda (।) versus Latin full stop (.) usage as sentence-final punctuation, genuinely mixed in real corpora and not a Unicode-representation issue at all — this is corpus-level inconsistency, addressed by the same kind of rule-based cleanup as text normalisation, not by
unicodedata.
Key references
- The Unicode Consortium. Unicode Standard Annex #15: Unicode Normalization Forms. unicode.org/reports/tr15
- The Unicode Consortium. The Unicode Standard, Chapter 12: South and Central Asia-I — the Devanagari block specification, including the nukta letters and their composition exclusion status.
- Kunchukuttan, A. (2020). The IndicNLP Library. github.com/anoopkunchukuttan/indic_nlp_library — documents the normalisation rules used in the developer demo above, beyond plain Unicode normalisation.
- Sharma, P. et al. (2018). Unicode Standardization of Indic Scripts and Challenges. IJCA. — a survey of representation inconsistencies specific to Indic text at web scale.
Current state and open problems
Unicode-level normalisation is a solved, mechanical problem — unicodedata.normalize handles it correctly and has for years. What remains genuinely unresolved is corpus-level consistency: real Indic-language web text mixes danda and period, mixes ZWJ/ZWNJ conventions across sources, and mixes nukta usage by regional and individual preference, none of which any single normalisation pass fixes uniformly.
Toolkit-level normalisers such as the one demonstrated above make defensible, documented choices about these corpus-level inconsistencies, but those choices are not universal standards the way Unicode normalisation forms are — a different toolkit may make a different, equally defensible choice. Pipelines that mix text cleaned by different normalisers should expect residual inconsistency, and treat this as a known, bounded risk rather than an oversight to eliminate entirely.
What to learn next
- The AI4Bharat and IndicNLP stack — the full toolkit this lesson's second demo is drawn from.
- Unicode, UTF-8 and mojibake — the general Unicode and encoding foundations this lesson builds on.
- Why Hindi costs three times more tokens than English — a downstream cost of inconsistent Indic text representation.