Multilingual and Indic NLP

Normalising Indic scripts

Two pieces of Devanagari text can look identical and still be stored as different bytes, and normalisation is the step that makes them match again.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Script normalisation makes two identical-looking pieces of text actually match, even when typed or saved in different ways.

Think about writing the number "100" versus "one hundred." Same value, different representation. A cash register that only recognises one form will refuse the other. A human sees no difference at all.

Devanagari and other Indic scripts have this problem inside single letters, not only whole words. The same visible character can be stored as different underlying data. It depends on which keyboard or software typed it.

Why it exists

Unicode, the standard that assigns every character a number, allows more than one way to represent certain Indic letters. Take a consonant with a nukta mark — the small dot under a letter, changing its sound, as in क़. It can be stored as one single combined character. Or it can be stored as the plain consonant plus a separate nukta mark, typed right after it.

Both render identically on screen. To a computer comparing raw text, they are different strings. A search, a spellcheck, or a machine-learning model may quietly treat them as two unrelated words.

Normalisation picks one consistent representation and converts everything to it, so identical-looking text is actually identical underneath.

How it works

  Typed one way:    क  +  ़  (nukta typed separately)     -->  two Unicode characters
  Typed another way: क़  (nukta built into one character)  -->  one Unicode character

  Both look the same on screen.
  Normalisation converts both to one agreed form, so string comparison works correctly.

Where you have already seen it

  • Search boxes on Indian-language websites failing to find a page. The query and the page text encoded the same word differently.
  • Spellcheckers flagging correctly-spelled Devanagari words as errors, because the dictionary stored a different Unicode form.
  • Duplicate detection on multilingual forums missing genuine duplicates, because visually identical posts were typed through different input methods.

Remember this

  • The same visible Indic character can exist as more than one sequence of Unicode code points.
  • This is invisible on screen — it only matters when a computer compares, searches or counts text.
  • Normalisation converts every representation of the same character to one standard form before any further processing.

What to learn next

Developer — Code and libraries.

Two demos: Python's built-in Unicode normalisation on a Devanagari nukta case, and the AI4Bharat toolkit's dedicated Indic normaliser confirming the same result.

Setup

bash
pip install indic-nlp-library

unicodedata needs no install — it ships with Python.

The same visible letter, two different byte sequences

nukta_normalise.py
import unicodedata

# Two ways to write the same visible letter "क़" (qa, ka with a nukta):
decomposed = chr(0x0915) + chr(0x093C)   # KA, then a separate combining nukta mark
precomposed = chr(0x0958)                 # QA, one single pre-combined character

print("Do they compare equal as raw strings?", decomposed == precomposed)
print("Decomposed code points: ", [hex(ord(c)) for c in decomposed])
print("Precomposed code points:", [hex(ord(c)) for c in precomposed])

for form in ["NFC", "NFD", "NFKC", "NFKD"]:
    a = unicodedata.normalize(form, decomposed)
    b = unicodedata.normalize(form, precomposed)
    print(f"{form}: equal after normalising? {a == b}")
Output
Do they compare equal as raw strings? False
Decomposed code points:  ['0x915', '0x93c']
Precomposed code points: ['0x958']
NFC: equal after normalising? True
NFD: equal after normalising? True
NFKC: equal after normalising? True
NFKD: equal after normalising? True

Line by line

Every normalisation form agrees here, which is not guaranteed for every Unicode character. Devanagari nukta letters like this one are marked in Unicode as "composition exclusions" — even NFC, which normally merges separate marks back into one combined character, leaves this one split apart. All four forms converge on the decomposed, two-character version.

This is a genuinely surprising detail even for people who know Unicode well. Most developers expect NFC to always produce the shortest, most "combined" form. For this class of Devanagari letters, it deliberately does not.

The practical fix is simple, once you know about it: run unicodedata.normalize("NFC", text) — or any of the four forms, since they agree here — on all incoming text before comparing, searching or tokenizing it.

The AI4Bharat normaliser, for comparison

indicnlp_normalise.py
from indicnlp.normalize.indic_normalize import IndicNormalizerFactory

normalizer = IndicNormalizerFactory().get_normalizer("hi")

decomposed_word = chr(0x0915) + chr(0x093C) + chr(0x093E)   # क + nukta + aa-matra
precomposed_word = chr(0x0958) + chr(0x093E)                 # क़ + aa-matra

norm_a = normalizer.normalize(decomposed_word)
norm_b = normalizer.normalize(precomposed_word)
print("Equal after AI4Bharat normalisation?", norm_a == norm_b)
Output
Equal after AI4Bharat normalisation? True

Common mistakes

Assuming Unicode normalisation is one universal fix for all Indic-text quirks. It fixes representation differences for the same character. It does not fix whitespace inconsistencies, mixed danda (।) versus period (.) usage, or genuine spelling variation — those need the separate cleanup covered in text normalisation.

Comparing or hashing raw user-typed Indic text without normalising first. Two visibly identical strings from two different keyboards can hash differently, silently breaking deduplication and caching.

Expecting NFC to always shorten text. As shown above, some Devanagari sequences stay multi-character under every normalisation form. Do not assume "normalised" means "one character per visible letter."

Try it yourself

Pick a Devanagari word from any Hindi website and type it two different ways on your keyboard — once using an on-screen Devanagari keyboard, once using a transliteration input method like the one in transliteration between scripts. Compare the raw code points with [hex(ord(c)) for c in word] and check whether they match before and after unicodedata.normalize("NFC", word).

What to learn next

Researcher — Mathematics and papers.

Unicode normalisation forms

Unicode defines four normalisation forms, built from two operations: canonical decomposition (splitting a character into its defined base-plus-marks sequence) and canonical composition (recombining a decomposed sequence back into a single character, where one exists).

text
NFD  = decompose
NFC  = decompose, then recompose
NFKD = decompose using compatibility mappings, not only canonical ones
NFKC = decompose using compatibility mappings, then recompose

Recomposition in NFC and NFKC is governed by the Composition Exclusion Table (Unicode Standard Annex #15), a fixed list of characters that are canonically decomposable but explicitly excluded from automatic recomposition. Most Devanagari nukta letters — U+0958 through U+095F — sit on this list, which is why the developer demo above shows NFC leaving them decomposed rather than recombining them.

Why nukta letters are excluded

Nukta letters represent sounds borrowed into Hindi and related languages from Persian, Arabic and English — क़ (qa), ख़ (kha with aspiration shift), ग़ (ghayn), ज़ (za), फ़ (fa). Regional and individual spelling practice varies on whether writers use the nukta-marked letter or its plain base consonant, since many speakers do not distinguish the sounds in casual speech.

Excluding these from automatic recomposition preserves the author's original choice through round-trip normalisation, rather than silently merging what a writer may have deliberately kept distinct — a linguistically-motivated design decision, not an oversight.

Beyond code-point normalisation

Unicode normalisation addresses representation equivalence at the character level. It does not address a separate, higher-level class of problems specific to Indic scripts:

  • Matra reordering. In several Indic scripts, a vowel sign that appears visually before its consonant is stored in memory after it, per the script's defined logical order — a frequent source of rendering and processing bugs in naively-written text tools.
  • Zero-width joiner and non-joiner usage, controlling conjunct consonant formation, which affects rendering but is easy to apply inconsistently across different input methods.
  • Danda (।) versus Latin full stop (.) usage as sentence-final punctuation, genuinely mixed in real corpora and not a Unicode-representation issue at all — this is corpus-level inconsistency, addressed by the same kind of rule-based cleanup as text normalisation, not by unicodedata.

Key references

  • The Unicode Consortium. Unicode Standard Annex #15: Unicode Normalization Forms. unicode.org/reports/tr15
  • The Unicode Consortium. The Unicode Standard, Chapter 12: South and Central Asia-I — the Devanagari block specification, including the nukta letters and their composition exclusion status.
  • Kunchukuttan, A. (2020). The IndicNLP Library. github.com/anoopkunchukuttan/indic_nlp_library — documents the normalisation rules used in the developer demo above, beyond plain Unicode normalisation.
  • Sharma, P. et al. (2018). Unicode Standardization of Indic Scripts and Challenges. IJCA. — a survey of representation inconsistencies specific to Indic text at web scale.

Current state and open problems

Unicode-level normalisation is a solved, mechanical problem — unicodedata.normalize handles it correctly and has for years. What remains genuinely unresolved is corpus-level consistency: real Indic-language web text mixes danda and period, mixes ZWJ/ZWNJ conventions across sources, and mixes nukta usage by regional and individual preference, none of which any single normalisation pass fixes uniformly.

Toolkit-level normalisers such as the one demonstrated above make defensible, documented choices about these corpus-level inconsistencies, but those choices are not universal standards the way Unicode normalisation forms are — a different toolkit may make a different, equally defensible choice. Pipelines that mix text cleaned by different normalisers should expect residual inconsistency, and treat this as a known, bounded risk rather than an oversight to eliminate entirely.

What to learn next