Multilingual and Indic NLP

The AI4Bharat and IndicNLP stack

AI4Bharat and the IndicNLP library are open tools built specifically for Indian languages, instead of stretching general multilingual tools to cover them.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

AI4Bharat and the IndicNLP library are free tools built specifically for Indian languages. They do not treat those languages as an afterthought.

Think about a tailor who makes clothes for one specific body type. He measures real customers of that build, instead of stitching one universal size for everyone. The tailor's fit is better, because the sizing was never averaged away.

General multilingual tools split their attention across 100 languages. AI4Bharat, a research group based in India, built tools that specialise instead. Its tokenizers, normalisers, transliteration and translation models are trained and tuned specifically for India's languages and scripts.

Why it exists

A model trained on 100 languages at once, like XLM-RoBERTa, spreads its capacity across every one of them. Indian languages get a share of that capacity, not the majority of it. India alone has 22 constitutionally scheduled languages, and hundreds more spoken daily.

AI4Bharat's answer: build the tokenizer, the normaliser, and the models around Indian scripts and languages. Make them the main design target, not one entry in a much longer list. IndicBERT, IndicTrans2, IndicXlit and the IndicNLP library are the resulting pieces, trained and tuned specifically for this language family.

How it works

  General multilingual model:  100 languages sharing one budget of attention

  AI4Bharat stack:              Indian scripts and languages as the primary target,
                                 not one language among a hundred others

The IndicNLP library handles the everyday preprocessing work directly — tokenizing, normalising, transliterating — for a dozen or more Indian scripts. It needs no large pretrained model for these basic steps.

Where you have already seen it

  • IndicTrans2, the translation model already introduced in translating a low-resource language, covering all 22 scheduled Indian languages.
  • Government and civic-tech projects processing Indian-language text, frequently building on this open stack rather than starting from scratch.
  • Academic NLP research on Indian languages, where this stack is close to a default starting point.

Remember this

  • AI4Bharat builds NLP tools specifically for Indian languages. It does not treat them as one entry among many in a general model.
  • The IndicNLP library handles everyday preprocessing — tokenizing, normalising, transliterating — without needing a large model download.
  • Specialised tooling and general multilingual models are complements, not rivals — production systems commonly use both.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install indic-nlp-library

This installs a pure-Python package for basic tokenizing and normalising — no large model download for the pieces used below.

Tokenizing and normalising Hindi with IndicNLP

indicnlp_basics.py
from indicnlp.tokenize import indic_tokenize
from indicnlp.normalize.indic_normalize import IndicNormalizerFactory

text = "मुझे कल सुबह 9:30 बजे मीटिंग है, कृपया याद दिलाना।"

tokens = indic_tokenize.trivial_tokenize(text, lang="hi")
print("tokens:", tokens)

normalizer = IndicNormalizerFactory().get_normalizer("hi")
normalised = normalizer.normalize(text)
print("normalisation changed the text:", normalised != text)
Output
tokens: ['मुझे', 'कल', 'सुबह', '9:30', 'बजे', 'मीटिंग', 'है', ',', 'कृपया', 'याद', 'दिलाना', '।']
normalisation changed the text: False

Line by line

trivial_tokenize splits on script-aware rules, not only whitespace. It correctly keeps "9:30" as one token, and separates the trailing danda (।, the Devanagari sentence-end mark) from "दिलाना" as its own token, the way a general whitespace-splitter would get wrong.

The normaliser reported no change here, because this sentence was already in one consistent Unicode form. Normalising Indic scripts shows the same normaliser actually changing text, on an input deliberately constructed with an inconsistent nukta encoding.

Neither function downloaded a model. Both are rule-based, script-aware code — fast, deterministic, and useful even on a machine with no GPU and no internet connection after installation.

Common mistakes

Reaching for a full pretrained model when a rule-based tool already does the job. Tokenizing and normalising Indic script text does not need IndicBERT or any neural model — the code above handles it directly, and faster.

Using trivial_tokenize for a language it was not told about. The lang argument changes real behaviour — script-specific punctuation and abbreviation handling differ by language, even across languages sharing the Devanagari script.

Assuming this library covers every Indian language equally well. Coverage, tooling maturity and testing depth vary by language inside this stack too, generally tracking the same resource patterns discussed in translating a low-resource language — a widely-spoken, well-resourced language gets more mature tooling than a less-documented one.

Try it yourself

Run indic_tokenize.trivial_tokenize on a sentence that mixes Hindi and English words, similar to the examples in Hinglish and code-mixed text. Check whether the English words come out as clean separate tokens too, and whether punctuation attached to them gets split correctly.

What to learn next

  • Normalising Indic scripts — the normalisation logic used above, explained in full.
  • Tokenization — the general subword tokenization this script-aware word tokenizer is a different, earlier stage from.
  • Hugging Face — where AI4Bharat's pretrained models, like IndicBERT and IndicTrans2, are actually hosted and downloaded.

Researcher — Mathematics and papers.

The AI4Bharat stack, piece by piece

IndicNLP Library (Kunchukuttan, 2020). Rule-based preprocessing: script-aware tokenization, Unicode normalisation per the nukta-exclusion behaviour detailed in normalising Indic scripts, script conversion tables across the Brahmic script family, and basic morphological analysis for several languages. No neural network involved in the components used above.

IndicBERT (Kakwani et al., 2020). A BERT-family encoder pretrained specifically on 12 major Indian languages plus English, using the IndicCorp corpus. Structurally similar to XLM-RoBERTa but with vocabulary and training data concentrated on this language set rather than spread across 100 languages.

IndicTrans2 (Gala et al., 2023). The translation model already covered in translating a low-resource language, spanning all 22 scheduled Indian languages, evaluated with particular attention to the lower-resource languages in that set.

IndicXlit. A learned transliteration model, the model-based counterpart to the rule-based approach shown in transliteration between scripts, trained on real romanisation data rather than fixed formal-scheme rules.

The specialisation argument, precisely stated

A shared multilingual model's per-language quality on language l depends on that language's effective share of training signal and model capacity — the mechanism behind the "curse of multilinguality" discussed for XLM-RoBERTa. Restricting a model's language scope to a related family, as IndicBERT and IndicTrans2 do, increases each included language's effective share at fixed model size, at the cost of losing any transfer benefit from typologically distant languages outside that family — English apart, which both projects retain given its practical necessity for Indian-language NLP workflows.

Whether specialisation or broad coverage wins is empirically language- and task-dependent, not settled by architecture alone. Reported comparisons (Kakwani et al., 2020) generally favour IndicBERT over multilingual baselines like mBERT and XLM-RoBERTa-base on Indian-language benchmarks, though the gap narrows or reverses against larger general multilingual checkpoints, since scale itself is a confound alongside specialisation.

Key references

  • Kunchukuttan, A. (2020). The IndicNLP Library. github.com/anoopkunchukuttan/indic_nlp_library
  • Kakwani, D. et al. (2020). IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pretrained Multilingual Language Models for Indian Languages. EMNLP Findings. — IndicBERT and IndicCorp.
  • Gala, J. et al. (2023). IndicTrans2. arXiv:2305.16307
  • Madhani, Y. et al. (2023). Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users. arXiv:2205.03018 — IndicXlit and its training data.
  • Kunchukuttan, A. & Bhattacharyya, P. (2020). Utilizing Language Relatedness to Improve Machine Translation. arXiv:2003.08925

Current state and open problems

This entire stack depends on IndicCorp and related web-crawled corpora for its training signal, which inherits the same uneven per-language volume the rest of this section discusses — a well-resourced scheduled language like Hindi or Bengali has dramatically more IndicCorp text behind it than languages like Bodo or Santali, even within a project specifically built to help under-resourced Indian languages.

Active work in this space is shifting from encoder-only models (IndicBERT) toward generative and instruction-tuned Indic LLMs, following the field-wide move away from encoder-only architectures — but training data volume, not architecture choice, remains the binding constraint on quality for the least-resourced languages in this set.

What to learn next