The AI4Bharat and IndicNLP stack
AI4Bharat and the IndicNLP library are open tools built specifically for Indian languages, instead of stretching general multilingual tools to cover them.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
AI4Bharat and the IndicNLP library are free tools built specifically for Indian languages. They do not treat those languages as an afterthought.
Think about a tailor who makes clothes for one specific body type. He measures real customers of that build, instead of stitching one universal size for everyone. The tailor's fit is better, because the sizing was never averaged away.
General multilingual tools split their attention across 100 languages. AI4Bharat, a research group based in India, built tools that specialise instead. Its tokenizers, normalisers, transliteration and translation models are trained and tuned specifically for India's languages and scripts.
Why it exists
A model trained on 100 languages at once, like XLM-RoBERTa, spreads its capacity across every one of them. Indian languages get a share of that capacity, not the majority of it. India alone has 22 constitutionally scheduled languages, and hundreds more spoken daily.
AI4Bharat's answer: build the tokenizer, the normaliser, and the models around Indian scripts and languages. Make them the main design target, not one entry in a much longer list. IndicBERT, IndicTrans2, IndicXlit and the IndicNLP library are the resulting pieces, trained and tuned specifically for this language family.
How it works
General multilingual model: 100 languages sharing one budget of attention
AI4Bharat stack: Indian scripts and languages as the primary target,
not one language among a hundred othersThe IndicNLP library handles the everyday preprocessing work directly — tokenizing, normalising, transliterating — for a dozen or more Indian scripts. It needs no large pretrained model for these basic steps.
Where you have already seen it
- IndicTrans2, the translation model already introduced in translating a low-resource language, covering all 22 scheduled Indian languages.
- Government and civic-tech projects processing Indian-language text, frequently building on this open stack rather than starting from scratch.
- Academic NLP research on Indian languages, where this stack is close to a default starting point.
Remember this
- AI4Bharat builds NLP tools specifically for Indian languages. It does not treat them as one entry among many in a general model.
- The IndicNLP library handles everyday preprocessing — tokenizing, normalising, transliterating — without needing a large model download.
- Specialised tooling and general multilingual models are complements, not rivals — production systems commonly use both.
What to learn next
- Normalising Indic scripts — the specific normalisation work this toolkit performs.
- XLM-RoBERTa and multilingual encoders — the general-purpose alternative this toolkit specialises away from.
- Translating a low-resource language — where IndicTrans2, part of this same stack, was already introduced.
Developer — Code and libraries.
Setup
pip install indic-nlp-libraryThis installs a pure-Python package for basic tokenizing and normalising — no large model download for the pieces used below.
Tokenizing and normalising Hindi with IndicNLP
from indicnlp.tokenize import indic_tokenize
from indicnlp.normalize.indic_normalize import IndicNormalizerFactory
text = "मुझे कल सुबह 9:30 बजे मीटिंग है, कृपया याद दिलाना।"
tokens = indic_tokenize.trivial_tokenize(text, lang="hi")
print("tokens:", tokens)
normalizer = IndicNormalizerFactory().get_normalizer("hi")
normalised = normalizer.normalize(text)
print("normalisation changed the text:", normalised != text)tokens: ['मुझे', 'कल', 'सुबह', '9:30', 'बजे', 'मीटिंग', 'है', ',', 'कृपया', 'याद', 'दिलाना', '।'] normalisation changed the text: False
Line by line
trivial_tokenize splits on script-aware rules, not only whitespace. It correctly keeps "9:30" as one token, and separates the trailing danda (।, the Devanagari sentence-end mark) from "दिलाना" as its own token, the way a general whitespace-splitter would get wrong.
The normaliser reported no change here, because this sentence was already in one consistent Unicode form. Normalising Indic scripts shows the same normaliser actually changing text, on an input deliberately constructed with an inconsistent nukta encoding.
Neither function downloaded a model. Both are rule-based, script-aware code — fast, deterministic, and useful even on a machine with no GPU and no internet connection after installation.
Common mistakes
Reaching for a full pretrained model when a rule-based tool already does the job. Tokenizing and normalising Indic script text does not need IndicBERT or any neural model — the code above handles it directly, and faster.
Using trivial_tokenize for a language it was not told about. The lang argument changes real behaviour — script-specific punctuation and abbreviation handling differ by language, even across languages sharing the Devanagari script.
Assuming this library covers every Indian language equally well. Coverage, tooling maturity and testing depth vary by language inside this stack too, generally tracking the same resource patterns discussed in translating a low-resource language — a widely-spoken, well-resourced language gets more mature tooling than a less-documented one.
Try it yourself
Run indic_tokenize.trivial_tokenize on a sentence that mixes Hindi and English words, similar to the examples in Hinglish and code-mixed text. Check whether the English words come out as clean separate tokens too, and whether punctuation attached to them gets split correctly.
What to learn next
- Normalising Indic scripts — the normalisation logic used above, explained in full.
- Tokenization — the general subword tokenization this script-aware word tokenizer is a different, earlier stage from.
- Hugging Face — where AI4Bharat's pretrained models, like IndicBERT and IndicTrans2, are actually hosted and downloaded.
Researcher — Mathematics and papers.
The AI4Bharat stack, piece by piece
IndicNLP Library (Kunchukuttan, 2020). Rule-based preprocessing: script-aware tokenization, Unicode normalisation per the nukta-exclusion behaviour detailed in normalising Indic scripts, script conversion tables across the Brahmic script family, and basic morphological analysis for several languages. No neural network involved in the components used above.
IndicBERT (Kakwani et al., 2020). A BERT-family encoder pretrained specifically on 12 major Indian languages plus English, using the IndicCorp corpus. Structurally similar to XLM-RoBERTa but with vocabulary and training data concentrated on this language set rather than spread across 100 languages.
IndicTrans2 (Gala et al., 2023). The translation model already covered in translating a low-resource language, spanning all 22 scheduled Indian languages, evaluated with particular attention to the lower-resource languages in that set.
IndicXlit. A learned transliteration model, the model-based counterpart to the rule-based approach shown in transliteration between scripts, trained on real romanisation data rather than fixed formal-scheme rules.
The specialisation argument, precisely stated
A shared multilingual model's per-language quality on language l depends on that language's effective share of training signal and model capacity — the mechanism behind the "curse of multilinguality" discussed for XLM-RoBERTa. Restricting a model's language scope to a related family, as IndicBERT and IndicTrans2 do, increases each included language's effective share at fixed model size, at the cost of losing any transfer benefit from typologically distant languages outside that family — English apart, which both projects retain given its practical necessity for Indian-language NLP workflows.
Whether specialisation or broad coverage wins is empirically language- and task-dependent, not settled by architecture alone. Reported comparisons (Kakwani et al., 2020) generally favour IndicBERT over multilingual baselines like mBERT and XLM-RoBERTa-base on Indian-language benchmarks, though the gap narrows or reverses against larger general multilingual checkpoints, since scale itself is a confound alongside specialisation.
Key references
- Kunchukuttan, A. (2020). The IndicNLP Library. github.com/anoopkunchukuttan/indic_nlp_library
- Kakwani, D. et al. (2020). IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pretrained Multilingual Language Models for Indian Languages. EMNLP Findings. — IndicBERT and IndicCorp.
- Gala, J. et al. (2023). IndicTrans2. arXiv:2305.16307
- Madhani, Y. et al. (2023). Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users. arXiv:2205.03018 — IndicXlit and its training data.
- Kunchukuttan, A. & Bhattacharyya, P. (2020). Utilizing Language Relatedness to Improve Machine Translation. arXiv:2003.08925
Current state and open problems
This entire stack depends on IndicCorp and related web-crawled corpora for its training signal, which inherits the same uneven per-language volume the rest of this section discusses — a well-resourced scheduled language like Hindi or Bengali has dramatically more IndicCorp text behind it than languages like Bodo or Santali, even within a project specifically built to help under-resourced Indian languages.
Active work in this space is shifting from encoder-only models (IndicBERT) toward generative and instruction-tuned Indic LLMs, following the field-wide move away from encoder-only architectures — but training data volume, not architecture choice, remains the binding constraint on quality for the least-resourced languages in this set.
What to learn next
- Translating a low-resource language — the specific problem IndicTrans2 was built to address.
- XLM-RoBERTa and multilingual encoders — the general-purpose model family IndicBERT specialises away from.
- Normalising Indic scripts — the preprocessing layer underneath every model in this stack.