Multilingual and Indic NLP

How neural machine translation works

Neural machine translation reads a whole sentence into one model, then generates the translation word by word, using attention to look back at the right source words.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Neural machine translation reads an entire sentence first. Then it writes out the translation one word at a time, checking back on the original as it goes.

Picture a skilled interpreter at a wedding. She listens to the full Hindi sentence before speaking a single word of English. She does not translate word by word as she hears them. She waits, understands the whole idea, then speaks fluently in the other language. She glances back at specific parts of what was said, whenever she needs to.

Understand everything first. Then produce word by word, checking back as you go. That two-step rhythm is exactly how neural machine translation works.

Why it exists

Older translation systems worked phrase by phrase. They used large tables of phrase-to-phrase mappings. The tables came from statistics counted across millions of sentence pairs. They produced translations that were often locally correct and globally awkward — right phrases, wrong flow.

Neural machine translation replaced the phrase table with a single model. That model reads the whole source sentence, builds one understanding of it, and generates fluent output from there. Fluency improved sharply, because the model plans the whole sentence rather than stitching phrases together.

How it works

  Source:  "मुझे चाय पसंद है"
              |
              v
        [ ENCODER ]  -->  one understanding of the whole sentence
              |
              v
        [ DECODER ]  -->  "I"  -->  "like"  -->  "tea"  -->  (done)
              ^              ^          ^
              |              |          |
        looks back at the source sentence at every step (this is attention)

The encoder reads the source sentence and builds a representation of it. The decoder generates the translation one word at a time. It uses attention to look back at whichever source words matter most, right now.

Where you have already seen it

  • Google Translate, translating full sentences with much better word order than older phrase-based tools ever managed.
  • Live captions with translation, generating fluent target-language text from speech in real time.
  • In-app translate buttons on social media, on comments written in a language you do not read.

Remember this

  • Neural machine translation has two parts. An encoder reads the source. A decoder writes the target, one word at a time.
  • Attention lets the decoder look back at the right source words for each word it produces.
  • This replaced older phrase-table systems, mainly by producing far more fluent, natural-sounding output.

What to learn next

Developer — Code and libraries.

This trains a genuine, tiny encoder-decoder-with-attention translator from scratch — English number words to Hindi number words. Ten sentence pairs, sixteen numbers per word, done in under a second on CPU.

Setup

bash
pip install torch --index-url https://download.pytorch.org/whl/cpu

A complete encoder-decoder with attention, trained from scratch

tiny_nmt.py
import torch
import torch.nn as nn
import torch.nn.functional as F

torch.manual_seed(0)

pairs = [
    ("one", "ek"), ("two", "do"), ("three", "teen"), ("four", "chaar"),
    ("five", "paanch"), ("six", "chhah"), ("seven", "saat"),
    ("eight", "aath"), ("nine", "nau"), ("ten", "das"),
]

SRC = sorted({w for w, _ in pairs})
TGT = sorted({w for _, w in pairs})
src_vocab = {w: i + 2 for i, w in enumerate(SRC)}
tgt_vocab = {w: i + 3 for i, w in enumerate(TGT)}   # 0=pad, 1=bos, 2=eos
PAD, BOS, EOS = 0, 1, 2
tgt_vocab_size = len(tgt_vocab) + 3
src_vocab_size = len(src_vocab) + 2
inv_tgt = {i: w for w, i in tgt_vocab.items()}
d_model = 16

class TinyTranslator(nn.Module):
    def __init__(self):
        super().__init__()
        self.src_emb = nn.Embedding(src_vocab_size, d_model)
        self.tgt_emb = nn.Embedding(tgt_vocab_size, d_model)
        self.encoder = nn.GRU(d_model, d_model, batch_first=True)
        self.decoder = nn.GRUCell(d_model, d_model)
        self.out = nn.Linear(d_model, tgt_vocab_size)

    def forward(self, src_ids, tgt_in_ids):
        enc_out, h = self.encoder(self.src_emb(src_ids))
        h = h.squeeze(0)
        logits = []
        for t in range(tgt_in_ids.shape[1]):
            h = self.decoder(self.tgt_emb(tgt_in_ids[:, t]), h)
            scores = (enc_out @ h.unsqueeze(-1)).squeeze(-1)   # attention scores
            weights = F.softmax(scores, dim=-1)
            context = (weights.unsqueeze(-1) * enc_out).sum(1)
            logits.append(self.out(h + context))
        return torch.stack(logits, dim=1)

model = TinyTranslator()
optimiser = torch.optim.Adam(model.parameters(), lr=0.01)

def encode(word, vocab):
    return torch.tensor([[vocab[word]]])

for epoch in range(300):
    total_loss = 0.0
    for src_word, tgt_word in pairs:
        src_ids = encode(src_word, src_vocab)
        tgt_ids = torch.tensor([[BOS, tgt_vocab[tgt_word], EOS]])
        logits = model(src_ids, tgt_ids[:, :-1])
        loss = F.cross_entropy(logits.view(-1, tgt_vocab_size), tgt_ids[:, 1:].reshape(-1))
        optimiser.zero_grad()
        loss.backward()
        optimiser.step()
        total_loss += loss.item()
    if epoch % 50 == 0 or epoch == 299:
        print(f"epoch {epoch:3d}  loss {total_loss / len(pairs):.4f}")

@torch.no_grad()
def translate(word):
    src_ids = encode(word, src_vocab)
    enc_out, h = model.encoder(model.src_emb(src_ids))
    h = h.squeeze(0)
    token = torch.tensor([BOS])
    result = []
    for _ in range(5):
        h = model.decoder(model.tgt_emb(token), h)
        scores = (enc_out @ h.unsqueeze(-1)).squeeze(-1)
        weights = F.softmax(scores, dim=-1)
        context = (weights.unsqueeze(-1) * enc_out).sum(1)
        token = model.out(h + context).argmax(-1)
        if token.item() == EOS:
            break
        result.append(inv_tgt.get(token.item(), "?"))
    return " ".join(result)

print()
for src_word, expected in pairs:
    print(f"{src_word:6} -> {translate(src_word):8}  (expected: {expected})")
Output
epoch   0  loss 2.4800
epoch  50  loss 0.0025
epoch 100  loss 0.0008
epoch 150  loss 0.0004
epoch 200  loss 0.0002
epoch 250  loss 0.0002
epoch 299  loss 0.0001

one    -> ek        (expected: ek)
two    -> do        (expected: do)
three  -> teen      (expected: teen)
four   -> chaar     (expected: chaar)
five   -> paanch    (expected: paanch)
six    -> chhah     (expected: chhah)
seven  -> saat      (expected: saat)
eight  -> aath      (expected: aath)
nine   -> nau       (expected: nau)
ten    -> das       (expected: das)

Loss values above come from this exact seeded run and will repeat identically on any machine, since every random source is fixed by torch.manual_seed(0).

Line by line

The encoder is one GRU, reading the source word into a single hidden state h. For a one-word source, this is almost trivial — the real work in a full system happens when enc_out holds one vector per source word instead of one.

The decoder loop is the interesting part. At every step, scores = enc_out @ h compares the decoder's current state against every encoder output — precisely the query-key dot product from attention, without separate learned projection matrices for query and key.

EOS is the model's way of saying "I am done." Real translation output has no fixed length. Generation stops the moment the model itself predicts the end-of-sequence token, not after a hardcoded number of words.

Training reached near-zero loss because the task is extremely small. Ten pairs, memorised perfectly, is expected here — it demonstrates the mechanism working, not translation quality at any real scale.

Common mistakes

Forgetting the EOS token entirely. Without it, a decoder has no learned signal for when to stop, and either runs forever or needs an arbitrary length cutoff — exactly the bug fixed by adding EOS above.

Using the same embedding table for source and target. They are different vocabularies here, in different languages. A shared vocabulary only makes sense when source and target share subword pieces, as in the multilingual models covered in XLM-RoBERTa.

Reading 0.0001 training loss as "the model translates well." It means the model memorised ten examples perfectly. A real system is judged on sentences it never saw during training — this toy has no held-out test set at all.

Try it yourself

Add an eleventh pair the model has never seen, say ("eleven", "gyarah"), to the pairs list and retrain. Then try calling translate("eleven") before retraining — it fails with a KeyError, because the word literally is not in the vocabulary yet. That failure is the subject of translating a low-resource language.

What to learn next

Researcher — Mathematics and papers.

The sequence-to-sequence objective

text
P(y_1, ..., y_m | x_1, ..., x_n) = product over t = 1..m of P(y_t | y_{<t}, x_1, ..., x_n)
  • x_1, ..., x_n are source tokens, y_1, ..., y_m target tokens.
  • Training maximises this likelihood over a parallel corpus via teacher forcing: the true previous target token y_{t-1} is fed to the decoder at training time, rather than the model's own prediction, so errors do not compound during training.
  • At inference, y_{<t} is the model's own previously generated output — greedy argmax in the demo above, beam search in production systems.

From RNN encoder-decoder to attention to transformers

Sutskever, Vinyals & Le (2014) introduced the encoder-decoder pattern for translation, compressing the entire source sentence into one fixed-size vector — the same bottleneck problem motivating attention's development, described in Bahdanau, Cho & Bengio (2014): instead of one fixed vector, the decoder attends over every encoder position at every output step, exactly as the toy decoder above does with enc_out @ h.

Wu et al. (2016), Google's Neural Machine Translation system, scaled this to production: 8-layer LSTM encoder-decoder with attention, replacing Google Translate's phrase-based system and reporting substantial fluency gains over the phrase-based baseline it replaced, measured by both automatic metrics and human evaluation.

Vaswani et al. (2017) replaced the recurrent encoder and decoder with self-attention entirely — the Transformer — trading the RNN's sequential per-token computation for parallel computation across the whole sequence, at the O(n^2) cost detailed in attention's researcher block. This is the architecture behind essentially every production translation system built since.

Training data and evaluation

Parallel corpora — sentence pairs with verified matching meaning across languages — are the fuel for this entire approach. OPUS (Tiedemann, 2012) aggregates the largest open collection, pooling parliamentary proceedings, subtitles, and web-crawled parallel text across hundreds of language pairs, with wildly unequal size per pair.

BLEU (Papineni et al., 2002) remains the standard automatic metric, scoring n-gram overlap between generated and reference translations, though it is now understood to correlate imperfectly with human judgments of fluency and correlates especially poorly for morphologically rich languages, where a correct translation can use different but equally valid word forms than the reference. COMET (Rei et al., 2020) and chrF (Popović, 2015) are increasingly used alongside it for exactly this reason.

Key references

  • Sutskever, I., Vinyals, O. & Le, Q. (2014). Sequence to Sequence Learning with Neural Networks. arXiv:1409.3215
  • Bahdanau, D., Cho, K. & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv:1409.0473
  • Wu, Y. et al. (2016). Google's Neural Machine Translation System. arXiv:1609.08144
  • Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
  • Papineni, K. et al. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. ACL.

Current state and open problems

Production translation has moved from one model per language pair (MarianMT-style) toward single models covering hundreds of languages and pairs at once (NLLB-200, IndicTrans2) — see running translation models yourself for the practical shape of that shift.

Long documents remain harder than single sentences: translating sentence-by-sentence loses cross-sentence coherence (consistent pronoun and terminology choice across a paragraph), and document-level translation is an active research area without a settled standard architecture.

Evaluation is arguably the least settled part of the field. Automatic metrics disagree with each other and with human judgment often enough that serious translation system comparisons still require human evaluation, which is expensive and slow — a genuine bottleneck on how fast the field can iterate.

What to learn next