How neural machine translation works
Neural machine translation reads a whole sentence into one model, then generates the translation word by word, using attention to look back at the right source words.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Neural machine translation reads an entire sentence first. Then it writes out the translation one word at a time, checking back on the original as it goes.
Picture a skilled interpreter at a wedding. She listens to the full Hindi sentence before speaking a single word of English. She does not translate word by word as she hears them. She waits, understands the whole idea, then speaks fluently in the other language. She glances back at specific parts of what was said, whenever she needs to.
Understand everything first. Then produce word by word, checking back as you go. That two-step rhythm is exactly how neural machine translation works.
Why it exists
Older translation systems worked phrase by phrase. They used large tables of phrase-to-phrase mappings. The tables came from statistics counted across millions of sentence pairs. They produced translations that were often locally correct and globally awkward — right phrases, wrong flow.
Neural machine translation replaced the phrase table with a single model. That model reads the whole source sentence, builds one understanding of it, and generates fluent output from there. Fluency improved sharply, because the model plans the whole sentence rather than stitching phrases together.
How it works
Source: "मुझे चाय पसंद है"
|
v
[ ENCODER ] --> one understanding of the whole sentence
|
v
[ DECODER ] --> "I" --> "like" --> "tea" --> (done)
^ ^ ^
| | |
looks back at the source sentence at every step (this is attention)The encoder reads the source sentence and builds a representation of it. The decoder generates the translation one word at a time. It uses attention to look back at whichever source words matter most, right now.
Where you have already seen it
- Google Translate, translating full sentences with much better word order than older phrase-based tools ever managed.
- Live captions with translation, generating fluent target-language text from speech in real time.
- In-app translate buttons on social media, on comments written in a language you do not read.
Remember this
- Neural machine translation has two parts. An encoder reads the source. A decoder writes the target, one word at a time.
- Attention lets the decoder look back at the right source words for each word it produces.
- This replaced older phrase-table systems, mainly by producing far more fluent, natural-sounding output.
What to learn next
- Running translation models yourself — using a real, pretrained version of this architecture.
- Attention — the exact mechanism the decoder uses to look back at the source.
- Translating a low-resource language — what happens when there is not enough data to train this well.
Developer — Code and libraries.
This trains a genuine, tiny encoder-decoder-with-attention translator from scratch — English number words to Hindi number words. Ten sentence pairs, sixteen numbers per word, done in under a second on CPU.
Setup
pip install torch --index-url https://download.pytorch.org/whl/cpuA complete encoder-decoder with attention, trained from scratch
import torch
import torch.nn as nn
import torch.nn.functional as F
torch.manual_seed(0)
pairs = [
("one", "ek"), ("two", "do"), ("three", "teen"), ("four", "chaar"),
("five", "paanch"), ("six", "chhah"), ("seven", "saat"),
("eight", "aath"), ("nine", "nau"), ("ten", "das"),
]
SRC = sorted({w for w, _ in pairs})
TGT = sorted({w for _, w in pairs})
src_vocab = {w: i + 2 for i, w in enumerate(SRC)}
tgt_vocab = {w: i + 3 for i, w in enumerate(TGT)} # 0=pad, 1=bos, 2=eos
PAD, BOS, EOS = 0, 1, 2
tgt_vocab_size = len(tgt_vocab) + 3
src_vocab_size = len(src_vocab) + 2
inv_tgt = {i: w for w, i in tgt_vocab.items()}
d_model = 16
class TinyTranslator(nn.Module):
def __init__(self):
super().__init__()
self.src_emb = nn.Embedding(src_vocab_size, d_model)
self.tgt_emb = nn.Embedding(tgt_vocab_size, d_model)
self.encoder = nn.GRU(d_model, d_model, batch_first=True)
self.decoder = nn.GRUCell(d_model, d_model)
self.out = nn.Linear(d_model, tgt_vocab_size)
def forward(self, src_ids, tgt_in_ids):
enc_out, h = self.encoder(self.src_emb(src_ids))
h = h.squeeze(0)
logits = []
for t in range(tgt_in_ids.shape[1]):
h = self.decoder(self.tgt_emb(tgt_in_ids[:, t]), h)
scores = (enc_out @ h.unsqueeze(-1)).squeeze(-1) # attention scores
weights = F.softmax(scores, dim=-1)
context = (weights.unsqueeze(-1) * enc_out).sum(1)
logits.append(self.out(h + context))
return torch.stack(logits, dim=1)
model = TinyTranslator()
optimiser = torch.optim.Adam(model.parameters(), lr=0.01)
def encode(word, vocab):
return torch.tensor([[vocab[word]]])
for epoch in range(300):
total_loss = 0.0
for src_word, tgt_word in pairs:
src_ids = encode(src_word, src_vocab)
tgt_ids = torch.tensor([[BOS, tgt_vocab[tgt_word], EOS]])
logits = model(src_ids, tgt_ids[:, :-1])
loss = F.cross_entropy(logits.view(-1, tgt_vocab_size), tgt_ids[:, 1:].reshape(-1))
optimiser.zero_grad()
loss.backward()
optimiser.step()
total_loss += loss.item()
if epoch % 50 == 0 or epoch == 299:
print(f"epoch {epoch:3d} loss {total_loss / len(pairs):.4f}")
@torch.no_grad()
def translate(word):
src_ids = encode(word, src_vocab)
enc_out, h = model.encoder(model.src_emb(src_ids))
h = h.squeeze(0)
token = torch.tensor([BOS])
result = []
for _ in range(5):
h = model.decoder(model.tgt_emb(token), h)
scores = (enc_out @ h.unsqueeze(-1)).squeeze(-1)
weights = F.softmax(scores, dim=-1)
context = (weights.unsqueeze(-1) * enc_out).sum(1)
token = model.out(h + context).argmax(-1)
if token.item() == EOS:
break
result.append(inv_tgt.get(token.item(), "?"))
return " ".join(result)
print()
for src_word, expected in pairs:
print(f"{src_word:6} -> {translate(src_word):8} (expected: {expected})")epoch 0 loss 2.4800 epoch 50 loss 0.0025 epoch 100 loss 0.0008 epoch 150 loss 0.0004 epoch 200 loss 0.0002 epoch 250 loss 0.0002 epoch 299 loss 0.0001 one -> ek (expected: ek) two -> do (expected: do) three -> teen (expected: teen) four -> chaar (expected: chaar) five -> paanch (expected: paanch) six -> chhah (expected: chhah) seven -> saat (expected: saat) eight -> aath (expected: aath) nine -> nau (expected: nau) ten -> das (expected: das)
Loss values above come from this exact seeded run and will repeat identically on any machine, since every random source is fixed by torch.manual_seed(0).
Line by line
The encoder is one GRU, reading the source word into a single hidden state h. For a one-word source, this is almost trivial — the real work in a full system happens when enc_out holds one vector per source word instead of one.
The decoder loop is the interesting part. At every step, scores = enc_out @ h compares the decoder's current state against every encoder output — precisely the query-key dot product from attention, without separate learned projection matrices for query and key.
EOS is the model's way of saying "I am done." Real translation output has no fixed length. Generation stops the moment the model itself predicts the end-of-sequence token, not after a hardcoded number of words.
Training reached near-zero loss because the task is extremely small. Ten pairs, memorised perfectly, is expected here — it demonstrates the mechanism working, not translation quality at any real scale.
Common mistakes
Forgetting the EOS token entirely. Without it, a decoder has no learned signal for when to stop, and either runs forever or needs an arbitrary length cutoff — exactly the bug fixed by adding EOS above.
Using the same embedding table for source and target. They are different vocabularies here, in different languages. A shared vocabulary only makes sense when source and target share subword pieces, as in the multilingual models covered in XLM-RoBERTa.
Reading 0.0001 training loss as "the model translates well." It means the model memorised ten examples perfectly. A real system is judged on sentences it never saw during training — this toy has no held-out test set at all.
Try it yourself
Add an eleventh pair the model has never seen, say ("eleven", "gyarah"), to the pairs list and retrain. Then try calling translate("eleven") before retraining — it fails with a KeyError, because the word literally is not in the vocabulary yet. That failure is the subject of translating a low-resource language.
What to learn next
- Running translation models yourself — the real, pretrained version of this exact architecture shape.
- Attention — the full mechanism this toy decoder implements a simplified version of.
- Translating a low-resource language — what breaks when the training data is this small for a real task.
Researcher — Mathematics and papers.
The sequence-to-sequence objective
P(y_1, ..., y_m | x_1, ..., x_n) = product over t = 1..m of P(y_t | y_{<t}, x_1, ..., x_n)x_1, ..., x_nare source tokens,y_1, ..., y_mtarget tokens.- Training maximises this likelihood over a parallel corpus via teacher forcing: the true previous target token
y_{t-1}is fed to the decoder at training time, rather than the model's own prediction, so errors do not compound during training. - At inference,
y_{<t}is the model's own previously generated output — greedy argmax in the demo above, beam search in production systems.
From RNN encoder-decoder to attention to transformers
Sutskever, Vinyals & Le (2014) introduced the encoder-decoder pattern for translation, compressing the entire source sentence into one fixed-size vector — the same bottleneck problem motivating attention's development, described in Bahdanau, Cho & Bengio (2014): instead of one fixed vector, the decoder attends over every encoder position at every output step, exactly as the toy decoder above does with enc_out @ h.
Wu et al. (2016), Google's Neural Machine Translation system, scaled this to production: 8-layer LSTM encoder-decoder with attention, replacing Google Translate's phrase-based system and reporting substantial fluency gains over the phrase-based baseline it replaced, measured by both automatic metrics and human evaluation.
Vaswani et al. (2017) replaced the recurrent encoder and decoder with self-attention entirely — the Transformer — trading the RNN's sequential per-token computation for parallel computation across the whole sequence, at the O(n^2) cost detailed in attention's researcher block. This is the architecture behind essentially every production translation system built since.
Training data and evaluation
Parallel corpora — sentence pairs with verified matching meaning across languages — are the fuel for this entire approach. OPUS (Tiedemann, 2012) aggregates the largest open collection, pooling parliamentary proceedings, subtitles, and web-crawled parallel text across hundreds of language pairs, with wildly unequal size per pair.
BLEU (Papineni et al., 2002) remains the standard automatic metric, scoring n-gram overlap between generated and reference translations, though it is now understood to correlate imperfectly with human judgments of fluency and correlates especially poorly for morphologically rich languages, where a correct translation can use different but equally valid word forms than the reference. COMET (Rei et al., 2020) and chrF (Popović, 2015) are increasingly used alongside it for exactly this reason.
Key references
- Sutskever, I., Vinyals, O. & Le, Q. (2014). Sequence to Sequence Learning with Neural Networks. arXiv:1409.3215
- Bahdanau, D., Cho, K. & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv:1409.0473
- Wu, Y. et al. (2016). Google's Neural Machine Translation System. arXiv:1609.08144
- Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
- Papineni, K. et al. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. ACL.
Current state and open problems
Production translation has moved from one model per language pair (MarianMT-style) toward single models covering hundreds of languages and pairs at once (NLLB-200, IndicTrans2) — see running translation models yourself for the practical shape of that shift.
Long documents remain harder than single sentences: translating sentence-by-sentence loses cross-sentence coherence (consistent pronoun and terminology choice across a paragraph), and document-level translation is an active research area without a settled standard architecture.
Evaluation is arguably the least settled part of the field. Automatic metrics disagree with each other and with human judgment often enough that serious translation system comparisons still require human evaluation, which is expensive and slow — a genuine bottleneck on how fast the field can iterate.
What to learn next
- Attention — the mechanism this lesson's decoder implements in miniature.
- Running translation models yourself — real pretrained checkpoints implementing this exact pattern at scale.
- Translating a low-resource language — what this pipeline needs when parallel data is scarce.