How it works

How Google Translate works

Modern translation reads the whole source sentence into numbers with an encoder, then a decoder writes the target language one token at a time, guided by attention.

On this page 7
  1. The pipeline at a glance
  2. Stage 1 — chop the sentence into pieces
  3. Stage 2 — the encoder reads the whole sentence
  4. Stage 3 — the decoder writes the translation
  5. Stage 4 — keep several drafts, pick the best
  6. One model, many languages
  7. Which lessons teach each stage

Type a Hindi sentence, get English back in under a second — for any of 100+ languages, in either direction. No dictionary lookup could do this, because translation is not word swapping. "मुझे भूख लगी है" is not "to-me hunger attached is". The system has to read the whole thought, then say it again in another language. Here is how.

The pipeline at a glance

 "मुझे भूख लगी है"
        |
        v
 [1. tokenize]   ->  chop into subword pieces
        |
        v
 [2. encode]     ->  read the WHOLE sentence into numbers
        |
        v
 [3. decode]     ->  write English one token at a time,
        |            glancing back at the source via attention
        v
 [4. search]     ->  keep several draft translations, pick the best
        |
        v
 "I am hungry"

Stage 1 — chop the sentence into pieces

The text is cut into subword tokens — chunks smaller than words, learned from data. "Unhappiness" might become "un", "happi", "ness". This one trick lets a fixed vocabulary of maybe 30,000 pieces cover every word in every supported language, including names and words the system never saw. Rare word? It falls apart into familiar pieces.

Stage 2 — the encoder reads the whole sentence

An encoder — a stack of transformer layers — converts the token sequence into numbers that represent the sentence's meaning. Crucially, it reads everything before anything gets translated.

That matters because languages disagree about order. Hindi and Japanese put the verb last; English wants it early. German splits verbs in half. A system translating word-by-word as it reads would commit to English structure before the Hindi verb even arrives. Reading everything first is what lets the output be reordered freely.

Inside the encoder, each token looks at every other token through attention — the mechanism that lets "लगी" connect to "भूख" no matter the distance between them.

Stage 3 — the decoder writes the translation

A decoder now generates the English, one token per step. At every step it consults two things: what it has written so far, and — through attention again — the encoded source sentence. Writing "hungry", its attention leans on "भूख". This soft, learned lookup replaced the old idea of aligning word 3 to word 5.

Nobody hand-wrote grammar rules for any of this. The model learned from hundreds of millions of translated sentence pairs — books, websites, official documents published in multiple languages. Grammar emerged from patterns in the data.

Stage 4 — keep several drafts, pick the best

Greedily taking the top token at each step can paint the decoder into a corner: a word that looks best now may lead to an awkward sentence later. So translation systems run a beam search — keeping a handful of candidate translations alive in parallel, extending each, and finally picking the highest-scoring complete sentence.

One model, many languages

Google's system trains one shared model on many language pairs at once, with a tag telling the decoder which language to produce. A striking side effect: the model can translate between pairs it never saw directly, by leaning on the shared meaning space in the middle. Hindi to Japanese works without any Hindi–Japanese training data. This design also lifts quality for low-resource languages, which borrow strength from data-rich cousins.

The honest caveat: fluent is not the same as faithful. The decoder is rewarded for natural-sounding output, so its errors read smoothly. It can drop a negation or pick the wrong sense of a word, and nothing in the output warns you. Never rely on raw machine translation where a mistranslation has legal or medical consequences.

Which lessons teach each stage