LLM (large language model)
In one sentence A large language model is a program trained on enormous amounts of text to predict what comes next, which turns out to be enough to answer, write and translate.
Updated
A large language model is a program trained on an enormous amount of text to predict the next piece of text — and that one ability turns out to cover answering questions, writing code and translating.
The autocomplete on your phone guesses your next word from the last two or three. Now imagine that same idea, trained on a very large slice of the public internet, with hundreds of billions of adjustable numbers instead of a small table. To predict the next word of a physics answer well, it has to have absorbed something about physics. Prediction, done well enough, starts to look like understanding.
At the core the loop is small: text comes in as tokens, the model outputs a probability for every possible next token, one is picked, it is appended to the input, and the whole thing runs again.
"The capital of France is" → Paris 0.94 | the 0.02 | a 0.01 | located 0.01 ...
→ pick "Paris", append, predict againWhat follows from that
Three consequences catch people out, and all three come from this one design.
The model has no database. It has weights. It cannot look anything up, which is why it invents plausible details when the training text was thin — see hallucination. It also has no memory between requests: everything it appears to remember about your conversation was re-sent inside the context window. And it does not see letters, it sees tokens, which is why counting the r's in a word can go wrong while writing a working SQL query does not.
The architecture underneath is the transformer, and "large" means something concrete: model size counted in billions of parameters, training measured in trillions of tokens. Open-weight families like Llama, Mistral, Qwen and Gemma can run on your own hardware; hosted models from OpenAI, Anthropic and Google run behind an API.
Where to go next
- Full lesson: What is a large language model?
- Related terms: token, transformer, context-window, hallucination