How AI agents work
An agent is an LLM in a loop — it reads a goal, picks a tool, acts, reads the result, and repeats until the job is done or a guardrail stops it.
Translation is unavailable right now. Your browser’s own “Translate page” usually works.
Short pieces that open up one idea at a time. Read one when a lesson mentions something you want to see from the inside, rather than in passing.
15 explainers
An agent is an LLM in a loop — it reads a goal, picks a tool, acts, reads the result, and repeats until the job is done or a guardrail stops it.
Image models start from pure noise and remove it step by step, steered by your text prompt, until a picture that matches the words emerges.
ChatGPT chops your message into tokens, scores every possible next token, and streams the winners back — shaped by training that taught it which answers people prefer.
Serving an LLM means one parallel pass over your prompt, then a one-token-at-a-time loop — kept fast by a cache, kept affordable by batching many users onto one GPU.
Modern translation reads the whole source sentence into numbers with an encoder, then a decoder writes the target language one token at a time, guided by attention.
OCR cleans up the photo, finds where the text lines are, reads each line as a sequence rather than letter by letter, and uses language knowledge to fix what the pixels left unclear.
RAG searches your documents first, then hands the best passages to the model so it answers from what it found instead of from memory.
Face recognition finds a face, straightens it, compresses it into a list of numbers, and decides identity by measuring how close that list is to a stored one.
Fine-tuning takes a model that already knows language, shows it your examples, and nudges its weights gently until it behaves your way without forgetting what it knew.
Training is one loop run millions of times — guess, measure the error, trace blame backwards through the network, nudge every weight, repeat until the guesses get good.
Recommenders log what you do, narrow millions of items to a few hundred candidates, score each one for you personally, and learn from what you click next.
A search engine reads the web ahead of time into an index, understands your query, pulls thousands of candidate pages, and ranks them with hundreds of signals — all before you finish blinking.
A spam filter turns each email into measurable features, scores it with a classifier trained on millions of labelled messages, and keeps learning as spammers adapt.
A microphone turns your voice into thousands of numbers per second, a spectrogram turns those into a picture of sound, and a model reads that picture into words.
A tiny always-on model listens for the wake word, speech recognition turns your request into text, an understanding layer picks the action, and synthesis speaks the reply.