AI glossary

Inference

In one sentence Inference is the moment you use a trained model to get an answer, as opposed to training, where the model is still learning.

By Updated

Inference is using a finished model to produce an answer. Training is study time; inference is exam day.

The student who spent six months studying does not learn anything new during the two hours of the exam. They read each question and produce an answer using what they already know. A model at inference time works the same way: the weights are frozen, nothing is updated, and the only work is turning your input into an output.

That difference shows up in the code. Training needs to remember every intermediate value so gradients can flow backwards; inference does not, which makes it much lighter on memory.

python
model.eval()                    # switches dropout off and batch-norm to running stats
with torch.no_grad():           # stops PyTorch storing values for backpropagation
    prediction = model(x)

Both lines are needed and they do different jobs. model.eval() changes layer behaviour; torch.no_grad() changes memory bookkeeping. Leaving out the second is a frequent cause of CUDA out of memory during evaluation.

Why the industry cares so much about it

Training happens once, or a few times a year. Inference happens on every single request, forever. A model trained for two weeks might then serve requests for three years, so almost all of the lifetime compute bill, and all of the latency your users feel, lives here. That is the entire reason for quantization, distillation, batching, KV caching and specialised servers like vLLM.

For language models, inference has two distinct phases with different costs. Reading your prompt processes all tokens in parallel, so it is fast. Generating the reply produces one token at a time, and cannot be parallelised in the same way. That is why a long prompt is cheap compared to a long answer.

Where to go next