AI glossary

Reasoning model

In one sentence A reasoning model is an LLM trained to think through a long internal chain of steps before answering, trading time and tokens for accuracy.

By Updated

A reasoning model is a language model trained to produce an extended chain of internal thinking before its final answer, spending extra compute at answer time to get harder problems right.

The chess-world distinction applies: blitz versus classical. The same grandmaster plays stronger with an hour on the clock than with ten seconds — not from knowing more, but from having time to explore lines, catch blunders, and reconsider. Standard LLMs answer in blitz mode. Reasoning models are granted classical time: they generate thousands of hidden "thinking" tokens — trying approaches, checking arithmetic, backtracking — then produce the visible answer.

The shift is in training, not prompting. Where chain-of-thought asks a standard model to show steps, reasoning models are trained — largely with reinforcement-learning on problems with checkable answers, like maths and code — to discover that long, self-correcting deliberation earns reward. OpenAI's o1 (2024) opened the category; DeepSeek-R1 (2025) published an open recipe; hybrid models now offer thinking as a switchable mode.

What this buys and costs, practically:

gains : maths, competitive coding, multi-step planning, tricky debugging
costs : seconds-to-minutes latency, thinking tokens billed, higher price per query
skip  : lookup, summaries, chat — deliberation adds cost, not quality

This is also "test-time scaling" — a new axis of improvement: spend more compute per answer, not only more per training run. One caution transfers from humans: visible deliberation is persuasive, and a model can reason its way, articulately, to a wrong answer.

Where to go next