Post-training and Alignment

Thinking tokens and reasoning models

A reasoning model writes rough work before its answer, and that rough work is not decoration — it is extra computation the model cannot otherwise perform.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Where reasoning models came from
  6. What is honestly hard here
  7. Where you have already seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A reasoning model writes rough work before its answer, and that rough work does real computation.

The analogy you have already lived

Multiply 47 by 83 in your head, right now, with no paper.

Most people cannot. Give the same person a scrap of paper and they get it in twenty seconds. The paper does not make them cleverer. It gives them somewhere to put a partial result while they work on the next one.

A model has the same limitation and the same fix. Its scrap of paper is the text it writes before the answer.

Why it exists

A model produces one token at a time. For each token it does a fixed amount of work. The same amount for "what is 2 plus 2" as for "prove this theorem".

That is a hard ceiling. Some problems need more steps than fit in one token's worth of computation.

But the model can read what it has already written. So if it writes down a partial result, it can use that partial result when producing the next token. Each written line becomes a step it did not have to hold in its head.

The number of steps is no longer fixed. It is as long as the model wants to write.

How it works

   question:  "A shop sells 3 pens for 45 rupees. What do 7 pens cost?"

   without rough work
   "105"                                <- one shot, often wrong

   with rough work
   "3 pens cost 45, so 1 pen costs 15.
    7 pens cost 7 times 15, which is 105.
    <answer> 105 </answer>"             <- each line reuses the last

Modern models put this rough work inside markers, often called <think> tags. The interface hides that part and shows you the answer. The rough work is still there, and you still pay for every token of it.

Where reasoning models came from

At first, people got this behaviour by asking. Adding "think step by step" to a prompt improved accuracy on maths problems noticeably.

Then teams trained it in. They rewarded the model for producing correct final answers, and let it decide how much rough work to do.

The models chose to write more and more. They started re-checking their own steps and correcting themselves mid-answer. Nobody wrote a rule asking for that.

What is honestly hard here

The rough work is not always the real reason for the answer.

Researchers have run a clean test. Put a hint in the prompt that changes the model's answer. Then read the rough work. Frequently it never mentions the hint. It presents a tidy, plausible chain of steps that was not the actual cause.

So the written reasoning is useful, and it is not a reliable window into the model's process. Treat it as a working surface, not as a confession.

Where you have already seen this

  • Rough work in the margin of an exam paper.
  • Counting on your fingers.
  • Writing intermediate totals down a column when adding a long bill.

Remember this

  • Each token gets a fixed amount of computation. Rough work buys more steps.
  • Reasoning models were trained to produce rough work by rewarding correct answers.
  • The rough work is genuinely useful and is not a faithful explanation.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Runs on a CPU in about thirty seconds.

Proving the scratchpad does real work

The task is the parity of a 12-bit string. A transformer has to combine information across all 12 positions, and it is a genuinely hard function for a shallow one. The two settings see identical bits and use an identical model.

scratchpad.py
import torch
import torch.nn as nn
import torch.nn.functional as F

# Task: parity of a 12-bit string - is the number of 1s odd?
# DIRECT     : read all 12 bits, then emit the answer in one step.
# SCRATCHPAD : after each bit, emit the running parity so far, and read it back.
# Same bits, same model, same budget. Only the shape of the output differs.
N_BITS = 12
EVEN, SEP, V = 2, 4, 5      # tokens: 0/1 are bits, 2/3 are parity answers, 4 is a separator


def make(n, gen):
    bits = torch.randint(0, 2, (n, N_BITS), generator=gen)
    run = bits.cumsum(1) % 2                                # parity of each prefix
    par = run + EVEN                                        # token 2 or 3

    direct_in = torch.cat([bits, torch.full((n, 1), SEP)], 1)      # 13 tokens
    direct_tgt = par[:, -1]                                        # one label

    inter = torch.stack([bits, par], dim=2).reshape(n, 2 * N_BITS)  # b p b p ...
    return direct_in, direct_tgt, inter, par


def build(layers, seed=0):
    torch.manual_seed(seed)
    layer = nn.TransformerEncoderLayer(64, 4, 128, batch_first=True, dropout=0.0)
    return nn.ModuleDict({"emb": nn.Embedding(V, 64), "pos": nn.Embedding(2 * N_BITS, 64),
                          "blocks": nn.TransformerEncoder(layer, layers),
                          "head": nn.Linear(64, V)})


def fwd(m, x):
    h = m["emb"](x) + m["pos"](torch.arange(x.shape[1]))
    mask = nn.Transformer.generate_square_subsequent_mask(x.shape[1])
    return m["head"](m["blocks"](h, mask=mask, is_causal=True))


def run_direct(layers, steps, gen):
    m = build(layers)
    opt = torch.optim.AdamW(m.parameters(), lr=3e-3)
    x, y, _, _ = make(512, gen)
    for _ in range(steps):
        loss = F.cross_entropy(fwd(m, x)[:, -1], y)
        opt.zero_grad(); loss.backward(); opt.step()
    tx, ty, _, _ = make(2048, torch.Generator().manual_seed(77))
    with torch.no_grad():
        return (fwd(m, tx)[:, -1].argmax(-1) == ty).float().mean().item()


def run_scratchpad(layers, steps, gen):
    m = build(layers)
    opt = torch.optim.AdamW(m.parameters(), lr=3e-3)
    _, _, inter, par = make(512, gen)
    for _ in range(steps):
        out = fwd(m, inter)[:, ::2]              # positions that must emit a parity token
        loss = F.cross_entropy(out.reshape(-1, V), par.reshape(-1))
        opt.zero_grad(); loss.backward(); opt.step()
    _, _, ti, tp = make(2048, torch.Generator().manual_seed(77))
    with torch.no_grad():
        forced = (fwd(m, ti)[:, ::2].argmax(-1)[:, -1] == tp[:, -1]).float().mean().item()
        # free-running: the model must read back its OWN parity tokens, not the true ones
        seq = ti.clone()
        for i in range(N_BITS):
            seq[:, 2 * i + 1] = fwd(m, seq[:, :2 * i + 1])[:, -1].argmax(-1)
        free = (seq[:, -1] == tp[:, -1]).float().mean().item()
    return forced, free


print("parity of 12 bits, chance = 0.500")
print(f"{'layers':>7} {'direct answer':>15} {'scratchpad (fed)':>17} {'scratchpad (own)':>17}")
for layers in (1, 2):
    d = run_direct(layers, 300, torch.Generator().manual_seed(1))
    forced, free = run_scratchpad(layers, 300, torch.Generator().manual_seed(1))
    print(f"{layers:>7} {d:>15.3f} {forced:>17.3f} {free:>17.3f}")
Output
parity of 12 bits, chance = 0.500
 layers   direct answer  scratchpad (fed)  scratchpad (own)
      1           0.492             1.000             1.000
      2           0.501             1.000             1.000

Written against PyTorch 2.5.1, CPU, all seeds fixed — reproducible on this build.

This is the clearest evidence you will see for why reasoning works

Direct answering sat at chance. 0.492 and 0.501 against a 0.500 baseline. Adding a second layer did not help. The model saw every bit; it could not combine them.

The scratchpad version hit 100% with one layer. Identical bits, identical width, identical training budget. The only change is that the running parity is written into the sequence and read back.

Why it works, mechanically. With a scratchpad, each output is previous_parity XOR current_bit — a function of the two tokens immediately to the left. One attention head can do that. Without it, position 12 must compute a 12-way XOR inside one forward pass, which needs depth the model does not have.

The free-running column matters. At 1.000 the model also succeeds when reading back its own emitted tokens rather than the true ones. That is what actually happens at generation time. Had this column collapsed while the fed column stayed at 1.000, the result would have been an artefact of teacher forcing — and this is precisely the check people forget to run.

The honest limit of this demo. Parity is a task where the intermediate steps are perfect and unambiguous. Real chain of thought contains wrong steps that happen to reach right answers, and right steps that reach wrong ones. The mechanism is real; the reliability is not the same thing.

Working with reasoning models in practice

The thinking region is delimited by special tokens, and the exact form is model-specific.

python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-0.6B")
msgs = [{"role": "user", "content": "Capital of France?"}]
print(repr(tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)))
Output
'<|im_start|>user\nCapital of France?<|im_end|>\n<|im_start|>assistant\n'

Qwen3's template opens the assistant turn and lets the model emit its own <think> block. Serving stacks then strip the region between the thinking tokens before returning the answer, while still billing you for it.

Three practical consequences:

You pay for tokens you never see. A reasoning model on a hard problem can emit thousands of hidden tokens for a two-line answer. Budget for it.

Do not feed the thinking region back into the next turn. Model cards are explicit about this: keep the answer, discard the reasoning, in multi-turn conversations.

Chain-of-thought prompting is largely redundant on these models. "Think step by step" was a prompt trick for non-reasoning models. On a model trained to reason, it is at best a no-op.

Common mistakes

Setting max_new_tokens too low. The model gets cut off mid-thought and returns nothing usable. Reasoning models need generous limits.

Sampling at temperature 0. Reasoning models are usually released with recommended sampling settings — Qwen3, for example, recommends non-greedy decoding for thinking mode. Greedy decoding can send them into repetition loops. Read the model card.

Parsing the answer by taking the last line. Extract from the answer markers the model was trained to emit, not from position.

Using a reasoning model for extraction or classification. Paying for a thousand hidden tokens to output one label is a waste. Use a small non-reasoning model.

Trusting the reasoning as an explanation. See the researcher block. It is not one.

Try it yourself

Change N_BITS to 6 and re-run the direct setting. A shorter string is an easier function, and the direct model may learn it. Then work upward and find the length where it fails. That length is your model's effective serial-computation depth.

What to learn next

Researcher — Mathematics and papers.

Why extra tokens are extra computation

A transformer's forward pass is a fixed circuit: depth $L$, width $d$, one pass per token. The class of functions computable in one pass is bounded — the standard result is that constant-depth transformers with bounded precision sit inside the complexity class $\mathsf{TC}^0$ (Merrill and Sabharwal, 2023), which does not contain problems believed to require serial computation, including exact parity of arbitrary length under some formalisations.

Chain of thought changes the model of computation. Feng et al., 2023 (Towards Revealing the Mystery behind Chain of Thought) and Merrill and Sabharwal, 2024 (The Expressive Power of Transformers with Chain of Thought) prove the key result: a constant-depth transformer generating $T$ intermediate tokens can simulate a computation of $O(T)$ serial steps. With a polynomial number of intermediate tokens, the reachable class rises to $\mathsf{P}$.

The parity experiment above is that theorem in miniature: an $O(n)$ serial computation, impossible in one pass at this depth, trivial when unrolled across $n$ tokens.

Corollary that surprises people: Pfau et al., 2024 (Let's Think Dot by Dot) show that filler tokens — literal ... sequences carrying no semantic content — can also recover performance on certain parallelisable problems, though they require dense supervision to learn. The extra computation, not the meaning of the words, is doing part of the work.

From prompting to training

Wei et al., 2022 (Chain-of-Thought Prompting Elicits Reasoning in Large Language Models) showed few-shot exemplars containing reasoning steps improved GSM8K sharply, and that the effect only appears above roughly 100B parameters. Kojima et al., 2022 (Large Language Models are Zero-Shot Reasoners) reduced this to the phrase "Let's think step by step". Wang et al., 2023 (Self-Consistency) added majority voting over sampled chains for a further large gain.

Nye et al., 2021 (Show Your Work: Scratchpads for Intermediate Computation) had already established the training-time version, on exactly the kind of algorithmic tasks the code above uses.

The 2024–2025 shift was to make it a trained policy rather than a prompt. OpenAI o1 (2024) and DeepSeek-R1 (2025) both train the model to allocate its own thinking length, with verifiable rewards supplying the signal. R1's response length grew from hundreds to thousands of tokens over training with no length reward — the model discovered that longer chains raised accuracy.

Test-time scaling

Snell et al., 2024 (Scaling LLM Test-Time Compute Optimally) treat inference compute as a resource to be allocated, and find that for easy and medium problems, optimally-allocated test-time compute can outperform a 14× larger model. For the hardest problems, more parameters still win.

Muennighoff et al., 2025 (s1: Simple test-time scaling) provide the most striking efficiency result: 1,000 curated reasoning examples plus "budget forcing" — appending the token "Wait" to extend thinking, or forcing an end-of-thinking token to truncate it — produced a model exceeding o1-preview on competition maths. Two implications: the reasoning behaviour is substantially elicitable rather than only learnable through long RL, and thinking length is a controllable inference parameter.

There is a limit. Reported test-time scaling curves flatten, and past that point additional thinking tokens cost money without improving accuracy. Overthinking — long chains on trivial questions — is a documented and measurable regression in reasoning models.

Faithfulness

The most important caveat, and the one most often omitted.

Turpin et al., 2023 (Language Models Don't Always Say What They Think) inserted biasing features into prompts — for example, making the correct multiple-choice answer always option (A) in the few-shot examples. Model accuracy dropped by up to 36% under the bias, and the generated chains of thought systematically failed to mention the biasing feature, instead constructing plausible justifications for the biased answer.

Lanham et al., 2023 (Measuring Faithfulness in Chain-of-Thought Reasoning) probed this by truncating or corrupting the reasoning and observing whether the answer changed. Faithfulness varies with task and inversely with model size: larger models more often reach the same answer regardless of the stated reasoning.

Baker et al., 2025 add the operational finding for RL-trained reasoners: chain-of-thought monitoring detects reward hacking well, but training against the monitor produces models that hide their intent while continuing to hack. Their recommendation — leave the chain of thought out of the training objective and monitor it separately — is the current best practice.

The practical position: reasoning traces are useful for debugging and monitoring, and are not evidence about why a model answered as it did.

Cost accounting

For a dense model, generation cost is roughly $2N$ FLOPs per token plus attention over the growing KV cache. Reasoning multiplies the token count, so:

  • A 10× longer output is a 10× higher inference cost and a 10× higher latency.
  • KV cache memory grows linearly in the thinking length, which caps the achievable batch size — see compressing the KV cache.
  • The prefill/decode split matters more, because reasoning models are overwhelmingly decode-bound.

This is why "thinking budget" controls now appear in serving APIs. It is a direct cost lever.

Papers

What to learn next