Constrained decoding and grammars
Instead of asking a model nicely for valid JSON, you delete every token that would break the format before it can be chosen, which makes malformed output impossible rather than unlikely.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Constrained decoding blocks every word that would break the required format, so the model can only produce something valid.
Think about the boxes on a bank form. One box takes a digit and nothing else. You physically cannot write a letter there. The box is one small square, and it allows nothing else.
Compare that to a blank sheet with "please write your account number" at the top. Most people get it right. Some write it with spaces, some add a note, one person writes it in words.
Constrained decoding turns the blank sheet into boxes.
The problem it solves
You ask a model for JSON — a strict text format that programs can read. You write "reply with only JSON" in capital letters. Nine hundred and ninety times out of a thousand it works.
The other ten times you get a friendly sentence before the JSON, or a missing bracket, or a trailing comma. Your program crashes.
Retrying is the usual fix. Retrying costs money, adds delay, and does not actually guarantee anything.
How it works
Remember that a model produces a score for every word it knows, at every step. Sampling picks from those scores.
Constrained decoding adds one step in between. Before picking, it asks: given what has been written so far, which of these are even legal? Everything else has its score set so low that it can never be chosen.
we have written: {"name":"
legal next tokens: Asha Ravi
everything else: blocked
the model's own preference decides between Asha and Ravi.
the format decides that nothing else is possible.The model keeps all its judgement about what to say. It loses only the ability to break the shape.
Why this is better than asking politely
It cannot fail. Malformed output is not unlikely, it is unreachable.
It is free. Blocking tokens costs a fraction of a millisecond per step, against a whole extra round trip for a retry.
It works on small models. Draw the boxes and a tiny model becomes reliable. On instructions alone it could never be trusted to produce clean JSON.
The catch, stated honestly
You are overruling the model. That is the point, and it has a cost.
If the format demands a number and the model wanted to say "unknown", it now has to invent a number. The output is valid and the content is worse.
So a constraint is not a substitute for a model that understands the task. Give the format an honest escape route. A field that can be null, or an option meaning "not found". Then the constraint stops forcing lies.
Where you have already seen this
- An API setting that guarantees JSON matching a schema you supply.
- A tool-calling feature where arguments always arrive in the right shape.
- A local model tool with a "grammar" file for output format.
- A form-filling assistant that never produces an invalid date.
Remember this
- Constrained decoding deletes illegal tokens before the model chooses.
- Invalid output becomes impossible, not only unlikely.
- Forcing a shape can force a wrong answer. Leave the format an honest way to say "I do not know".
What to learn next
- Structured output — the application layer built on this mechanism.
- Function calling — tool arguments, which are constrained decoding in disguise.
- Reading logprobs — inspecting the distribution a constraint has reshaped.
Developer — Code and libraries.
Setup
pip install numpyThe mechanism is a state machine plus a mask. Building it once makes every library in this space obvious afterwards.
A grammar, a mask, and a model with bad instincts
import numpy as np
VOCAB = ['{', '}', '"', ':', ',', 'name', 'age', 'Asha', 'Ravi',
'1', '2', '3', '4', '5', '6', '7', '8', '9', '0', ' Sure,', '</s>']
ID = {t: i for i, t in enumerate(VOCAB)}
DIGITS = [ID[d] for d in '0123456789']
# The grammar for {"name": "<one of two names>", "age": <1-2 digits>}
# state -> {token id: next state}
GRAMMAR = {
0: {ID['{']: 1},
1: {ID['"']: 2},
2: {ID['name']: 3},
3: {ID['"']: 4},
4: {ID[':']: 5},
5: {ID['"']: 6},
6: {ID['Asha']: 7, ID['Ravi']: 7},
7: {ID['"']: 8},
8: {ID[',']: 9},
9: {ID['"']: 10},
10: {ID['age']: 11},
11: {ID['"']: 12},
12: {ID[':']: 13},
13: {d: 14 for d in DIGITS if d != ID['0']},
14: {**{d: 15 for d in DIGITS}, ID['}']: 16},
15: {ID['}']: 16},
16: {ID['</s>']: 17},
}
rng = np.random.default_rng(11)
def model_scores(step):
"""Stands in for a real model. It has opinions, and they are not JSON-shaped."""
s = rng.normal(size=len(VOCAB))
s[ID[' Sure,']] += 3.0 # every chat model's favourite opening
s[ID['</s>']] += 1.0
return s
def decode(constrained, max_steps=20):
state, out = 0, []
for step in range(max_steps):
scores = model_scores(step)
if constrained:
mask = np.full(len(VOCAB), -np.inf)
allowed = GRAMMAR.get(state, {})
if not allowed:
break
for tok in allowed:
mask[tok] = 0.0
scores = scores + mask # every illegal token becomes impossible
nxt = int(scores.argmax())
if VOCAB[nxt] == '</s>':
break
out.append(VOCAB[nxt])
state = GRAMMAR.get(state, {}).get(nxt, -1)
return ''.join(out)
print("unconstrained greedy:", repr(decode(constrained=False)))
rng = np.random.default_rng(11) # same model, same scores, same seed
print("constrained greedy:", repr(decode(constrained=True)))
# What the mask looks like at one step
rng = np.random.default_rng(11)
state = 0
print("\nlegal tokens at each state, which is the entire mechanism:")
for state in (0, 6, 13, 14):
legal = [VOCAB[t] for t in GRAMMAR[state]]
print(f" state {state:>2}: {len(legal):>2} of {len(VOCAB)} tokens legal -> {legal}")unconstrained greedy: ' Sure, Sure, Sure,4 Sure, Sure, Sure,3 Sure, Sure, Sure, Sure, Sure, Sure,1 Sure, Sure, Sure, Sure, Sure,'
constrained greedy: '{"name":"Ravi","age":51}'
legal tokens at each state, which is the entire mechanism:
state 0: 1 of 21 tokens legal -> ['{']
state 6: 2 of 21 tokens legal -> ['Asha', 'Ravi']
state 13: 9 of 21 tokens legal -> ['1', '2', '3', '4', '5', '6', '7', '8', '9']
state 14: 11 of 21 tokens legal -> ['0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '}']Reading the output
Same model, same seed, same scores. Two completely different results. The unconstrained run produced garbage. The constrained run produced {"name":"Ravi","age":51}, which parses. Nothing about the model changed; only the set of legal moves did.
The model still made every real decision. State 6 offered Asha and Ravi, and the model's own scores picked Ravi. State 13 offered nine digits and the model picked 5. Constraints removed choices about form, never choices about content.
State 13 forbids 0. JSON has no leading zeros on numbers, so age: 05 is invalid. The state machine encodes that. It is the sort of rule nobody remembers to put in a prompt, and every real JSON grammar has to handle.
State 14 allows a digit or a closing brace. That branch is what makes both 5 and 51 valid ages. A grammar is a graph, not a checklist, and this branching is where the interesting bugs live.
The unconstrained output is a caricature and also the real failure. A chat model's most likely opening for any request is a polite acknowledgement. That is exactly why "reply with only JSON" fails, and why the failure gets worse on smaller and more heavily chat-tuned models.
The hard part, which the toy hides
Real vocabularies have a hundred thousand or more tokens, and tokens are not neat single characters. One token might be ":" or ,"na. The grammar is written over characters; the mask has to be over tokens.
Bridging that is the whole engineering problem, and it is why you use a library:
- Outlines (Willard and Louf, 2023) compiles a regular expression or JSON schema into a finite-state machine, then precomputes, for each state, the set of allowed token ids. Runtime cost becomes a table lookup.
- XGrammar (Dong et al., MLSys 2025) uses a pushdown automaton for full context-free grammars. It splits the vocabulary into context-independent tokens, over ninety-nine percent of them, whose validity is decided at compile time. Only a tiny remainder is checked per step.
- llama.cpp ships GBNF, a grammar file format, with the same idea implemented in C.
Using it for real
# vLLM (v0.12+ API), offline
from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams
schema = {
"type": "object",
"properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
"required": ["name", "age"],
}
params = SamplingParams(
temperature=0.0,
structured_outputs=StructuredOutputsParams(json=schema),
)# Any OpenAI-compatible endpoint, including a vLLM server
resp = client.chat.completions.create(
model="...",
messages=[{"role": "user", "content": "Extract name and age."}],
response_format={"type": "json_schema", "json_schema": {"name": "person",
"schema": schema,
"strict": True}},
)No output blocks for these two, deliberately. Both need a served model, and the text depends on which model. Any output printed here would be invented rather than observed.
One API note that will bite you: vLLM renamed these fields. guided_json, guided_regex, guided_choice, guided_grammar and guided_decoding_backend were removed in v0.12.0 in favour of StructuredOutputsParams with json, regex, choice and grammar. Tutorials written before that release will not run. Pin your vLLM version and read its own docs — this corner of the API has moved more than once.
Common mistakes
Constraining without an escape hatch. A schema requiring "age": integer forces a number even when the document has no age. Allow ["integer", "null"] and the model can be honest.
Assuming valid means correct. Constrained output always parses. It can still be wrong. Schema validity and factual accuracy are unrelated properties, and shipping the first while measuring nothing about the second is a common failure.
Using a regex where a grammar is needed. Regular expressions cannot count nested brackets. Arbitrary nested JSON needs a context-free grammar, which is why XGrammar's pushdown automaton exists.
Leaving penalties on. JSON needs " and , repeatedly. A repetition penalty fights the grammar and produces strange field names. Set all penalties to zero for structured output.
Forgetting the whitespace explosion. JSON permits whitespace almost anywhere, so a permissive grammar lets the model emit spaces and newlines indefinitely without ever violating the format. Restrict whitespace in the grammar, and always set a token limit.
Try it yourself
Add a third name to state 6 and a "city" field after age. Then break it on purpose: make state 14 loop to itself on digits with no exit to }, and watch the decoder run to max_steps producing a number that never ends. A grammar with no reachable terminal state is the most common bug in hand-written grammars.
What to learn next
- Structured output — the application layer built on this mechanism.
- Function calling — tool arguments, which are constrained decoding in disguise.
- Reading logprobs — inspecting the distribution a constraint has reshaped.
Researcher — Mathematics and papers.
The formulation
Let $\mathcal{L} \subseteq V^{*}$ be the target language and $y_{<t}$ the tokens produced so far. Define the set of tokens that keep the prefix extensible:
$$ A(y_{<t}) = { v \in V : \exists\, s \in V^{*} \text{ such that } y_{<t} \cdot v \cdot s \in \mathcal{L} } $$
Constrained decoding replaces $P(\cdot \mid y_{<t})$ with
$$ P'(v \mid y_{<t}) = \frac{P(v \mid y_{<t}) \cdot \mathbb{1}[v \in A(y_{<t})]}{\sum_{u \in A(y_{<t})} P(u \mid y_{<t})} $$
The condition uses extensibility, not local validity. A token that produces a valid prefix from which no complete string is reachable must also be masked, or the decoder walks into a dead end and cannot terminate. Hand-rolled implementations get this wrong constantly.
Why the token boundary is the hard part
The constraint is naturally expressed over characters; the model emits tokens. A single token may span a grammar boundary. ": is one token in many byte-level BPE vocabularies. So the automaton must be advanced by a token's entire character string, and a token is legal only if every intermediate transition is legal.
Outlines (Willard and Louf, 2023, arxiv.org/abs/2307.09702) solve this by indexing the FSM over the vocabulary once at compile time: for each state, walk every token's character string through the automaton and record which tokens are accepted and where they land. Runtime per step is then $O(1)$ in vocabulary size, with the cost moved into compilation.
XGrammar (Dong et al., MLSys 2025, arxiv.org/abs/2411.15100) extends this to context-free grammars using a pushdown automaton. It adds a key observation: the overwhelming majority of vocabulary tokens are context-independent, meaning their validity does not depend on the stack. Those are precomputed; only a small context-dependent remainder is checked at runtime, with a persistent execution stack to handle backtracking. Reported mask-generation time is under 40 microseconds for JSON schemas.
The distributional objection
Masking and renormalising does not sample from "the model's distribution restricted to $\mathcal{L}$" in the sense people usually assume. The correct conditional would be
$$ P(y \mid y \in \mathcal{L}) = \frac{P(y) \cdot \mathbb{1}[y \in \mathcal{L}]}{\sum_{y' \in \mathcal{L}} P(y')} $$
which requires normalising over complete sequences. Greedy per-step masking normalises locally instead. The two differ whenever a prefix's mass on now-illegal continuations is redistributed differently from how a global conditional would do it.
The practical consequence is measurable. Constrained generation can degrade content quality relative to an unconstrained model that happens to produce valid output. Tam et al., Let Me Speak Freely?, 2024 (arxiv.org/abs/2408.02442) report reasoning degradation under format restriction; other work reports the opposite on other tasks. The honest summary is that the effect is real, task-dependent, and confounded with prompt changes made alongside the constraint. Measure it on your own task; do not assume either direction.
A partial mitigation that costs nothing: allow a free-text reasoning field before the constrained fields in the schema. The model gets to think in unconstrained tokens, and the machine-readable part still cannot be malformed.
Related mechanisms
- Token healing. A prompt ending mid-token (
"http) forces the model to continue from a tokenisation it would never have produced itself. Token healing backs up to the last complete token boundary and re-masks. Present in Guidance and intransformersastoken_healing. - Structural tags. Constrain only the region inside a marker, leaving the rest free. vLLM exposes this as
structural_tag, and it is how modern tool-calling formats are enforced without constraining the assistant's prose. - Speculative decoding interaction. Draft tokens must be validated against the automaton as well as against the target model, so an aggressive grammar lowers acceptance rates. Combining the two is supported but the speed-up shrinks.
- Semantic constraints. Grammars enforce syntax. Foreign-key validity, unit consistency and cross-field arithmetic are not regular or context-free properties, and they still need validation after generation.
Papers and tools
- Willard and Louf, Efficient Guided Generation for Large Language Models, 2023 — arxiv.org/abs/2307.09702
- Dong et al., XGrammar: Flexible and Efficient Structured Generation Engine for LLMs, MLSys 2025 — arxiv.org/abs/2411.15100
- Tam et al., Let Me Speak Freely? A Study on the Impact of Format Restrictions, 2024 — arxiv.org/abs/2408.02442
- Geng et al., Grammar-Constrained Decoding for Structured NLP Tasks, 2023 — arxiv.org/abs/2305.13971
What to learn next
- Structured output — the application layer built on this mechanism.
- Function calling — tool arguments, which are constrained decoding in disguise.
- Reading logprobs — inspecting the distribution a constraint has reshaped.