LLM Development

Structured output

Structured output means forcing a model to answer inside a shape you defined, so your program can read the answer instead of guessing at a paragraph.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The thing that surprises everyone
  6. Where you have already seen it
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Structured output means the model must answer inside a shape you chose, instead of writing free paragraphs.

The analogy you have already lived

Think of a bank form. Date of birth is not a blank line. It is eight little boxes: D D M M Y Y Y Y.

You cannot write "sometime in the monsoon of 2003" there. The boxes force one shape, and that is exactly why the clerk can read a thousand forms quickly.

Now think of an exam OMR sheet. Round bubbles, one per option. The machine reads it in a second because it never has to interpret handwriting.

Structured output puts a model inside those boxes.

Why it exists

A model's normal reply is written for a human. It is friendly, and that friendliness is the problem.

Ask for the details of a support email as data, and you often get a reply shaped like this:

Sure! Here's the ticket for you:

    (a line of three backticks, then the word json)
    { "subject": "Refund not received", "urgency": "high" }
    (another line of three backticks)

Let me know if you'd like me to change anything!

A person reads that fine. Your program does not. It tries to read the whole thing as data and crashes on the word "Sure".

There are three separate problems in that one reply.

Wrapping text. "Sure! Here's..." and "Let me know..." are not part of the answer.

Code fences. Those backtick lines are decoration for a chat window, not data.

Wrong value type. "urgency": "high" — your database column expects a number from 1 to 5.

People used to attack this with clever prompt wording. "Reply with JSON only. No explanation. Do not use markdown." It works about 95% of the time. Then it fails on request 4,000 at midnight.

Structured output replaces begging with a guarantee.

How it works

  WITHOUT a shape
  ---------------
  question  ->  [ model ]  ->  Sure! Here's the JSON, wrapped in backticks.
                               Hope that helps!
                                       |
                                  your code
                                       |
                                    CRASH


  WITH a shape
  ------------
  question + the exact shape you want
            ->  [ model ]  ->  {"subject": "...", "category": "billing", "urgency": 4}
                                       |
                                  your code
                                       |
                                   works, every time

You send the shape along with the question. The serving software then only allows the model to produce characters that keep the answer valid. A closing brace becomes the only legal choice at the point where the shape says it should close.

That is a real mechanical guarantee, not a polite request. It is called constrained decoding.

The thing that surprises everyone

A valid shape is not a true answer.

Asked for facts about Pune inside a shape with a "population in millions" number field, a small model returned this:

{ "city": "Pune", "country": "India", "population_millions": 3.5 }

Perfectly shaped. The number is a guess. A different run gave a different number.

The boxes on a bank form do not make what you write in them true. They only make it readable. Keep those two ideas apart and you will avoid the most expensive mistake in this whole area.

Where you have already seen it

  • An app that reads a photo of a bill and fills in amount, date and shop name.
  • A hiring tool that turns a resume into fields.
  • A support system that tags an incoming email with a category and priority.
  • Any AI feature whose result appears in a table rather than a chat bubble.

Remember this

  • Structured output makes the model answer inside a shape you defined.
  • It guarantees the shape, never the truth of the values.
  • Always check the values in your own code afterwards.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pydantic

Pydantic is how Python programmers describe a shape. It generates the JSON Schema you send to the model, and it validates what comes back. Both directions from one class definition.

For the live example you also need Ollama running.

The problem, then the fix

This script needs no model. Every line of its output is fixed.

validate.py
import json
from typing import Literal
from pydantic import BaseModel, Field, ValidationError


class Ticket(BaseModel):
    subject: str
    category: Literal["billing", "delivery", "technical"]
    urgency: int = Field(ge=1, le=5)


# Exactly what llama3.2:1b returned when asked for "JSON only", with no schema attached.
unconstrained = '```json\n{\n  "city": "Pune",\n  "country": "India"\n}\n```'

try:
    json.loads(unconstrained)
except json.JSONDecodeError as exc:
    print("json.loads failed:", exc)

print()

# Three replies to validate: one good, two broken in different ways.
replies = [
    '{"subject": "Refund not received", "category": "billing", "urgency": 4}',
    '{"subject": "App keeps crashing", "category": "bugs", "urgency": 9}',
    '{"subject": "Parcel missing"}',
]

for raw in replies:
    try:
        print("OK  ", Ticket.model_validate_json(raw))
    except ValidationError as exc:
        print("FAIL", [(e["loc"], e["type"]) for e in exc.errors()])
Output
json.loads failed: Expecting value: line 1 column 1 (char 0)

OK   subject='Refund not received' category='billing' urgency=4
FAIL [(('category',), 'literal_error'), (('urgency',), 'less_than_equal')]
FAIL [(('category',), 'missing'), (('urgency',), 'missing')]

What each line is telling you

Expecting value: line 1 column 1 (char 0). This is the single most common error in this whole topic, and the message is misleading. Nothing is wrong with the JSON. The parser choked on the backtick at position zero. Print the first 50 characters of any reply before you parse it and this stops being mysterious.

literal_error and less_than_equal. The model chose "bugs", which is not in your three allowed categories, and 9, which is outside 1 to 5. Both were caught. Notice that both errors came back from one call. Pydantic reports every problem, not only the first, which is what makes automatic repair possible.

missing. Two fields absent. Truncated output does this, usually because max_tokens was too low.

e["type"] rather than str(exc). Pydantic's full message includes a documentation URL with the library version in it. Reading the stable machine-readable fields keeps your code working across upgrades.

Now force the shape

Ollama accepts a JSON Schema in a format field, and Pydantic generates that schema for you.

extract.py
import json
import urllib.request
from typing import Literal
from pydantic import BaseModel, Field


class Ticket(BaseModel):
    subject: str
    category: Literal["billing", "delivery", "technical"]
    urgency: int = Field(ge=1, le=5)


EMAIL = "Hi, I was charged twice for order 8812 and nobody has replied for a week."

body = {
    "model": "llama3.2:1b",
    "messages": [{"role": "user", "content": f"Turn this support email into a ticket.\n\n{EMAIL}"}],
    "stream": False,
    "format": Ticket.model_json_schema(),      # the shape the model must fill
    "options": {"temperature": 0},
}

request = urllib.request.Request(
    "http://localhost:11434/api/chat",
    data=json.dumps(body).encode(),
    headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request) as response:
    raw = json.loads(response.read())["message"]["content"]

print("raw text from the model:")
print(raw)
print()
print("after validation:", Ticket.model_validate_json(raw))
Output
raw text from the model:
{"subject": "Order 8812 Charge and Delay in Response", "category": "billing", "urgency": 2}

after validation: subject='Order 8812 Charge and Delay in Response' category='billing' urgency=2

The subject wording and the urgency number will differ on your machine. They change with the model, the version and the run. What does not change is that the output parses, has all three fields, and passes validation. That is the guarantee you bought.

Now argue with the result. A customer charged twice and ignored for a week is not urgency 2. The shape was enforced perfectly and the judgement was poor. This is the point from the Beginner tab, in your own terminal.

The same idea on the paid APIs

The mechanism differs slightly, the concept does not.

python
# OpenAI-compatible servers, including Ollama's /v1 endpoint and vLLM
completion = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": EMAIL}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "ticket", "strict": True,
                        "schema": Ticket.model_json_schema()},
    },
)

strict: True is what turns a suggestion into a guarantee. Without it you are back to hoping.

Anthropic's models take the same shape through a tool definition: describe one tool whose input schema is your object, and force its use.

Repairing instead of crashing

When validation fails, you have three moves. Try them in this order.

  1. Retry with the error text attached. Send the model its own broken output and the validation errors. Most models fix it on the second attempt.
  2. Fall back to a stricter mode. If you were prompting for JSON, switch to a schema-constrained call.
  3. Give up cleanly. Return a null result your application can handle. A wrong record is worse than no record.
python
def extract(email, attempts=2):
    messages = [{"role": "user", "content": email}]
    for _ in range(attempts):
        raw = call_model(messages)
        try:
            return Ticket.model_validate_json(raw)
        except ValidationError as exc:
            # The model reads its own mistake and corrects it on the next turn.
            messages += [{"role": "assistant", "content": raw},
                         {"role": "user", "content": f"That failed validation: {exc.errors()}. Fix it."}]
    return None

Count your retries and watch that number. A repair rate creeping upward is a schema problem, not a model problem.

Common mistakes

Parsing without looking. Print the raw reply before json.loads. The backtick problem is visible in one glance and invisible in a stack trace.

Confusing valid with correct. Constrained decoding cannot produce a fact. Add range checks, cross-field rules and business logic on top. Pydantic validators are the right place.

Deeply nested schemas. Accuracy falls with nesting depth. Three flat calls beat one deeply nested call, most of the time.

Optional fields everywhere. Making everything Optional stops validation errors and hides the failure instead. You then discover the empty fields much later, in your database.

max_tokens too small. The model gets cut off mid-object. You see missing errors on the trailing fields and blame the schema.

Free-text fields with no length limit. A model asked for a "summary" with no bound will sometimes write four paragraphs. Use Field(max_length=200).

Assuming every server supports it. Older or unusual servers ignore format and response_format silently. Test with a deliberately impossible schema and confirm it actually fails.

Try it yourself

Add a refund_amount: float | None field and a mentioned_order_id: str | None field to Ticket. Run extract.py on five support emails you write yourself, including one with no order number at all.

Then add a Pydantic validator that rejects a refund amount above 50,000. That single rule is the line between "the shape was right" and "the answer was right", and writing it will make the distinction permanent.

What to learn next

Researcher — Mathematics and papers.

Three mechanisms, three different guarantees

The phrase "structured output" covers approaches with genuinely different failure probabilities.

Prompt instruction. Ask for JSON in the system prompt. No guarantee. Empirically 90–99% valid depending on model and schema complexity, and the failures cluster on exactly the inputs you care about — long, unusual, adversarial.

JSON mode. A coarse constraint that the output be syntactically valid JSON. Removes prose wrappers and fences. Says nothing about which keys or types appear.

Grammar-constrained decoding. The schema is compiled to a formal grammar, and at each decoding step the logit vector is masked so only tokens that can continue a valid string retain non-zero probability. Validity is then a structural property of the sampler, not a behaviour of the model. This is what strict: true and Ollama's format: <schema> invoke.

How constrained decoding is implemented

Outlines (Willard and Louf, 2023) compiles a regular expression — or a JSON Schema lowered to one — into a deterministic finite automaton, then precomputes, for each FSM state, the set of vocabulary tokens whose string continues a valid path. Per-step cost is then an index lookup and a mask, effectively free relative to a forward pass.

The subtlety is tokenisation. Constraints are defined over characters; sampling happens over tokens, and a token may span a state boundary. The FSM must therefore be built over token boundaries rather than characters, which is why these libraries maintain vocabulary-specific compiled indices.

XGrammar and llama.cpp's GBNF push further to context-free grammars, needed for recursive schemas that a regular language cannot express — arbitrarily nested arrays and objects, for example.

What masking does to the distribution

Masking renormalises over the allowed set:

p'(t) = p(t) · 1[t ∈ A_s] / Σ_{t' ∈ A_s} p(t')

A_s is the allowed token set in FSM state s. This is exact conditioning on the constraint only if the constraint set is closed under the model's own preferences at every step. It is not, in general.

The practical consequence: forcing a low-probability token early shifts the conditional distribution for everything after it. Reported effects range from negligible to material. Two mitigations are well supported.

Order schema fields so that reasoning-bearing fields precede conclusion fields — a reasoning string before a category enum lets the model condition its answer on its own generated argument, and measurably improves accuracy on classification-style extraction.

And prefer field names and enum values that appear naturally in the model's output distribution. "billing" costs less probability mass than "CAT_BILLING_V2".

Validity and correctness are orthogonal

A grammar constrains the codomain of the generator. It cannot constrain the map.

Every extraction system therefore needs a second layer, and the layers catch different things:

  • Schema validation — types, required keys, enum membership, numeric bounds.
  • Semantic validation — cross-field consistency, referential integrity against your own data, business rules.
  • Confidence estimation — token log-probabilities over the value span, or self-consistency across k samples with disagreement as an abstention signal.

Self-consistency (Wang et al., 2022) is the most reliable of the three for detecting fabricated values: sample k times at non-zero temperature, and treat field-level disagreement as a flag for human review. It costs k× inference, which is why it is usually applied only to high-stakes fields.

Evaluation

Measure these separately, because interventions move them in opposite directions.

  • Schema-validity rate — should be 1.0 under true constrained decoding. Anything less means the constraint is not actually engaged.
  • Field-level exact match against a labelled set, reported per field. Aggregate accuracy hides that one field carries all the error.
  • Hallucination rate on nullable fields — how often a value is invented where the correct answer is null. Constrained decoding makes this worse if the field is required, because the grammar forbids omission and the model must write something.

That last point is the most under-appreciated result in this area. Marking a field required does not make the information present in the input. It converts a missing value into a fabricated one. Make fields optional when the source may genuinely lack them.

References

What to learn next