Agentic Frameworks

DSPy — programming, not prompting

DSPy replaces hand-written prompts with declared inputs and outputs, then tunes the actual prompt wording automatically against a score you define.

On this page 6
  1. Why it exists
  2. How it works
  3. The part you have already lived
  4. The honest part
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

DSPy is a framework where you declare what a model call should take in and give out, and the framework writes and tunes the actual prompt for you.

Think about getting a kurta stitched. One way: stand beside the tailor and dictate every cut — a little more here, redo that seam. You will be there all day, and with a new tailor you start from zero. The other way: hand over your measurements, show a photo of the fit you want, and come for one fitting. The tailor's own craft finds the cuts.

Hand-written prompting is the first way. DSPy is the second — you specify the outcome and the checks, not the wording.

Why it exists

Anyone who has built with prompt engineering knows the cycle. You fiddle with the wording, it works, you ship it. The model gets updated, and your carefully fiddled prompt quietly gets worse. Or you switch providers, and everything breaks.

The deeper problem: the prompt wording is a setting that was never measured, only vibed. DSPy's founders at Stanford asked — what if the wording were like a model's weights? Not written by hand, but found by a search that chases a score.

How it works

Three pieces, and the vocabulary is worth memorising.

  • A signature — the declaration: "this takes a review in, gives a sentiment out". No prompt text. Only the shape.
  • A module — a strategy for filling that signature. Predict asks plainly. ChainOfThought makes the model reason step by step first.
  • An optimizer — the tuner. Give it examples and a metric (a function that scores an answer), and it rewrites instructions and picks worked examples until the score climbs.
   signature: "review -> sentiment"
        │
        ▼
   [ module ]  ──► a first, machine-written prompt
        │
        ▼
   [ optimizer + your examples + your metric ]
        │
        ▼
   a better prompt — found, not hand-written

The part you have already lived

You have played the optimizer yourself. Every time you rephrased a request to a chatbot four times until it behaved, you were searching the space of wordings by hand. DSPy's claim is that this search is a machine's job — the same shift as a camera's autofocus replacing the focus ring. You still choose what to point at. The fiddling is delegated.

The honest part

DSPy shifts the work; it does not remove it. Instead of writing prompts, you now write the metric — the definition of "good" — and collect real examples. That is genuinely harder than it sounds, and a sloppy metric gets optimised into sloppy behaviour, efficiently. Teams that thrive with DSPy are the ones willing to measure. If you will not build a test set, the tailor has no fitting to work from.

Remember this

  • A signature declares in and out; the prompt text is generated, not written.
  • A module is the asking strategy; swapping strategy does not change your declaration.
  • The optimizer tunes wording against your metric and examples — you define good, it chases it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install dspy

Written against dspy 3.3. The install is heavier than you might expect — roughly a quarter of a gigabyte with dependencies, most of it the LiteLLM layer DSPy uses to talk to a hundred-odd providers.

DSPy ships a scripted stand-in model, DummyLM, used in its own test suite. The examples below use it so they run offline, free, and identically for everyone. Where a real model changes things, the lesson says so.

Signatures and modules

sig_demo.py
import dspy
from dspy.utils.dummies import DummyLM

# a stand-in model that returns scripted answers, so this runs offline
lm = DummyLM([
    {"sentiment": "positive"},
    {"reasoning": "The words 'cold' and 'watery' describe a bad meal.",
     "sentiment": "negative"},
])
dspy.configure(lm=lm)

classify = dspy.Predict("review -> sentiment")
out = classify(review="The dosa was crisp and the chutney was fresh.")
print("Predict:        ", out.sentiment)

classify_cot = dspy.ChainOfThought("review -> sentiment")
out = classify_cot(review="The sambar was cold and watery.")
print("ChainOfThought: ", out.sentiment)
print("its reasoning:  ", out.reasoning)
Output
Predict:         positive
ChainOfThought:  negative
its reasoning:   The words 'cold' and 'watery' describe a bad meal.

Look at what is absent: no prompt string anywhere. "review -> sentiment" is the entire specification. And note what ChainOfThought did — the same signature grew a reasoning output field, because the module's strategy is "reason first, then answer". Swap modules, keep the declaration.

See the prompt DSPy wrote

The prompt exists — DSPy generated it. inspect_history shows the last real exchange:

prompt_peek.py
import dspy
from dspy.utils.dummies import DummyLM

dspy.configure(lm=DummyLM([{"sentiment": "positive"}]))
classify = dspy.Predict("review -> sentiment")
classify(review="Great filter coffee.")
dspy.inspect_history(n=1)

The output below is what appears in the terminal, minus its colour codes and timestamp:

Output
System message:

Your input fields are:
1. `review` (str):
Your output fields are:
1. `sentiment` (str):
All interactions will be structured in the following way, with the appropriate values filled in.

[[ ## review ## ]]
{review}

[[ ## sentiment ## ]]
{sentiment}

[[ ## completed ## ]]
In adhering to this structure, your objective is: 
        Given the fields `review`, produce the fields `sentiment`.


User message:

[[ ## review ## ]]
Great filter coffee.

Respond with the corresponding output fields, starting with the field `[[ ## sentiment ## ]]`, and then ending with the marker for `[[ ## completed ## ]]`.


Response:

[[ ## sentiment ## ]]
positive

This is the machine-written prompt your field names turned into — which is also why field names matter: review -> sentiment produces a sensible prompt, x -> y produces a baffling one. Name your fields the way you would name them for a colleague.

Plugging in a real model

DummyLM swaps for a real model in one line:

python
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))                 # hosted, needs an API key
dspy.configure(lm=dspy.LM("ollama_chat/llama3.2",                # local via Ollama, free
                          api_base="http://localhost:11434", api_key=""))

The provider string format is provider/model-name; the DSPy docs list the exact string for each provider. Everything else in your program stays the same — that portability is a large part of the pitch. See Ollama for running local models.

Optimizers, honestly

An optimizer takes your program, a training set of examples, and a metric, and returns the same program with tuned prompts:

python
trainset = [
    dspy.Example(review="Crisp dosa, fresh chutney.", sentiment="positive").with_inputs("review"),
    dspy.Example(review="Cold, watery sambar.", sentiment="negative").with_inputs("review"),
]   # a real project needs dozens of these, not two

def metric(gold, pred, trace=None):
    return gold.sentiment == pred.sentiment

optimizer = dspy.BootstrapFewShot(metric=metric)
tuned = optimizer.compile(classify, trainset=trainset)

BootstrapFewShot runs your program on training examples, keeps the attempts the metric approves as worked examples in the prompt. MIPROv2 goes further and also rewrites the instruction text, searching combinations. GEPA, the newest in the family, evolves prompts using the model's own reflections on its failures.

No output block for this one, deliberately: optimizer results depend entirely on your model, your data and your metric, and a scripted demo would be theatre. What to expect in practice: real gains on tasks with a clear metric and some dozens of examples — and the same overfitting risks as any other fit-to-data process.

Common mistakes

Optimising a bad metric, efficiently. metric is the definition of success. If it only checks that the output is non-empty, the optimizer will find prompts that produce confident non-empty nonsense. Write the metric as carefully as you would write a test.

A training set of five examples. The optimizer will happily specialise your prompt to those five. Use dozens at minimum, keep a held-out set the optimizer never sees, and compare before-and-after scores on that set only.

Reading reasoning as the model's actual thought process. The reasoning field is generated text that improves answers on average. It is a technique, not a window into the model. Do not build audit trails on it.

Fighting the format instead of declaring it. If you need a list or a number out, say so with types — "review -> scores: list[int]" — instead of writing prompt-like instructions in field descriptions. Typed fields are the same idea as structured output, and DSPy handles the parsing.

Try it yourself

Change the signature to "review -> sentiment, dish" and extend the DummyLM script with a dish key. Run inspect_history again and watch the generated prompt grow a second output field. You have now edited a prompt without touching a prompt — the entire DSPy experience in one small move.

What to learn next

Researcher — Mathematics and papers.

The formal claim

DSPy models an LLM pipeline as a program $P$ composed of modules $m_1, \dots, m_k$, each parameterised by a prompt configuration $\theta_i = (I_i, D_i)$. Where:

  • $I_i$ — the instruction text of module $i$.
  • $D_i$ — its set of few-shot demonstrations.
  • $\mu(x, P(x)) \in [0, 1]$ — a task metric over input $x$ and program output.

Compilation is the search $\arg\max_{\theta} \; \mathbb{E}{x \sim \mathcal{D}}[\mu(x, P\theta(x))]$ over a discrete, non-differentiable space — prompt text — using only forward evaluations. The framing deliberately mirrors training: prompts as parameters, metric as loss, trainset/valset split, overfitting included.

The papers

  • Khattab et al., 2023, DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines (arxiv.org/abs/2310.03714) — the framework paper; grew out of Demonstrate-Search-Predict (Khattab et al., 2022, arxiv.org/abs/2212.14024).
  • Opsahl-Ong et al., 2024, Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs (arxiv.org/abs/2406.11695) — MIPROv2: proposes instruction candidates from data summaries and program traces, then Bayesian-optimises over instruction-demo assignments across stages, addressing credit assignment when only the whole pipeline is scored.
  • Agrawal et al., 2025, GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning (arxiv.org/abs/2507.19457) — evolutionary search where the model reflects in natural language on execution traces to propose mutations, with a Pareto frontier over examples; reports beating GRPO-style RL fine-tuning on several tasks with far fewer rollouts.

Cost and search behaviour

Optimisation cost is dominated by program evaluations: bootstrap-style optimizers cost $O(N \cdot k)$ model calls for $N$ training examples and $k$ modules per pass; MIPROv2 multiplies by the number of candidate configurations trialled (its auto presets cap the budget). Where $N$ is small, variance in metric estimates swamps candidate differences — the formal restatement of "five examples is not a training set". Overfitting to the trainset is empirically real and the standard defence is unchanged from classical ML: a held-out validation set and no peeking.

Where it sits, and open questions

DSPy's bet — the model interface is a compiler target, not an authoring surface — is one answer to prompt brittleness across model versions; re-compiling against a new model replaces re-fiddling. The open questions are live: how far metric-driven search generalises beyond tasks with checkable answers; whether optimised prompts transfer across models or must be recompiled per target (evidence favours recompilation); and how prompt optimisation composes with weight tuning — fine-tuning and prompt search optimise overlapping capacity, and the division of labour between them is not settled. The GEPA line, positioning natural-language reflection as a rival to RL for pipeline improvement, is where that debate is currently sharpest.

What to learn next