Skip to content
Back to student guides
DSPyLLMsFrameworks & agents3 levels84 sectionsCovers DSPy 3.4

The Complete DSPy Guide

Program LLMs instead of prompting them, and let optimisers tune prompts and weights for you. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
28examples

This is part one of three. It covers everything you need to do real work with DSPy, not a teaser. By the end you can install DSPy, point it at a model, declare a task as a Python signature, run it, read the exact prompt DSPy built for you, score the results against a small dataset, and let an optimizer improve the prompt automatically. Mid-level and Senior take the same topics further; nothing here is wasted.

Everything below is written against DSPy 3.4.0 (released 25 September 2026). DSPy moves fast and a lot of blog posts you will find are written against versions that removed the APIs they teach, so where something changed recently this guide says so and shows the current form.

Each section ends with a Try it task. Do them as you go. They take a few minutes, and these ideas only stick once you have watched your own program run, print a wrong answer, and then print a better one after you changed a single line.

What DSPy is, and the problem it solves

DSPy — the name expands to "Declarative Self-improving Python" — is a Python library for programming language models instead of prompting them. You describe a task by declaring its inputs and outputs in Python. DSPy writes the prompt, sends it to the model, parses the reply back into typed Python values, and gives you an object whose attributes are the fields you asked for. Then, if you give it a handful of examples and a function that scores an answer, DSPy's optimizers will rewrite the instructions and pick few-shot examples for you, measured against that score.

To see why that matters, look at what it replaces. The ordinary way to build a feature on a language model is to write a long string. You ask for JSON, you add "respond only with valid JSON", you add an example, you add "do not include markdown fences", you add a note about the edge case a tester found on Tuesday. Then you call the model, get a string back, and write parsing code that handles the three shapes the model actually produces. The prompt becomes a 600-word artefact that nobody dares edit, because nobody can tell which sentence is load-bearing.

The deeper problem is not that the string is ugly. It is that the string is tangled up with everything else. Three separate decisions live inside it: what the task is (classify this ticket into one of five categories), how to attack it (think step by step first, or just answer), and how to talk to this particular model (this one needs the JSON schema repeated, that one does not). Because the three are glued together, you cannot change one without risking the others. Swap gpt-4o-mini for a local Llama and your careful formatting instructions are suddenly wrong. Decide to add chain-of-thought reasoning and you have to rewrite the output contract by hand.

DSPy separates those three decisions into three different objects, and that separation is the whole value proposition.

SIGNATUREwhat the task is
→
MODULEhow to attack it
→
ADAPTERhow to talk to this model
→
LMthe model itself

Because the pieces are separate, each one can be changed on its own. You keep the signature and switch ChainOfThought for Predict to make the program cheaper. You keep both and switch the LM from OpenAI to a model running on your own machine. And — the part people come to DSPy for — you keep all of it and hand the program to an optimizer, which searches over instructions and examples and gives you back a measurably better program without you editing a single word of prompt text.

🧩

Typed outputs

You declare sentiment: bool and get a Python bool, not a string you have to interpret.

🔁

Swappable models

One line changes the provider. The task description does not move.

📈

Measured, not vibed

A metric and a dev set turn "the prompt feels better" into a number you can defend.

🤖

Automatic tuning

Optimizers write the instructions and choose the examples against that number.

One thing to be clear about before you start: DSPy has no command line tool of its own. There is no dspy run. Everything is the Python API, which is why this guide is full of Python rather than shell. The only shell commands you need are for installing things and setting a key.

Try it
  1. Find the longest prompt string in a project you have worked on (or search GitHub for "respond only with valid JSON").
  2. Mark each sentence as task, strategy, or formatting.
  3. Count how many sentences are formatting rather than task.
usually more than half. Every one of those lines is something DSPy's adapter layer generates for you, and the reason the rest of this guide is worth your afternoon.

Prompting versus programming: the shift in mindset

It is worth slowing down on the mental shift, because people who skip it write DSPy that looks like prompt engineering with extra imports, and then wonder why the optimizers do nothing for them.

When you prompt, you are the optimizer. You read an output, decide it is bad, guess which words caused it, edit them, and re-read. Your loop is intuition-driven, your sample size is however many examples you looked at this morning, and your record of what you tried is in your head. It works surprisingly well for the first prototype and stops working the moment the task gets specific enough that two sensible people disagree about what a good answer looks like.

When you program with DSPy, you write down the task once and move your effort to two other places: the data (twenty or fifty examples of input and the answer you want) and the metric (a Python function that looks at an example and a prediction and returns a number). Those two artefacts are the real deliverable. Given them, the library can do the editing loop for you, thousands of times, with a sample size you choose.

This is a genuine trade. You give up the illusion of control over the exact words and you take on the obligation to define success numerically. Beginners resist the second half, and it is the half that pays. If you cannot write a metric, you do not yet know what you are asking the model for — which means you also cannot tell whether yesterday's prompt edit helped.

Programming (DSPy)

  • Task declared once, in Python
  • Prompt text is generated output, not source
  • Progress is a score on a dev set
  • Changing model is one line
  • Improvement loop can run unattended

Prompting by hand

  • Task, strategy and formatting in one string
  • Prompt text is hand-maintained source
  • Progress is a feeling about recent outputs
  • Changing model means rewriting the string
  • Improvement loop needs a human in it

The practical consequence for your first week: resist writing long instructions. Write a short, honest signature, run it, and look at what DSPy produced. Nine times out of ten the output is already reasonable, and the interesting work is building the twenty examples that will tell you whether it is right.

Keep a scratch dev set from day one

Every time you notice an input the program gets wrong, add it to a list with the answer you wanted. Twenty of those is enough to run an optimizer. Collecting them as you go costs nothing; collecting them in a hurry two weeks later costs an afternoon.

Try it
  1. Pick a small task you would normally prompt for: classify a support message, extract a date, summarise in one sentence.
  2. Write five input/answer pairs by hand in a text file. Do not write a prompt.
  3. Write, in one English sentence, the rule that decides whether an answer is correct. That sentence becomes your metric later.
five pairs and a one-sentence rule. That is the raw material DSPy works from, and you now have it before writing any code.

The four nouns you need: signature, module, adapter, prediction

DSPy has a large API surface but a small vocabulary. Four nouns carry most of the weight, and once you can say what each one is, the documentation stops being confusing.

A signature is a declaration of a task's inputs and outputs: their names, their types, and optionally a sentence describing each. It is not a prompt. The shortest form is a string, "question -> answer", where the names before the arrow are inputs and the names after it are outputs. You can add types: "context: list[str], question: str -> answer: str, confidence: float". For anything beyond a line or two you write a class instead, which gives you a place to put the instructions (the class docstring) and per-field descriptions.

A module is a strategy for satisfying a signature. dspy.Predict is the simplest: ask the model, parse the answer. dspy.ChainOfThought adds a reasoning output field before your declared fields, so the model works through the problem before committing to an answer. dspy.ReAct gives the model tools it can call in a loop. All of them take the same signature, which is the point — you change the strategy without rewriting the task. Modules compose: you subclass dspy.Module, build submodules in __init__, and write the logic in forward().

An adapter sits between the signature and the model and does the mechanical work: turn fields into messages the model will understand, then turn the reply back into typed values. The default is ChatAdapter, which uses [[ ## field_name ## ]] markers in the prompt and, if parsing fails, falls back to JSONAdapter. There are also XMLAdapter, TwoStepAdapter and BAMLAdapter. You will rarely touch adapters as a beginner; knowing they exist explains where the prompt's formatting instructions come from.

A prediction is what comes back: a dspy.Prediction whose attributes are your output fields. For "question -> answer" you read pred.answer. With ChainOfThought you also get pred.reasoning.

Beneath those four sits the LM: dspy.LM("openai/gpt-4o-mini"), a configured client holding the model name, sampling parameters, retry count and a cache flag. Model strings follow LiteLLM's naming convention — openai/..., anthropic/..., gemini/..., vertex_ai/..., ollama_chat/..., azure/<deployment> — which is handy, because it means the same string works in LiteLLM itself.

you write this
Signaturefield names, types, instructions. No prompt text
DSPy runs this
ModulePredict, ChainOfThought, ReAct: chooses the strategy
AdapterBuilds the messages, parses the reply, retries the format
the result
LMprovider client, sampling params, cache, history
Predictiontyped attributes: pred.answer, pred.reasoning

Three more words appear constantly and are worth learning now. A predictor is a dspy.Predict instance — the leaf that actually calls the model, and the thing that holds the learnable state (instructions plus few-shot examples). An example is a dspy.Example, one data point. Demos are the few-shot examples attached to a predictor; optimizers mostly work by choosing good ones.

Try it
  1. Write the signature string for: given a product review, return a one-to-five rating and a one-line justification.
  2. Say out loud which noun each part of your answer is.
  3. Now name the module you would start with, and why.
something like "review -> rating: int, justification: str" with dspy.Predict first, because a rating rarely needs visible reasoning and Predict is the cheaper call.

Installing DSPy and checking your setup

DSPy is a pure-Python wheel, so installation is identical on Linux, macOS and Windows, and there is nothing to compile. It supports Python 3.10 through 3.14. Create a virtual environment first — DSPy pulls in a reasonably large dependency tree (openai, litellm, pydantic, diskcache, tenacity and others) and you do not want those in your system Python.

BASH
python -m venv .venv
source .venv/bin/activate          # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -U dspy

If you prefer uv, uv add dspy does the same thing. Pin the version in anything you will come back to, with pip install "dspy==3.4.0", because DSPy has shipped breaking changes in minor releases and a floating dependency means your program can stop working after an unrelated pip install -U.

Install dspy, not dspy-ai

The package was once published as dspy-ai. That name is legacy. Tutorials that tell you to pip install dspy-ai are old enough that other things in them will also be wrong. The package to install, and the module to import, is dspy.

Several features live behind optional extras, because the maintainers moved heavyweight dependencies out of the base install. If you hit a ModuleNotFoundError for numpy or optuna, this is why — nothing is broken, you just need the extra. Quote the brackets; zsh and PowerShell both treat bare square brackets as pattern syntax.

BASH
pip install "dspy[numpy]"      # embeddings, KNN / KNNFewShot, SIMBA
pip install "dspy[optuna]"     # MIPROv2, BootstrapFewShotWithOptuna
pip install "dspy[deno]"       # managed Deno runtime for the sandboxed interpreter
pip install "dspy[mcp]"        # Model Context Protocol tools

Next, credentials. DSPy reads the standard provider environment variables, so you do not pass keys in code: OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, or for Azure the trio AZURE_API_KEY, AZURE_API_BASE and AZURE_API_VERSION. You can also pass api_key= to dspy.LM(...) if you are loading it from a secret manager.

BASH
export OPENAI_API_KEY=sk-...        # Windows PowerShell: $env:OPENAI_API_KEY="sk-..."

Now verify, in two stages. First that the library imports and is the version you think it is:

BASH
python -c "import dspy; print(dspy.__version__)"     # 3.4.0

Then that a real call works end to end. Save this as check.py and run it:

check.py
import dspy

lm = dspy.LM("openai/gpt-4o-mini")      # reads OPENAI_API_KEY
dspy.configure(lm=lm)

print(lm("Say this is a test!"))        # ['This is a test!']
print(dspy.Predict("question -> answer")(question="2+2?").answer)

Two details in that output are worth noticing now. Calling the LM directly with a plain string returns a list of strings, not a string — one entry per sample — so print(lm("...")) shows brackets. And the Predict call returns a prediction, which is why you read .answer off it.

Do not copy lm(messages=[...]) from older material

The OpenAI-style lm(messages=[{"role": ...}]) call is deprecated in 3.4 and removed in 3.5. Some pages of the official docs still show it. For a quick check use lm("hello"), which stays supported in 3.5; when you need explicit system and user messages, use the request object shown later in this guide. To see which of your own calls are affected, run your program with python -W default::DeprecationWarning your_program.py.

If you have no API key, you can still follow every example locally with Ollama. Run ollama run llama3.2 once to pull the model, then point DSPy at it. Expect weaker answers than a frontier model gives, which actually makes the evaluation sections of this guide more interesting, not less.

PYTHON
lm = dspy.LM("ollama_chat/llama3.2", api_base="http://localhost:11434", api_key="")
dspy.configure(lm=lm)
Try it
  1. Create a virtualenv, install dspy, and print dspy.__version__.
  2. Run check.py against either a hosted model or Ollama.
  3. Run it a second time and notice how fast it is. That is the cache, and the next sections explain it.
a version number, a bracketed list, and an answer. The second run returning instantly is correct behaviour, not a bug.

Your first real program, step by step

Let us build something small but complete: a classifier for incoming support messages. The task is to read a message and return a category, an urgency flag, and a one-line reason. We will build it in four passes, each of which teaches one thing.

Pass one: the shortest thing that works. An inline signature with types.

triage.py
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))

triage = dspy.Predict("message -> category: str, urgent: bool, reason: str")

pred = triage(message="My card was charged twice for the same order this morning.")
print(pred.category, pred.urgent, pred.reason)

That already does more than it looks. urgent comes back as a genuine Python bool, because the adapter told the model what shape to produce and parsed it on the way back. You wrote no formatting instruction and no parsing code.

Pass two: constrain the category. A free-form str category is useless for anything downstream, because the model will invent synonyms. Switch to a class-based signature and a Literal, which pins the allowed values.

triage.py
from typing import Literal
import dspy

class Triage(dspy.Signature):
    """Classify an incoming customer support message for routing."""

    message: str = dspy.InputField(desc="The raw message as the customer wrote it")
    category: Literal["billing", "technical", "account", "shipping", "other"] = dspy.OutputField()
    urgent: bool = dspy.OutputField(desc="True only if the customer is blocked right now")
    reason: str = dspy.OutputField(desc="One short sentence justifying the category")

triage = dspy.Predict(Triage)
pred = triage(message="My card was charged twice for the same order this morning.")
print(pred.category, pred.urgent, pred.reason)

Three things changed. The docstring became the task instructions. The Literal became an enumerated constraint the adapter enforces in the prompt and the parser. And each desc= became a per-field hint — note that urgent has a definition attached, which is exactly where a human reviewer's disagreement would otherwise live.

Pass three: change the strategy without touching the task. Suppose urgency is being judged badly. Give the model room to think first:

PYTHON
triage = dspy.ChainOfThought(Triage)
pred = triage(message="My card was charged twice for the same order this morning.")
print(pred.reasoning)      # the model's working, added as an extra output field
print(pred.category, pred.urgent)

The signature did not move. ChainOfThought prepends a reasoning output field, so the model produces its working before the fields you declared, and you get it back on the prediction. This costs more tokens and takes longer. Whether it is worth it is an empirical question, which is the subject of two sections' time.

Pass four: wrap it in a module. Real programs do more than one model call. A module is where that composition lives: submodules in __init__, logic in forward().

triage.py
class SupportTriage(dspy.Module):
    def __init__(self):
        super().__init__()
        self.classify = dspy.ChainOfThought(Triage)
        self.reply = dspy.Predict("message, category -> draft_reply")

    def forward(self, message):
        t = self.classify(message=message)
        draft = self.reply(message=message, category=t.category)
        return dspy.Prediction(
            category=t.category, urgent=t.urgent,
            reason=t.reason, draft_reply=draft.draft_reply,
        )

program = SupportTriage()
out = program(message="I still have not received order 4471, it was due Sunday.")
print(out.category, out.urgent)
print(out.draft_reply)
Call the module, never forward()

Write program(message=...), not program.forward(message=...). Calling the module runs DSPy's surrounding machinery — callbacks, usage tracking, and the trace recording that optimizers depend on. Calling forward directly skips all of it, and the symptom is an optimizer that mysteriously finds nothing to learn from.

Notice what the module gave us: a single object that takes a message and returns everything the rest of the application needs, with two tunable predictors inside it. When we optimize later, both predictors get tuned, and we will not have to think about which.

Try it
  1. Build the four passes above for a task of your own choosing.
  2. Feed it one input that is genuinely ambiguous and read the reasoning field.
  3. Change one desc= string and rerun the ambiguous input. Did the answer move?
a working module, and a first feel for how much a field description changes behaviour. That sensitivity is why we measure rather than guess.

Signatures in depth

Signatures are where beginners get the most value from a little extra care, so here is the fuller picture.

Inline signatures are strings of the form "inputs -> outputs". Names are comma-separated and may carry types after a colon. A name with no type is treated as a string. The names matter: they are shown to the model, so "passage, question -> answer" communicates more than "a, b -> c". Use inline signatures for anything you can describe in one line, and for throwaway experiments.

Class-based signatures earn their extra lines as soon as you want instructions or per-field descriptions. Subclass dspy.Signature, put the task instruction in the docstring, and declare fields with dspy.InputField() and dspy.OutputField(). The docstring is real: it is the instruction block in the prompt, and it is the thing optimizers rewrite.

Output types can be plain Python types, Literal[...] for an enumeration, a list such as list[str], or a Pydantic model when you need nested structure. DSPy also ships its own field types — dspy.Image, dspy.History for conversation turns, dspy.Tool, dspy.Code — which behave as fields the adapter knows how to render.

PYTHON
from pydantic import BaseModel
import dspy

class LineItem(BaseModel):
    description: str
    amount: float

class ParseInvoice(dspy.Signature):
    """Extract the supplier and all line items from an invoice's plain text."""

    invoice_text: str = dspy.InputField()
    supplier: str = dspy.OutputField()
    items: list[LineItem] = dspy.OutputField(desc="One entry per billed line")

parse = dspy.Predict(ParseInvoice)
pred = parse(invoice_text=open("invoice.txt").read())
for item in pred.items:
    print(item.description, item.amount)      # real LineItem objects

pred.items is a list of validated LineItem instances. Pydantic 2.11 or newer is required in DSPy 3.4, and this is where you see the reason: the models in your signature are validated on the way out of the adapter, so a reply with a non-numeric amount fails loudly at the parse step instead of quietly three functions later.

Two habits make signatures work well. First, name the output you actually want, not a container for it — declare supplier: str rather than result: dict, because the field name is a hint to the model and a dict tells it nothing. Second, put the disputable definitions in desc=. Any rule a human reviewer would have to ask about ("does a refund request count as billing?") belongs in a description, where the optimizer can also see and refine it.

Removed and deprecated signature APIs

Three things from older tutorials no longer work. dspy.TypedPredictor and TypedChainOfThought are gone — typed signatures work directly with Predict and ChainOfThought. The prefix=, format= and parser= keyword arguments on InputField and OutputField are deprecated. And duplicate field names across inputs and outputs are now rejected rather than silently merged.

Try it
  1. Convert one of your inline signatures to a class, and move the rules you had in your head into the docstring and desc= fields.
  2. Add a Pydantic model as one output field and confirm you get objects back.
  3. Deliberately declare a Literal with a typo in one option, and see what the parse does.
a class signature you could hand to a colleague as the task's specification, plus a feel for how the enumeration is enforced.

The modules you will actually use

There are more modules than you need. Here is the short list, in the order you are likely to reach for them, with the reason each exists.

dspy.Predict is one call to the model: format the inputs, read back the output fields. Start here always. It is the cheapest, the fastest, and the easiest to debug, and for extraction and classification it is often as good as anything else.

dspy.ChainOfThought is Predict plus an automatically added reasoning output that the model fills in before your fields. Reach for it when the task involves a decision with steps — arithmetic, multi-condition rules, anything where you would want to see the working. The cost is extra output tokens on every call, and for a short classification that can double your bill for no measurable gain. Measure it rather than assuming.

dspy.ReAct gives the model tools and lets it loop: think, call a tool, read the result, think again, until it has an answer. A tool is any Python callable wrapped in dspy.Tool(func), which introspects the name, docstring and argument types to tell the model how to call it. The default iteration cap is max_iters=20.

PYTHON
import dspy

def get_order_status(order_id: str) -> str:
    """Look up the delivery status of an order by its ID."""
    return {"4471": "in transit, due Thursday"}.get(order_id, "no such order")

agent = dspy.ReAct("question -> answer", tools=[dspy.Tool(get_order_status)], max_iters=6)
print(agent(question="Where is order 4471?").answer)

Write the tool's docstring for the model, not for your colleagues: it is the only description the model gets. A tool called lookup whose docstring says "looks things up" will be called at the wrong times, and no amount of optimization fixes that.

dspy.BestOfN and dspy.Refine buy quality with repeated calls. Both take a module, a number N, a reward_fn and a threshold. BestOfN runs the module up to N times and keeps the highest-scoring result; Refine feeds the previous attempt back for improvement. They are the supported replacement for the old dspy.Assert and dspy.Suggest API, which is deprecated and no longer supported — if a tutorial shows you dspy.Assert, it predates this change.

dspy.Parallel and the .batch() method on any module run many inputs concurrently, which is how you evaluate a dev set in seconds instead of minutes.

PYTHON
# Run 200 inputs with 8 worker threads
examples = [dspy.Example(message=m).with_inputs("message") for m in messages]
results = program.batch(examples, num_threads=8)
Two modules you should not learn

dspy.CodeAct and dspy.ProgramOfThought emit a deprecation warning on construction and are scheduled for removal in 3.5. If you need the model to write and run Python, the current module is dspy.RLM, which is a Senior-level topic. Plenty of 2025-era tutorials still teach the deprecated two.

Try it
  1. Run the same signature through Predict and ChainOfThought on ten inputs and compare both the answers and the time taken.
  2. Give ReAct a single tool with a deliberately vague docstring, then rewrite the docstring precisely and rerun.
  3. Use .batch() with num_threads=8 on fifty inputs.
evidence about whether reasoning helps your task, and a direct demonstration that a tool docstring is a prompt.

Seeing what DSPy actually sent

This is the section that converts scepticism into confidence, and it is one line of code. DSPy builds prompts for you, which is wonderful until something goes wrong and you have no idea what the model was asked. So look.

PYTHON
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))
dspy.Predict("message -> category, urgent: bool")(message="My card was charged twice.")

dspy.inspect_history(n=1)

dspy.inspect_history(n) prints the last n LM interactions: the system message DSPy composed, the field markers it used, any few-shot demos attached to the predictor, your input, and the raw reply. Read it once in full. You will see where your docstring went, where each desc= went, and how the output fields were requested. After that, the library stops feeling magical and starts feeling like code.

For programmatic access, every LM keeps its own history:

PYTHON
lm = dspy.LM("openai/gpt-4o-mini")
dspy.configure(lm=lm)
dspy.Predict("question -> answer")(question="2+2?")

entry = lm.history[-1]
print(entry.keys())        # prompt, messages, kwargs, response, outputs, usage, cost, timestamp, model, ...
print(entry["usage"], entry["cost"])

usage gives you token counts and cost an estimate. Be careful with cost: DSPy reports unknown cost as unknown rather than as zero, so do not sum it blindly across a run. A cached call also reports empty usage, which is correct — nothing was spent — but surprising the first time you see a zero.

For tidier accounting, turn on usage tracking and read it off the prediction:

PYTHON
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"), track_usage=True)
pred = dspy.ChainOfThought("question -> answer")(question="Why is the sky blue?")
print(pred.get_lm_usage())

When you want this permanently, rather than at a debugging prompt, DSPy integrates with MLflow tracing via mlflow.dspy.autolog(), which records every module and LM call as a trace you can browse later. Third-party platforms including Langfuse and Arize Phoenix also have DSPy integrations; MLflow is the one the official DSPy docs cover.

Make inspect_history your first debugging move

When a program behaves oddly, the question is almost never "what is wrong with DSPy" and almost always "what did the model see". Printing the history answers that in a second, and in roughly half the cases the answer is visible immediately: a field description says something you did not mean, or the reply was truncated mid-field.

Try it
  1. Run any program and print dspy.inspect_history(n=1). Find your docstring in the output.
  2. Add a desc= to one field and find that too.
  3. Turn on track_usage=True and print pred.get_lm_usage() for a Predict and a ChainOfThought call on the same input.
the exact prompt, and a token-count comparison that puts a number on what reasoning costs you.

Data, metrics and evaluation

Now the part that makes DSPy more than a prompt wrapper. To improve a program you need a dataset and a score.

A data point is a dspy.Example. You build it with keyword arguments and then mark which keys are inputs with .with_inputs(...); everything you did not mark is treated as a label.

PYTHON
import dspy

trainset = [
    dspy.Example(message="My card was charged twice.", category="billing").with_inputs("message"),
    dspy.Example(message="The app crashes when I open settings.", category="technical").with_inputs("message"),
    dspy.Example(message="Order 4471 has not arrived.", category="shipping").with_inputs("message"),
]

The .with_inputs() call is not optional bookkeeping. It is how DSPy knows what to feed the program and what to hold back as the answer, so forgetting it is one of the most common first-day errors — your program receives the label as an input and scores suspiciously well.

A metric is a plain function with the signature metric(example, pred, trace=None) returning a float or a bool. Higher is better.

PYTHON
def category_match(example, pred, trace=None):
    return example.category == pred.category

The trace parameter deserves a word, because it looks like noise and is not. During ordinary evaluation trace is None. During an optimizer's bootstrapping phase, DSPy passes the run's trace, and the convention is to return a stricter boolean in that case — the optimizer is deciding whether a generated example is good enough to keep as a demo, and you want only clean successes. A metric that returns partial credit during evaluation and a hard pass/fail during bootstrapping is a good default:

PYTHON
def quality(example, pred, trace=None):
    correct = example.category == pred.category
    reasoned = len(pred.reason.split()) >= 4
    if trace is not None:          # bootstrapping: be strict
        return correct and reasoned
    return 0.8 * correct + 0.2 * reasoned

DSPy also ships metrics you can use directly: dspy.evaluate.answer_exact_match and answer_passage_match for question answering, and SemanticF1 and CompleteAndGrounded, which use a model to judge the answer. For evaluating retrieval-augmented systems in more depth, the companion toolkit is RAGAS.

With data and a metric, evaluation is three lines:

PYTHON
evaluate = dspy.Evaluate(devset=devset, metric=category_match, num_threads=8, display_progress=True)
result = evaluate(program)

print(result.score)          # a percentage, 0-100, rounded to two decimals
for example, prediction, score in result.results:
    if not score:
        print(example.message, "->", prediction.category)

Two details to internalise. result.score is a percentage, not a fraction: a program that gets two-thirds right scores 66.67, not 0.67. And the per-example detail is in result.results, a list of (example, prediction, score) triples. If you find a tutorial passing return_outputs=True, that argument now raises an error telling you the results are always in results.

Printing the failures, as the loop above does, is the highest-value habit in this guide. The score tells you where you are; the failures tell you what to do next. Ten minutes reading wrong predictions usually beats an hour of prompt speculation.

An empty dev set now raises

In 3.4, Evaluate rejects an empty devset with ValueError: devset must contain at least one example, got an empty devset. If you see that, your data loading returned nothing — check the filter or the file path, not DSPy.

Try it
  1. Turn the five pairs you wrote in section two into dspy.Example objects with .with_inputs().
  2. Write a metric, run dspy.Evaluate, and note the score.
  3. Print every failing example and write one sentence about what they have in common.
a baseline number and a hypothesis. You now have everything an optimizer needs.

Your first optimizer

An optimizer (older documentation calls it a teleprompter) takes your program, your training examples and your metric, and returns a new program with better instructions and better few-shot demos. Your original is not modified, which makes experimenting safe.

Start with dspy.BootstrapFewShot. It works by running your program on the training examples, keeping the runs the metric approves of, and attaching those successful runs to the predictors as demos. In other words, it writes your few-shot examples by using your own program's good days as the teacher.

optimize.py
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))

program = SupportTriage()

baseline = dspy.Evaluate(devset=devset, metric=category_match, num_threads=8)
print("before:", baseline(program).score)

optimizer = dspy.BootstrapFewShot(metric=category_match, max_bootstrapped_demos=4, max_labeled_demos=16)
optimized = optimizer.compile(program, trainset=trainset)

print("after:", baseline(optimized).score)

optimized.save("triage.json")

Three things to understand about what just happened.

It returned a new object. optimized is a separate program; program is untouched. Compare them on the same dev set, and keep the dev set separate from the training set, or you are measuring your own homework.

It cost real money and real time. Compiling runs your program many times. BootstrapFewShot is at the cheap end; the heavier optimizers are not. The official documentation is explicit that a GEPA or MIPROv2 run can cost hundreds of dollars in model calls on a large task. The economics of DSPy assume you compile once, offline, save the result, and serve it many times — not that you optimize per request.

The result is small. optimized.save("triage.json") writes state only: the tuned instructions, the chosen demos, the signature field metadata, and the LM configuration minus the API key, which is never saved. The file is human-readable, belongs in version control, and loads into a fresh instance of the same class:

PYTHON
program = SupportTriage()
program.load("triage.json")

Which optimizer next? The official cheat sheet gives a simple ladder, and it is worth following rather than jumping to the newest name. Starting out, use BootstrapFewShot. If the quality of your demos varies a lot between runs, move to BootstrapFewShotWithRandomSearch (aliased dspy.BootstrapRS), which tries several candidate demo sets and keeps the best. If the problem is that your instructions are wrong rather than your examples, use dspy.COPRO or dspy.GEPA, both of which rewrite instruction text. If both are weak and you have budget, dspy.MIPROv2 tunes instructions and demos together — it needs pip install "dspy[optuna]" and takes auto="light", "medium" or "heavy" to set the budget.

PYTHON
# A step up, once BootstrapFewShot has plateaued
optimizer = dspy.MIPROv2(metric=category_match, auto="light")
optimized = optimizer.compile(program, trainset=trainset, valset=valset)
Save the baseline score in your commit message

An optimized program.json is meaningless without the number it achieved and the dev set it achieved it on. Commit the dev set alongside the program and put the score in the commit message. Six weeks later, that is the only way to know whether a change helped.

Try it
  1. Split your examples into a trainset and a devset that do not overlap.
  2. Score the program, run BootstrapFewShot, score it again, and save the result.
  3. Open triage.json in an editor and find the demos the optimizer chose.
two numbers and a readable artefact. Reading the chosen demos often tells you more about your task than the score does.

Configuration, context and the cache

Three configuration mechanisms cover nearly everything a beginner needs.

dspy.configure(...) sets process-wide defaults: the LM, the adapter, usage tracking, the default thread count, the error budget. It is the thing you call once at the top of your program.

PYTHON
dspy.configure(
    lm=dspy.LM("openai/gpt-4o-mini", temperature=0.7, max_tokens=1000),
    track_usage=True,
    num_threads=8,
)

with dspy.context(...) makes a scoped override that is safe across threads and async tasks. Use it when one part of your program needs a different model — a cheap model for the bulk of the work, a strong one for a hard step.

PYTHON
with dspy.context(lm=dspy.LM("anthropic/claude-sonnet-4-5-20250929")):
    hard = careful_step(question=q)

And predictor.set_lm(other_lm) pins a model to one predictor permanently, which is the cleanest way to express "this step always uses the big model".

configure belongs to one thread

Only the thread (or async task) that first called dspy.configure may call it again. From a worker thread you get RuntimeError: dspy.settings can only be changed by the thread that initially configured it. and from another async task a similar message pointing you at dspy.context. The fix is in the error: configure once at the top, use dspy.context everywhere else.

Now the cache, because it will confuse you before it helps you. DSPy caches LM responses by default, in memory and on disk. The disk cache lives in ~/.dspy_cache unless you set DSPY_CACHEDIR, and its size limit is controlled by DSPY_CACHE_LIMIT (30 GB by default). The key is a hash of the request arguments, excluding credentials.

This is excellent for development: rerunning a script is instant and free. It is also the explanation for the most common "DSPy is broken" report — you call the model three times with the same input and get the same answer three times, even at a high temperature. Nothing is broken; you are reading the cache. Two ways to get genuinely fresh calls:

PYTHON
lm = dspy.LM("openai/gpt-4o-mini", cache=False)          # this LM never caches
lm("Tell me a joke", rollout_id=1, temperature=1.0)      # cached, but a distinct key per rollout_id

The rollout_id form is usually what you want: you still get caching (so a rerun is free) but each rollout_id is a separate cache entry, which is how you collect several independent samples reproducibly.

Try it
  1. Call an LM twice with the same input at temperature=1.0 and confirm the answers are identical.
  2. Add rollout_id=1 and rollout_id=2 and confirm they differ.
  3. Wrap one call in dspy.context with a different model and print dspy.inspect_history(n=2) to see both.
a concrete feel for the cache, which turns a future hour of confusion into a ten-second diagnosis.

Errors you will meet, and how to read them

DSPy's error messages are unusually good: most of them tell you the fix in the message. Here are the ones you will actually hit, and what is really going on behind each.

ValueError: No LM is loaded. The full message suggests dspy.configure(lm=dspy.LM('openai/gpt-4o-mini')), which is the fix. You called a module before configuring a model. In a notebook this usually means the configure cell has not been run since the kernel restarted.

ValueError: LM must be an instance of dspy.BaseLM, not a string. You wrote dspy.configure(lm="openai/gpt-4o-mini"). Wrap the string: dspy.configure(lm=dspy.LM("openai/gpt-4o-mini")). The model string names a model; dspy.LM is the client that talks to it.

dspy.LMAuthError. Missing or invalid key. Check the environment variable name matches the provider in your model string — an anthropic/ model will not read OPENAI_API_KEY. Worth knowing: since 3.3, DSPy normalises provider errors into its own exception hierarchy. Catch dspy.LMError or a subclass, not LiteLLM or OpenAI exceptions, because the provider exception will not reach you. The subclasses are descriptive: LMRateLimitError (with .retry_after), LMTimeoutError, LMServerError, LMInvalidRequestError, LMBillingError, ContextWindowExceededError and others, and dspy.is_retryable_lm_error(e) tells you whether retrying is sensible.

AdapterParseError: Adapter <Name> failed to parse the LM response ... Expected to find output fields in the LM response: [...] This is the characteristic DSPy error and it has three common causes. The reply was truncated because max_tokens is too low — the most frequent cause by a distance, especially with ChainOfThought, where the reasoning eats the budget before your fields are emitted. The model ignored the format, which smaller local models do more often; switching to dspy.JSONAdapter() or dspy.TwoStepAdapter(extraction_model=...) usually fixes it. Or the response was empty. In all three cases, dspy.inspect_history(n=1) shows you which it is in one glance.

dspy.LMRateLimitError. A provider 429. Lower num_threads, lean on the built-in retries (num_retries defaults to 3), and remember that an Evaluate with num_threads=16 against a new API key will trip limits that a single call never would.

dspy.ContextWindowExceededError. Your prompt is too long. For a beginner the usual culprit is an optimizer attaching more demos than you expected, or a long retrieved context. Reduce max_bootstrapped_demos, shorten the context, or move to a larger-context model.

LMConfigurationError for reasoning models. OpenAI's reasoning models (gpt-5 and the o-series) require temperature=1.0 or None and max_tokens of at least 16000, or None. The message says so explicitly. This catches everyone who copies a temperature=0.0, max_tokens=1000 line from a tutorial written for a chat model.

ModuleNotFoundError for numpy or optuna. Those are optional extras now, as covered in the install section: pip install "dspy[numpy]" or "dspy[optuna]".

ValueError: Loading .pkl files can run arbitrary code ... You tried to load a pickled program without opting in. Prefer .json saves. If you must load a pickle, and you trust where it came from, pass allow_pickle=True.

Deprecation warnings. If you see them for lm(messages=...), BaseLM.forward, CodeAct, ProgramOfThought or Image.from_file, you are using an API that goes away in 3.5. Fix them while they are warnings. Run python -W default::DeprecationWarning your_program.py to see the ones Python hides by default.

The three-step debugging loop

Read the error message to the end — DSPy's messages usually contain the fix. Then run dspy.inspect_history(n=1) to see what the model was asked. Only then change your code. Doing those in the other order is how an afternoon disappears.

Try it
  1. Cause a No LM is loaded error on purpose, and read the message.
  2. Set max_tokens=20 on a ChainOfThought call and read the AdapterParseError.
  3. Run a program under -W default::DeprecationWarning and fix anything it reports.
three errors you have now met deliberately, in a calm moment, rather than for the first time under pressure.

Explicit messages, when you need them

Most of the time you let modules talk to the model for you. Occasionally you want to send a raw request with a specific system message — for a sanity check, or to compare a DSPy program against a hand-written prompt. The 3.4 way to do that is an explicit request object, and it is worth seeing once because it is also the migration target for anyone with lm(messages=[...]) in their codebase.

PYTHON
import dspy
from dspy.lm15 import Config, Message, Request

lm = dspy.LM("openai/gpt-4o-mini")
response = lm(Request(
    model=lm.model,
    system="Be concise.",
    messages=(Message.user("What is DSPy?"),),
    config=Config(max_tokens=200),
))
print(response.text)

Two rules apply. The request's model must match the LM's model, and generation options go in Config — the LM's own defaults are not merged into an explicit request, so if you need max_tokens you must state it here. One more subtlety: Config.cache controls the provider's prompt caching, not DSPy's response cache, which is the cache= flag on dspy.LM.

The names lm15, Request and Response come from 3.4's biggest change. DSPy now ships native LM engines and dspy.LM(..., engine="auto") is the default: it prefers the native path where it can and falls back to LiteLLM before inference where it cannot. You can force either with engine="lm15" or engine="litellm". As a beginner you should leave engine alone — the default is correct — but knowing the word helps you read the release notes and the error messages. LiteLLM is still installed and is not deprecated.

Try it
  1. Send one explicit Request with a system message and print response.text.
  2. Omit config=Config(max_tokens=...) and see what changes.
  3. Compare the hand-written prompt's answer with your DSPy program's answer on the same input.
a working escape hatch, and usually a mild surprise at how competitive the generated prompt is.

Putting it all together

Here is the whole beginner arc in one file: a signature, a module, a dataset, a metric, a baseline, an optimizer, a saved artefact, and a reload. Run it, change one thing, run it again.

triage_project.py
from typing import Literal
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini", max_tokens=1000), track_usage=True)


class Triage(dspy.Signature):
    """Classify an incoming customer support message for routing."""

    message: str = dspy.InputField(desc="The raw message as the customer wrote it")
    category: Literal["billing", "technical", "account", "shipping", "other"] = dspy.OutputField()
    urgent: bool = dspy.OutputField(desc="True only if the customer is blocked right now")
    reason: str = dspy.OutputField(desc="One short sentence justifying the category")


class SupportTriage(dspy.Module):
    def __init__(self):
        super().__init__()
        self.classify = dspy.ChainOfThought(Triage)

    def forward(self, message):
        return self.classify(message=message)


def make(message, category, urgent):
    return dspy.Example(message=message, category=category, urgent=urgent).with_inputs("message")


trainset = [
    make("My card was charged twice for order 4471.", "billing", True),
    make("The app crashes when I open settings.", "technical", True),
    make("How do I change the email on my account?", "account", False),
    make("Order 4471 was due Sunday and has not arrived.", "shipping", True),
    make("Do you ship to Riyadh?", "shipping", False),
    make("Just wanted to say the new version is great.", "other", False),
]
devset = [
    make("I was billed in the wrong currency.", "billing", False),
    make("Login fails with a server error.", "technical", True),
    make("Please delete my account.", "account", False),
    make("Where is my parcel?", "shipping", True),
]


def score(example, pred, trace=None):
    correct = example.category == pred.category
    if trace is not None:
        return bool(correct and example.urgent == pred.urgent)
    return 0.7 * correct + 0.3 * (example.urgent == pred.urgent)


if __name__ == "__main__":
    program = SupportTriage()
    evaluate = dspy.Evaluate(devset=devset, metric=score, num_threads=4, display_progress=True)

    before = evaluate(program)
    print("baseline:", before.score)
    for example, prediction, s in before.results:
        if s < 1:
            print("  miss:", example.message, "->", prediction.category, prediction.urgent)

    optimizer = dspy.BootstrapFewShot(metric=score, max_bootstrapped_demos=3, max_labeled_demos=6)
    optimized = optimizer.compile(program, trainset=trainset)

    after = evaluate(optimized)
    print("optimized:", after.score)

    optimized.save("triage.json")

    reloaded = SupportTriage()
    reloaded.load("triage.json")
    print(reloaded(message="You charged me for a plan I cancelled.").category)

Run it and you get a baseline percentage, a list of misses, an optimized percentage, a triage.json on disk, and a prediction from the reloaded program. Then do the experiment that teaches the most: swap ChainOfThought for Predict and run it again. On a short classification like this one, reasoning often buys nothing and costs a third of your tokens — and now you have the number instead of an opinion.

Two notes for anyone building this for a real employer. Because the model string is the only thing naming a provider, the same program runs against an Azure OpenAI deployment in the UAE North region or against a model you host yourself, which is how teams in the Gulf and Egypt usually satisfy a data-residency requirement: dspy.LM("azure/<your-deployment>") with the AZURE_API_BASE of a regional endpoint, or dspy.LM("openai/...", api_base=...) pointed at a self-hosted OpenAI-compatible server. And because program.json contains no API key — dump_state excludes it unconditionally — the optimized artefact is safe to commit.

Try it
  1. Run triage_project.py as written and record both scores.
  2. Swap ChainOfThought for Predict, rerun, and compare score and token usage.
  3. Add four examples drawn from the misses, rerun the optimizer, and see whether the dev score moves.
a reproducible experiment loop. That loop, not any single prompt, is what DSPy is for.

What you can now do, and what comes next

You can install DSPy, point it at a hosted or local model, and confirm the setup works. You can declare a task as an inline or class-based signature with typed outputs including Literal enumerations and Pydantic models. You can choose between Predict, ChainOfThought and ReAct on evidence rather than fashion, and compose them into a custom dspy.Module. You can read the exact prompt DSPy built with inspect_history, and the token cost with get_lm_usage(). You can build a dataset of dspy.Example objects, write a metric that behaves correctly during bootstrapping, score a program with dspy.Evaluate, and read the per-example failures. You can run BootstrapFewShot, measure the improvement, save the result as JSON and load it back. And you can diagnose the dozen errors that account for most beginner time loss.

What you have not touched is everything to do with running this in production. The Mid-level guide picks up at the boundary: how settings, context and threads actually interact; how the adapters format and parse, precisely enough to predict a parse failure; the full optimizer ladder including GEPA and when each one is the right spend; caching strategy and usage accounting for a service rather than a script; serving a program behind FastAPI with dspy.asyncify and streaming with dspy.streamify; and integrating tracing properly rather than printing history. The Senior guide goes further into multi-tenancy, the security model around pickles and code interpreters, cost control at scale, and the upgrade path to 3.5, where the APIs deprecated in 3.4 disappear.

Two things to do before you move on. First, finish a small project end to end, with a committed dev set and a score in the commit message — the habit is worth more than any additional API you could learn. Second, since you will want to see your traces rather than print them, read the MLflow guide next and turn on mlflow.dspy.autolog(); if your team already runs an LLM observability platform, the Langfuse guide covers the same ground from that side. When you start evaluating retrieval quality seriously, RAGAS is the natural companion.

Sources